Skip to content
mohitrai810Public

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Migrato

A schema-aware migration engine that decides when to use deterministic code and when to use a specialized language model.


The Problem

Enterprise migrations rarely involve copying columns from one database to another.

The same business value can appear in completely different forms across legacy systems.

Legacy Employee System
"Works with the platform reliability and infrastructure group"
                              ↓
                    department = "engineering"


Legacy Product System
"Item was serviced and restored to working condition"
                              ↓
                    condition = "refurbished"


Legacy Claims System
"Estimated claim value is INR 2.5L"
                              ↓
                    claim_amount = 250000

These transformations look similar from the outside, but they are fundamentally different.

The first two require semantic understanding.

The last one is a deterministic normalization problem.

Using a large language model for all three adds unnecessary inference cost, latency and unpredictability. Using only rules becomes difficult when the transformation depends on meaning rather than syntax.

Migrato explores a hybrid approach:

Use the smallest reliable mechanism for each transformation.


How Migrato Works

flowchart LR
    A["Legacy Record"] --> B["Transformation Router"]

    B -->|"Predictable"| C["Deterministic Engine"]
    B -->|"Semantic"| D["Qwen3-0.6B + LoRA"]

    C --> E["Validation"]
    D --> E

    E -->|"Valid"| F["Canonical Record"]
    E -->|"Invalid / Uncertain"| G["Larger LLM<br/>Future Fallback"]

    G --> F
Loading

Deterministic Path

Used when the transformation has explicit rules.

"INR 2.5L"        → 250000
"252 thousand"    → 252000
"2.52 lakh"       → 252000

Semantic Path

Used when the value must be inferred from natural language.

"Product was serviced and restored to working condition"

                         ↓

                Qwen3-0.6B + LoRA

                         ↓

             condition = refurbished

The SLM is therefore one component of the migration engine, not a replacement for traditional migration logic.


The Experiment

I wanted to answer three questions:

  1. Can a 0.6B parameter model learn schema-specific semantic transformations?
  2. Does increasing LoRA capacity improve generalization?
  3. When should deterministic code be preferred over model inference?

Migrato evaluates these questions across Employee, Claims and Product migration tasks.

flowchart LR
    A["Qwen3-0.6B<br/>Base<br/><b>62.50%</b>"]
    --> B["LoRA r=8<br/><b>85.12%</b>"]
    --> C["Held-out Test<br/><b>83.50%</b>"]

    B --> D["LoRA r=16<br/><b>84.50%</b>"]
Loading

Fine-tuning improved validation accuracy by 22.62 percentage points while updating approximately 0.38% of model parameters.


Results

Experiment Validation Held-out Test
Qwen3-0.6B Base 62.50% —
LoRA r=8, α=16 85.12% 83.50%
LoRA r=16, α=32 84.50% —
Hybrid V1 89.75% 79.88%

The r=8 adapter was selected because doubling the LoRA rank did not improve validation accuracy.

LoRA r=8

Trainable parameters : ~2.29M
Trainable percentage : ~0.38%

Validation            : 85.12%
Held-out test         : 83.50%
Valid JSON            : 97.00%

The adapter was also loaded onto a fresh base model after training and reproduced the same 85.12% validation accuracy.


The Interesting Failure

The first hybrid result looked excellent:

LoRA Validation        85.12%
        ↓
Hybrid Validation      89.75%

Then I ran the held-out test.

Hybrid Validation      89.75%
        ↓
Hybrid Test            79.88%

Why did adding deterministic logic make the system worse?

The problem was claim_amount.

flowchart LR
    A["2.5 lakh"] -->|"✓"| D["250000"]
    B["250 thousand"] -->|"✓"| D
    C["INR 2.5L"] -->|"✗"| E["2"]
Loading

The original parser achieved 100/100 on validation, but only 46/100 on the first held-out test.

The test introduced a representation family the parser did not understand.

This exposed an important lesson:

Deterministic does not automatically mean robust.
Rule-based transformations need explicit representation coverage too.


Fixing the Transformation Layer

Instead of routing claim_amount back to the language model, the normalization grammar was expanded to handle representation families such as:

2.5 lakh
2.5 lakhs
2.5 lac
2.5L
250 thousand
250k
₹250000
Rs 250000

A new 200-example challenge set was generated and frozen before evaluation.

Amount Normalizer V2

Frozen challenge set       200 / 200   100%
Post-hoc regression check  100 / 100   100%

The regression check uses the previously inspected test set and is intentionally reported separately rather than presented as a new held-out result.


What The Experiments Showed

flowchart TD
    A{"What kind of<br/>transformation?"}

    A -->|"Explicit rule"| B["Deterministic Code"]
    A -->|"Requires meaning"| C["Specialized SLM"]

    B --> D["Fast + predictable"]
    C --> E["Semantic generalization"]

    D --> F["Validate"]
    E --> F

    F -->|"Pass"| G["Canonical Value"]
    F -->|"Fail / uncertain"| H["Fallback Strategy"]
Loading

Fine-tuning mattered.
Qwen3-0.6B improved from 62.50% → 85.12% validation accuracy.

More LoRA capacity did not.
Increasing rank from 8 → 16 slightly reduced accuracy from 85.12% → 84.50%.

Not every migration field should use an LLM.
Numeric normalization was better expressed as deterministic transformation logic once its representation contract was properly defined.


Dataset

Migrato uses synthetic legacy records across three domains.

Domain Transformations
Employee Department, Employment Type
Claims Incident Type, Claim Status, Claim Amount
Product Category, Condition, Availability
Training       4,998
Validation       800
Test             800

Train, validation and test use separate linguistic/template families to test generalization beyond exact training patterns.

Dataset integrity checks cover malformed records, duplicate IDs and cross-split duplicates.


Project Structure

slm-migration-engine/
│
├── data/
│   ├── domains/                 # domain + legacy generators
│   ├── generated/               # dataset splits
│   └── generate_dataset.py
│
├── evaluation/
│   └── results/                 # experiment outputs
│
├── src/
│   ├── amount_normalizer.py     # deterministic transformations
│   ├── transformer.py           # transformation routing
│   └── schemas/
│
├── training/                    # fine-tuning experiments
└── tests/

Tech Stack


What's Next

The current experiments establish the transformation strategy. The next stage is turning it into an end-to-end migration runtime:

  • schema-driven automatic routing
  • dynamic validation
  • batched SLM inference
  • larger-model fallback for uncertain transformations
  • checkpointed workers and retries
  • target-system bulk writes

Core Idea

Migrato started with a question:

How small can the model be if we stop asking the model to solve problems that code already solves better?

The result is a migration architecture where deterministic transformations and specialized language models complement each other instead of competing to handle every field.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages