A schema-aware migration engine that decides when to use deterministic code and when to use a specialized language model.
Enterprise migrations rarely involve copying columns from one database to another.
The same business value can appear in completely different forms across legacy systems.
Legacy Employee System
"Works with the platform reliability and infrastructure group"
↓
department = "engineering"
Legacy Product System
"Item was serviced and restored to working condition"
↓
condition = "refurbished"
Legacy Claims System
"Estimated claim value is INR 2.5L"
↓
claim_amount = 250000
These transformations look similar from the outside, but they are fundamentally different.
The first two require semantic understanding.
The last one is a deterministic normalization problem.
Using a large language model for all three adds unnecessary inference cost, latency and unpredictability. Using only rules becomes difficult when the transformation depends on meaning rather than syntax.
Migrato explores a hybrid approach:
Use the smallest reliable mechanism for each transformation.
flowchart LR
A["Legacy Record"] --> B["Transformation Router"]
B -->|"Predictable"| C["Deterministic Engine"]
B -->|"Semantic"| D["Qwen3-0.6B + LoRA"]
C --> E["Validation"]
D --> E
E -->|"Valid"| F["Canonical Record"]
E -->|"Invalid / Uncertain"| G["Larger LLM<br/>Future Fallback"]
G --> F
Used when the transformation has explicit rules.
"INR 2.5L" → 250000
"252 thousand" → 252000
"2.52 lakh" → 252000
Used when the value must be inferred from natural language.
"Product was serviced and restored to working condition"
↓
Qwen3-0.6B + LoRA
↓
condition = refurbished
The SLM is therefore one component of the migration engine, not a replacement for traditional migration logic.
I wanted to answer three questions:
- Can a 0.6B parameter model learn schema-specific semantic transformations?
- Does increasing LoRA capacity improve generalization?
- When should deterministic code be preferred over model inference?
Migrato evaluates these questions across Employee, Claims and Product migration tasks.
flowchart LR
A["Qwen3-0.6B<br/>Base<br/><b>62.50%</b>"]
--> B["LoRA r=8<br/><b>85.12%</b>"]
--> C["Held-out Test<br/><b>83.50%</b>"]
B --> D["LoRA r=16<br/><b>84.50%</b>"]
Fine-tuning improved validation accuracy by 22.62 percentage points while updating approximately 0.38% of model parameters.
| Experiment | Validation | Held-out Test |
|---|---|---|
| Qwen3-0.6B Base | 62.50% | — |
| LoRA r=8, α=16 | 85.12% | 83.50% |
| LoRA r=16, α=32 | 84.50% | — |
| Hybrid V1 | 89.75% | 79.88% |
The r=8 adapter was selected because doubling the LoRA rank did not improve validation accuracy.
Trainable parameters : ~2.29M
Trainable percentage : ~0.38%
Validation : 85.12%
Held-out test : 83.50%
Valid JSON : 97.00%
The adapter was also loaded onto a fresh base model after training and reproduced the same 85.12% validation accuracy.
The first hybrid result looked excellent:
LoRA Validation 85.12%
↓
Hybrid Validation 89.75%
Then I ran the held-out test.
Hybrid Validation 89.75%
↓
Hybrid Test 79.88%
Why did adding deterministic logic make the system worse?
The problem was claim_amount.
flowchart LR
A["2.5 lakh"] -->|"✓"| D["250000"]
B["250 thousand"] -->|"✓"| D
C["INR 2.5L"] -->|"✗"| E["2"]
The original parser achieved 100/100 on validation, but only 46/100 on the first held-out test.
The test introduced a representation family the parser did not understand.
This exposed an important lesson:
Deterministic does not automatically mean robust.
Rule-based transformations need explicit representation coverage too.
Instead of routing claim_amount back to the language model, the normalization grammar was expanded to handle representation families such as:
2.5 lakh
2.5 lakhs
2.5 lac
2.5L
250 thousand
250k
₹250000
Rs 250000
A new 200-example challenge set was generated and frozen before evaluation.
Amount Normalizer V2
Frozen challenge set 200 / 200 100%
Post-hoc regression check 100 / 100 100%
The regression check uses the previously inspected test set and is intentionally reported separately rather than presented as a new held-out result.
flowchart TD
A{"What kind of<br/>transformation?"}
A -->|"Explicit rule"| B["Deterministic Code"]
A -->|"Requires meaning"| C["Specialized SLM"]
B --> D["Fast + predictable"]
C --> E["Semantic generalization"]
D --> F["Validate"]
E --> F
F -->|"Pass"| G["Canonical Value"]
F -->|"Fail / uncertain"| H["Fallback Strategy"]
Fine-tuning mattered.
Qwen3-0.6B improved from 62.50% → 85.12% validation accuracy.
More LoRA capacity did not.
Increasing rank from 8 → 16 slightly reduced accuracy from 85.12% → 84.50%.
Not every migration field should use an LLM.
Numeric normalization was better expressed as deterministic transformation logic once its representation contract was properly defined.
Migrato uses synthetic legacy records across three domains.
| Domain | Transformations |
|---|---|
| Employee | Department, Employment Type |
| Claims | Incident Type, Claim Status, Claim Amount |
| Product | Category, Condition, Availability |
Training 4,998
Validation 800
Test 800
Train, validation and test use separate linguistic/template families to test generalization beyond exact training patterns.
Dataset integrity checks cover malformed records, duplicate IDs and cross-split duplicates.
slm-migration-engine/
│
├── data/
│ ├── domains/ # domain + legacy generators
│ ├── generated/ # dataset splits
│ └── generate_dataset.py
│
├── evaluation/
│ └── results/ # experiment outputs
│
├── src/
│ ├── amount_normalizer.py # deterministic transformations
│ ├── transformer.py # transformation routing
│ └── schemas/
│
├── training/ # fine-tuning experiments
└── tests/
The current experiments establish the transformation strategy. The next stage is turning it into an end-to-end migration runtime:
- schema-driven automatic routing
- dynamic validation
- batched SLM inference
- larger-model fallback for uncertain transformations
- checkpointed workers and retries
- target-system bulk writes
Migrato started with a question:
How small can the model be if we stop asking the model to solve problems that code already solves better?
The result is a migration architecture where deterministic transformations and specialized language models complement each other instead of competing to handle every field.