Data Migration
What we've learned so far
Data migration is the part of a project nobody is thanked for and everybody notices when it goes wrong. It is also where the risk concentrates, because the records being moved are the business: customers, orders, invoices, history that finance and support depend on and that cannot be regenerated if it arrives incomplete.
The work is mostly reconciliation rather than transfer. Moving records is straightforward. Proving that the destination holds what the source did, including the records that were already wrong, is where the time goes. Legacy data is never as described. Fields get repurposed, a status means something different for anything created before a certain year, duplicates exist under several identifiers, and the same customer appears three times in three systems. Those discoveries belong to the business, not to us, and they always take longer to decide than to surface, which is why we raise them in the first fortnight rather than the last.
Automation has made transformation and matching far faster, and it has made one failure mode more likely: a migration that runs cleanly, produces plausible output and is quietly wrong. Plausible is the specific risk with generated logic, because nothing complains. We rehearse the full migration at least twice against production data and compare counts, totals and a sampled set of individual records before anyone commits to a date. The verification is the deliverable. The transfer is the easy half.
What this can involve
ETL and Data Pipelines
Data Migration Services
Database Architecture
How we work
Audit and understand the data
Counts, formats, duplicates and the fields that have quietly been repurposed over the years. Legacy data is never as described, and finding that out early is cheaper than finding out at cutover.
Agree what happens to the awkward records
Duplicates, orphans, statuses that mean something different pre-2019. These decisions belong to the business, and they take longer to settle than to surface, so we raise them in the first fortnight.
Build the mapping and transformation
Field by field, with the rules written down rather than held in code, so anyone can check what was supposed to happen to each record.
Rehearse against production data
At least twice, on the real thing rather than a tidy sample. Each run surfaces something the mapping missed.
Reconcile before anyone commits to a date
Counts, financial totals and a sampled set of individual records checked against the source. A migration that produces plausible output and is quietly wrong is the failure worth engineering against.
Where this isn't the right fit
If the source system has a supported export and the destination has a matching import, use them. Standard tools handle standard cases well, and paying for custom work to move a clean catalogue between two platforms that already speak to each other is money badly spent.
If what you actually need is an ongoing sync between two systems that both stay live, that is integration work rather than migration, and it is built differently. And if nobody on your side can say what the data is supposed to mean, we will stall at the second step. We can find the anomalies. We cannot decide what they should become.

