Blog

Your legacy system is the answer key

A lot of organisations still run their core business on a system installed decades ago. It still processes transactions correctly. What has gone is the understanding of it: the people who built it have moved on, and the people running it inherited it. When those organisations ask for a data warehouse, they usually think the problem is "we can't get at our data". More often the real problem is "we don't know what our data means".

I recently worked on exactly this: a claims warehouse for an insurer whose system of record is a long-lived IBM i platform (case study). These are the principles I'd take to the next one. None of them depend on the industry, the cloud vendor or the legacy technology.

1. Establish what your numbers mean before you move them

Recovering meaning is its own project, and it comes first. Make the legacy application logic searchable, so that "what writes this column, and when?" takes minutes to answer. Work out how each table really behaves by observing the system itself: its job schedules and transaction journals, not the documentation. Write each definition down with its evidence. Moving data you don't understand just gives you a faster, more expensive version of the confusion you started with.

2. Bound the first build, and hold it to production standard

Pick one domain with real business users and prove it end to end, with full verification, access control, infrastructure as code and documented operations. An estate-wide programme delivers nothing until it delivers everything. A throwaway proof of concept is worse: it proves only that data can be copied, and it gives you a cost estimate that's too optimistic, because it skipped the work where the cost actually is.

3. Pick your answer key, and automate the comparison

The system your business already trusts is your answer key. Give every published measure a two-sided test: one side computes it from the warehouse, the other from the live source, and both are calculated at comparison time. Comparing against a captured baseline only proves the warehouse still agrees with its past self. Run the comparison after every load, not once at go-live.

4. Grade your tolerances, and never suppress a discrepancy

Hold money exact to the penny. Match structures on shape and keys where there's no stable amount to compare. Allow a narrow, documented freshness band for a source that moves while you compare. Then add one hard rule: a difference can only be accepted as a dated, attributed business decision, and it stays visible afterwards. If a red light can be quietly switched off, by the end of the quarter the reconciliation is just for show.

5. Size capacity against measured reality, and start small

Measure the real data footprint before choosing compute. Develop the logic on the smallest capacity that works, and scale up only when data volume, not how fast you can iterate, is what's actually limiting you.

6. Move only what changed, but never let cheap outrank correct

Reading the source's change journal is the biggest running-cost lever you have. On a typical day only about two rows in ten thousand change. It's also the most dangerous optimisation, because an expired journal returns "no changes", which looks identical to "nothing changed". Check retention before you believe an empty result, fall back to a full reload when in doubt, and never use the journal for tables that are rebuilt in place.

7. Decouple business decisions from build progress

Some questions are policy, not engineering. For example, does a voided payment still count towards paid totals? Give each open decision a safe default and a clean place to plug the answer in, then keep building. If the build waits for senior people to make decisions, it moves at the speed of their diaries.

8. Make silent failures loud

Loud failures are cheap. Silent ones destroy trust, because by the time anyone notices, every number the system has produced is suspect. List the ways your pipeline could report success while producing wrong data, and build a guard for each. Some real ones: reading a table mid-rebuild, a deploy step that quietly runs nothing, security roles that bind to nobody, and a schedule that shifts at the daylight-saving change.

9. Design the handover before you need it

Write the runbooks, the glossary and the contribution workflow before a second team needs them, and block hand-edits to generated code automatically. A system that only its builders can run recreates the original problem a generation later.

The test

If your warehouse started producing a wrong number tomorrow, how would you find out, and how long would it take? If the honest answer is "when a user complains", you haven't built an answer key yet.

Work with me

Need data you can actually trust?