Hi! I'm Abdul Afolabi
Data Engineer
I build data pipelines and warehouses that businesses can trust, with every published number reconciled back to its source.
01 · Sources
It starts with messy data
Legacy mainframes, SaaS platforms, CRMs and APIs. Decades of data that nobody fully understands any more.
02 · Ingest
Move only what changed
Read-only, journal-driven change capture. Cheap when it can be, a full reload whenever an empty result can’t be trusted.
03 · Bronze
Land it exactly as it was
A byte-faithful copy with load metadata. The raw evidence is always there to check against.
04 · Silver
Clean, typed, decoded
Every source is translated into one consistent language, kept separate from modelling the business.
05 · Gold
Modelled for the business
Kimball star schemas: facts, conformed dimensions and governed measures that slice the same way everywhere.
06 · Reconcile
Every number, proven
After every load, business measures are recomputed and checked against the live source. An unexplained red never gets promoted.
07 · Insight
Data you can trust
Dashboards and decisions built on numbers that reconcile all the way back to source.

About me
From analyst to engineer
I spent my early career as a data analyst, building the dashboards everyone relied on and finding out the hard way how often the numbers underneath them were wrong. That's why I moved into data engineering: to fix the problem at the source.
Today I build cloud data platforms on Microsoft Fabric, Snowflake, Azure and AWS, most recently for a US insurer, migrating a decades-old system into a governed warehouse where every number is reconciled back to source.
What I build
Ingestion & change capture
Read-only extraction from legacy and SaaS sources, journal-based CDC with safe fallbacks, idempotent and resumable loads.
Lakehouse & modelling
Medallion layers on Delta, Kimball star schemas, conformed dimensions and governed semantic models.
Data quality you can prove
Two-sided reconciliation against the source after every load, load audit and lineage, guards for silent failures.
Platform as code
Terraform, containerised jobs, CI/CD with secret scanning, spec-driven generation and deploy provenance.
Featured case study
A reconciled claims warehouse beside a legacy mainframe
Lakehouse for a medical liability insurer whose system of record is a decades-old IBM i platform that can't be switched off. Every load is proven against the live source.
- ~80% lower platform capacity cost after moving silver and gold to incremental loads on scheduled compute
- ~100M rows across ~160 source tables, refreshed daily in ~90 minutes
- ~70 business measures reconciled against the live source after every load
- Daily change capture moves ~0.02% of rows instead of a full reload
- Infrastructure, semantic model and security roles all deployed from code
- Microsoft Fabric
- PySpark
- Delta Lake
- Python
- Terraform
- Azure Container Apps
- GitLab CI
- Power BI / TMDL
- IBM i (DB2)
Experience
Full resume →- Jun 2026 – Sep 2026
Data Engineer (Contract Consultant)
Exequt · client: US medical professional liability insurer
- Led the technical audit and recovery of an AI-assisted migration from a legacy IBM i (DB2) system to a Microsoft Fabric warehouse, deployed entirely through CI/CD.
- Modelled the claims domain (claims, payments, reserves, recoveries, reinsurance, written premium) for actuarial and finance users, reconciled to the legacy system after every load.
- Cut Fabric capacity cost by ~80% by moving the always-on silver and gold layers to incremental loads on scheduled compute.
- Jul 2022 – May 2026
Data Engineer
Emeria, London
- Replaced Salesforce as the reporting source of truth for property data with a Snowflake warehouse consolidating Qube MRI, Reapit, DataStation and on-premises SQL Server, with Power BI reporting that saved 8 hours a week of manual processing.
- Built reusable, parameterised ELT with Fivetran metadata, dbt vars, macros and artifacts, and Python UDFs, on a medallion architecture across AWS and Microsoft stacks.
- Cut pipeline processing time by ~80% through table redesign, clustering, materialized views, parallel jobs and better orchestration.
- Jul 2021 – Jul 2022
Data Analyst Lead
Modulr Finance, UK
- Worked with data engineering on pipelines feeding management-information datasets; owned data governance through a CRM warehouse migration.
- Feb 2019 – Jul 2021
Data Analyst Specialist
Tutora, UK
- Automated manual data collection and reporting; built Python + Power BI customer segmentation.
Domains
- Insurance: claims, reserves, reinsurance
- Property management
- Fintech
Languages
- Python
- SQL
- PySpark
- DAX
- Jinja (dbt)
Platforms
- Microsoft Fabric
- Azure
- Snowflake
- AWS
- SQL Server
- IBM i (DB2)
Data
- dbt
- Fivetran
- Delta Lake
- Kimball modelling
- Change data capture
- Data governance
Delivery
- Terraform
- Docker
- GitLab CI
- Power BI / TMDL
- Claude agents & MCP
Writing
All posts →Your legacy system is the answer key
Nine principles for building a data warehouse beside an operational system you can't switch off, and proving every number against it.

Data project process documentation
Data project process documentation Having a standard process around the business is nice. However, having it documented is nicer as it helps the business reduce downtime when…

Framework for executing data projects
Framework for executing data projects Following a framework when executing a data project could help reduce downtime, plan ahead, and have an established process in the…
Earlier work: data analysis
Before data engineering I spent several years as a data analyst. These projects are from 2021–2022.


