Hi! I'm Abdul Afolabi

Data Engineer

I build data pipelines and warehouses that businesses can trust, with every published number reconciled back to its source.

  1. 01 · Sources

    It starts with messy data

    Legacy mainframes, SaaS platforms, CRMs and APIs. Decades of data that nobody fully understands any more.

  2. 02 · Ingest

    Move only what changed

    Read-only, journal-driven change capture. Cheap when it can be, a full reload whenever an empty result can’t be trusted.

  3. 03 · Bronze

    Land it exactly as it was

    A byte-faithful copy with load metadata. The raw evidence is always there to check against.

  4. 04 · Silver

    Clean, typed, decoded

    Every source is translated into one consistent language, kept separate from modelling the business.

  5. 05 · Gold

    Modelled for the business

    Kimball star schemas: facts, conformed dimensions and governed measures that slice the same way everywhere.

  6. 06 · Reconcile

    Every number, proven

    After every load, business measures are recomputed and checked against the live source. An unexplained red never gets promoted.

  7. 07 · Insight

    Data you can trust

    Dashboards and decisions built on numbers that reconcile all the way back to source.

Skip animation
Portrait of Abdul Afolabi

About me

From analyst to engineer

I spent my early career as a data analyst, building the dashboards everyone relied on and finding out the hard way how often the numbers underneath them were wrong. That's why I moved into data engineering: to fix the problem at the source.

Today I build cloud data platforms on Microsoft Fabric, Snowflake, Azure and AWS, most recently for a US insurer, migrating a decades-old system into a governed warehouse where every number is reconciled back to source.

What I build

Ingestion & change capture

Read-only extraction from legacy and SaaS sources, journal-based CDC with safe fallbacks, idempotent and resumable loads.

Lakehouse & modelling

Medallion layers on Delta, Kimball star schemas, conformed dimensions and governed semantic models.

Data quality you can prove

Two-sided reconciliation against the source after every load, load audit and lineage, guards for silent failures.

Platform as code

Terraform, containerised jobs, CI/CD with secret scanning, spec-driven generation and deploy provenance.

Featured case study

A reconciled claims warehouse beside a legacy mainframe

Lakehouse for a medical liability insurer whose system of record is a decades-old IBM i platform that can't be switched off. Every load is proven against the live source.

  • ~80% lower platform capacity cost after moving silver and gold to incremental loads on scheduled compute
  • ~100M rows across ~160 source tables, refreshed daily in ~90 minutes
  • ~70 business measures reconciled against the live source after every load
  • Daily change capture moves ~0.02% of rows instead of a full reload
  • Infrastructure, semantic model and security roles all deployed from code
  • Microsoft Fabric
  • PySpark
  • Delta Lake
  • Python
  • Terraform
  • Azure Container Apps
  • GitLab CI
  • Power BI / TMDL
  • IBM i (DB2)

Read the case study

Experience

Full resume →
  1. Jun 2026 – Sep 2026

    Data Engineer (Contract Consultant)

    Exequt · client: US medical professional liability insurer

    • Led the technical audit and recovery of an AI-assisted migration from a legacy IBM i (DB2) system to a Microsoft Fabric warehouse, deployed entirely through CI/CD.
    • Modelled the claims domain (claims, payments, reserves, recoveries, reinsurance, written premium) for actuarial and finance users, reconciled to the legacy system after every load.
    • Cut Fabric capacity cost by ~80% by moving the always-on silver and gold layers to incremental loads on scheduled compute.
  2. Jul 2022 – May 2026

    Data Engineer

    Emeria, London

    • Replaced Salesforce as the reporting source of truth for property data with a Snowflake warehouse consolidating Qube MRI, Reapit, DataStation and on-premises SQL Server, with Power BI reporting that saved 8 hours a week of manual processing.
    • Built reusable, parameterised ELT with Fivetran metadata, dbt vars, macros and artifacts, and Python UDFs, on a medallion architecture across AWS and Microsoft stacks.
    • Cut pipeline processing time by ~80% through table redesign, clustering, materialized views, parallel jobs and better orchestration.
  3. Jul 2021 – Jul 2022

    Data Analyst Lead

    Modulr Finance, UK

    • Worked with data engineering on pipelines feeding management-information datasets; owned data governance through a CRM warehouse migration.
  4. Feb 2019 – Jul 2021

    Data Analyst Specialist

    Tutora, UK

    • Automated manual data collection and reporting; built Python + Power BI customer segmentation.

Domains

  • Insurance: claims, reserves, reinsurance
  • Property management
  • Fintech

Languages

  • Python
  • SQL
  • PySpark
  • DAX
  • Jinja (dbt)

Platforms

  • Microsoft Fabric
  • Azure
  • Snowflake
  • AWS
  • SQL Server
  • IBM i (DB2)

Data

  • dbt
  • Fivetran
  • Delta Lake
  • Kimball modelling
  • Change data capture
  • Data governance

Delivery

  • Terraform
  • Docker
  • GitLab CI
  • Power BI / TMDL
  • Claude agents & MCP

Work with me

Need data you can actually trust?