HIPAA-grade ingestion for 40M patient events per day, with full lineage, revocable consent and reproducible cohort builds.
Where it started
Research teams could not reproduce their own cohorts. Data arrived from fourteen source systems in three FHIR versions plus several bespoke exports, and consent was tracked in a spreadsheet. Withdrawing a patient's consent required a manual hunt through downstream derivatives that nobody could guarantee was complete.
- 01
Standardised ingestion onto FHIR with a bronze-silver-gold lakehouse layering, keeping raw payloads immutable for replay.
- 02
Built a consent engine as a first-class service: consent state is a join key, so revocation propagates to every derivative on the next materialisation and is provably complete.
- 03
Captured column-level lineage across all transformations so any figure can be traced to the source records that produced it.
- 04
Made cohort definitions versioned artefacts pinned to a data snapshot, so a study rebuilt in three years returns identical results.
- Apache Iceberg
- Spark
- dbt
- Kafka
- Snowflake
- AWS
Where it landed
Consent withdrawal now completes automatically with a verifiable audit record. Cohort builds are reproducible by construction, and the platform sustains 40 million events per day with a five-minute freshness SLA.
Their evaluation harness became the standard across our whole ML org. That outlived the engagement by two years.
A short conversation with an engineer, not a sales qualification call. If we're the wrong people for it, we'll say so and point you somewhere better.