Correlated tracing and automated reconciliation across a payment estate, cutting mean time to restore from hours to minutes.
Where it started
A failed payment could involve seven internal services and three external providers, and nobody could follow a single transaction end to end. Incidents were diagnosed by reading logs in parallel across teams, and daily reconciliation breaks were investigated manually days after the fact.
- 01
Introduced a propagated correlation identifier across every service and provider boundary, including the ones that required a proxy to inject it.
- 02
Built a reconciliation engine that matches ledger, provider and internal event streams continuously rather than nightly.
- 03
Defined SLOs per payment corridor, with alerts tied to user-visible symptoms instead of resource metrics.
- 04
Ran game days against injected provider failures until the runbooks matched reality.
- OpenTelemetry
- Grafana
- Kafka
- Kubernetes
- Terraform
Where it landed
Any transaction can now be followed end to end in a single view. Median time to restore fell from just over three hours to under twenty minutes, and reconciliation breaks surface within minutes rather than at the next day's close.
A short conversation with an engineer, not a sales qualification call. If we're the wrong people for it, we'll say so and point you somewhere better.