Case studies
Platform programs, measured honestly
Three illustrative client stories from logistics, healthcare, and financial services. Each shows the starting conditions, the work we did together, and the numbers that changed. Results depend on scope and starting point — we set every baseline with the client before work begins.
Northline Freight Systems: From twelve hand-built environments to a platform that scales with freight volume
Client context
Northline Freight Systems operates a regional freight network with a 90-person engineering organization. Its dispatch, routing, and customer-tracking services had grown on AWS over six years through a series of independent team decisions.
Challenge
Twelve environments had been built by hand and had drifted far apart. Routing service releases required a coordinated weekend window, staging rarely matched production, and a routing outage during peak season exposed the absence of any rollback path. Cloud spend had risen 40% year over year with no allocation by team.
Solution
A five-month embedded delivery pod rebuilt the foundation: a standardized landing zone, Kubernetes runtime for the dispatch and routing services, and environment pipelines driven entirely by Terraform. Progressive delivery with automated rollback replaced the weekend release windows.
Key technical work
- Landing zone redesign: 12 drifted accounts consolidated into a structured organization with shared services
- EKS platform with workload standards, resource policies, and cluster autoscaling tuned for seasonal peaks
- Terraform module library and environment pipelines — dev, staging, and production from one source of truth
- Argo-based progressive delivery with canary analysis and verified automated rollback
- Cost tagging standards and per-team allocation reporting
Results
- Environment provisioning
- 2 weeks → 25 minutes
- Release windows
- Weekend coordinated → any weekday, canary-first
- Rollback time
- None available → under 8 minutes, automated
- Cost visibility
- 0% → 94% of spend allocated to teams
Engagement flow
Cedarwell Health Network: HIPAA-ready platform standards across eleven clinical application teams
Client context
Cedarwell Health Network runs clinical and patient-engagement applications for a multi-hospital network. A compliance review found that each of its eleven application teams had implemented logging, access control, and deployment practices differently.
Challenge
Audit evidence took an average of three weeks to assemble per control. PHI-adjacent services had inconsistent encryption and access patterns. New clinical applications took months to reach production because each team rebuilt compliance plumbing from scratch, and no two services produced comparable audit trails.
Solution
A seven-month program combining a focused assessment with an embedded pod. We standardized the Azure foundation, built compliant service templates, and moved control evidence collection into the platform itself so audit artifacts are produced continuously rather than assembled on demand.
Key technical work
- Assessment of all eleven application estates, risk-ranked against HIPAA Security Rule controls
- Azure landing zone hardening: private networking, managed identities, and customer-managed encryption keys
- Compliant service templates: pre-wired audit logging, access reviews, and PHI handling defaults
- Policy-as-code gates in CI/CD blocking non-compliant deployments before merge
- Continuous evidence pipeline producing auditor-ready artifacts on demand
Results
- Audit evidence assembly
- 3 weeks per control → same-day report
- New service time-to-production
- 4 months → 3 weeks with compliant template
- Non-compliant deploys reaching production
- Unmeasured → blocked at CI, zero in 6 months
- Teams on shared standards
- 3 of 11 → all 11
Engagement flow
Meridian Ledger Group: A reliability and governance program for a payments platform under regulatory scrutiny
Client context
Meridian Ledger Group provides ledger and payment-processing services to mid-market financial institutions. Following two extended outages, its regulators asked for evidence of operational resilience that the company could not readily produce.
Challenge
Incidents averaged 55 minutes to recover because on-call engineers lacked service maps and clear ownership. Dashboards disagreed with each other, alert volume had trained engineers to ignore pages, and the cloud bill was a single undifferentiated number that finance could not reconcile.
Solution
A six-month managed improvement engagement began with SLO definition across the payments core, then rebuilt observability around a shared telemetry schema, redesigned alerting to symptom-based rules, and installed cost governance with quarterly reviews that both engineering and finance attend.
Key technical work
- SLOs and error budgets defined with engineering and product for 23 payments-critical services
- Shared telemetry schema: structured logs, distributed tracing, and RED metrics across all services
- Alerting redesign: 80% reduction in page volume, symptom-based severities, documented on-call routing
- Service catalog with ownership, dependencies, and runbook links for every production service
- Cost governance: tagging enforcement, budgets, anomaly alerts, and per-service allocation
Results
- Mean time to recovery
- 55 minutes → 9 minutes
- Pages per on-call week
- 40+ → under 8, symptom-based
- Services with SLOs and owners
- 0 → all 23 payments-critical services
- Quarterly cloud spend growth
- 9% unexplained → flat, with per-service attribution
Engagement flow
A note on these stories: Northline Freight Systems, Cedarwell Health Network, and Meridian Ledger Group are illustrative client stories that represent the kind of platform programs we deliver. The engagement shapes, technical work, and metrics reflect realistic outcomes; specific results vary with scope, starting conditions, and team capacity.