← Case studies
Migration & observability

Zero-downtime migration & observability rebuild

Bounteous — Knitwell Group commerce platform
98%
ticket noise reduction
161→2
tickets in 3 weeks
$2.5B
GMV platform migrated

The problem

Five core services — OMS, PIM, Offers, Integration, and Payment — plus their microservices needed to move onto a new platform without a minute of downtime, on a commerce system running $2.5B in GMV. Inherited monitoring was noisy enough that real signals were getting buried.

The approach

Led hypercare across 35+ daily tickets while coordinating engineering, operations, and DevOps. Rebuilt the alerting architecture: 54 new health-check and anomaly monitors went in, and 250+ inherited monitors were rationalized down to 150. Authored the Datadog runbooks for the critical microservices — EKS, MongoDB, SQS, Identity, OMS — so the noise reduction would hold.

The impact

Ticket noise on critical services dropped from 161 to 2 within three weeks. Systems stabilized inside the 7–10 week hypercare window, with zero downtime on the migration itself.