Checkout p99 crossed 2s in us-east. Error rate stayed flat.
The 43-minute window
Checkout API p99, one-minute rollup. The rollback cleared queued connections without dropping writes.
Timeline
Declared SEV-2. Priya joined as first responder; agent opened this record.
Miss rate had moved, but pinning cache nodes did nothing. Wrong lead closed.
Connection wait time matched p99. All 30 checkout connections were pinned.
Morning push changed DB_POOL_MAX from 60 to 30 on every checkout task.
Inez reverted the pool parameter. New tasks entered rotation at the prior size.
Pool saturation dropped below 70%; p99 fell through 2s and kept falling.
p99 held at 320ms for five minutes. No lost orders or duplicate charges.
Status moved to resolved. Support received the final customer-impact note.
Close the loop
What we learned
Editable in place. The agent keeps these notes with the incident, even after the channel and call are gone.
Live documentLesson 01 We reviewed the config syntax, not the operational effect. A valid value can still be an unsafe change.
Lesson 02 Our first ten minutes went to cache churn because that graph was familiar. Connection wait time needs to sit beside checkout latency, not three dashboards away.