cartwheel / incident record
Last agent update 10:12 UTCSystems normal

INC-2114 / SEV-2

Checkout latency regression

Resolved. A connection-pool config change pushed checkout p99 from 320ms to 8.4s. Rolled back with no data loss.

Started18 Jul, 09:14 UTC
Duration43 minutes
Customer impact18.7% of checkouts over 3s

The 43-minute window

Checkout API p99, one-minute rollup. The rollback cleared queued connections without dropping writes.

8.4s peak p99
degradedbaseline
Checkout p99 latency during incident INC-2114 Latency rises after the configuration deploy, peaks at 8.4 seconds, and returns to a 320 millisecond baseline after rollback. 8.0s5.0s2.0s0.3s CONFIG DEPLOY · 09:08 ROLLBACK · 09:36 09:0809:2009:3209:4409:57
320ms → 8.4s after pool capacity was halved0 failed writes · 0 data loss

Timeline

Latency alert fired

Checkout p99 crossed 2s in us-east. Error rate stayed flat.

Mara took incident command

Declared SEV-2. Priya joined as first responder; agent opened this record.

Cache churn looked guilty

Miss rate had moved, but pinning cache nodes did nothing. Wrong lead closed.

Turning point: pool saturation found

Connection wait time matched p99. All 30 checkout connections were pinned.

Config delta confirmed

Morning push changed DB_POOL_MAX from 60 to 30 on every checkout task.

Turning point: rollback started

Inez reverted the pool parameter. New tasks entered rotation at the prior size.

Queues draining

Pool saturation dropped below 70%; p99 fell through 2s and kept falling.

Service recovered

p99 held at 320ms for five minutes. No lost orders or duplicate charges.

All clear

Status moved to resolved. Support received the final customer-impact note.

Close the loop

2 of 6 closed

What we learned

Editable in place. The agent keeps these notes with the incident, even after the channel and call are gone.

Live document

Lesson 01 We reviewed the config syntax, not the operational effect. A valid value can still be an unsafe change.

Lesson 02 Our first ten minutes went to cache churn because that graph was familiar. Connection wait time needs to sit beside checkout latency, not three dashboards away.