52 min
Customer impact
Detected after 4 min
Postmortem PM-112
A blameless review of the 14 September incident, its causes and the follow-up actions.
Impact
14 September, 09:12 to 10:04 UTC.
52 min
Customer impact
Detected after 4 min
21%
Checkouts failed at peak
3 410 orders retried
38%
Monthly error budget used
SLO 99.9%
0
Orders lost
All retried or refunded
Source: incident record PM-112
Timeline
All times UTC, 14 September.
09:12
Release rolls out to all regions at once.
09:16
Checkout error rate passes 5%.
09:24
Commander assigned and status page updated.
09:31
On-call reverts to release 4.17.
09:48
Error rate back under 0.1%.
10:04
Incident closed and review scheduled.
The takeawayDetection was fast. The 15 minutes before the rollback were not.
Blast radius
Failed checkouts during the incident.
| Europe West | 1,480 orders |
|---|---|
| North America | 1,120 orders |
| Europe North | 540 orders |
| Asia Pacific | 270 orders |
Source: checkout logs, 09:12-10:04 UTC
Root cause
The code was fine; the rollout was not.
Trigger
The pool size dropped from 50 to 5. Checkout exhausted it within minutes under normal traffic.
The release went to all regions in one step.
The setting was not reviewed as a config change.
The rollback needed a manual approval.
The takeawayAny change that reaches every region at once can take checkout down.
Follow-up actions
Tracked in the PM-112 board.
| Done | Owner | Due | |
|---|---|---|---|
| Roll out releases region by region with a canary | Platform team | 15 Oct | |
| Move pool settings to reviewed configuration | Checkout team | 30 Sep | |
| Allow on-call to roll back without approval | SRE lead | 25 Sep | |
| Alert on connection-pool saturation | Observability | 30 Sep | |
| Replay this incident in the next game day | SRE lead | 20 Nov |