Postmortem PM-112

Checkout failed for 52 minutes. Here is what we change.

A blameless review of the 14 September incident, its causes and the follow-up actions.

Incident commander Engineering review · 21 September

Impact

One in five checkouts failed for 52 minutes

14 September, 09:12 to 10:04 UTC.

52 min

Customer impact

Detected after 4 min

21%

Checkouts failed at peak

3 410 orders retried

38%

Monthly error budget used

SLO 99.9%

0

Orders lost

All retried or refunded

Source: incident record PM-112

Postmortem PM-112 / blameless02 / 07

Timeline

Detected in 4 minutes, resolved in 52

All times UTC, 14 September.

  1. 09:12

    Deploy 4.18

    Release rolls out to all regions at once.

  2. 09:16

    Alert fires

    Checkout error rate passes 5%.

  3. 09:24

    Incident declared

    Commander assigned and status page updated.

  4. 09:31

    Rollback starts

    On-call reverts to release 4.17.

  5. 09:48

    Traffic recovers

    Error rate back under 0.1%.

  6. 10:04

    Resolved

    Incident closed and review scheduled.

The takeawayDetection was fast. The 15 minutes before the rollback were not.

Postmortem PM-112 / blameless03 / 07

Blast radius

Every region failed, the busiest the most

Failed checkouts during the incident.

Every region failed, the busiest the most
Europe West1,480 orders
North America1,120 orders
Europe North540 orders
Asia Pacific270 orders

Source: checkout logs, 09:12-10:04 UTC

Postmortem PM-112 / blameless04 / 07

Root cause

A config change reached every region at once

The code was fine; the rollout was not.

DeployDetectMitigateLearn

Trigger

A connection-pool setting shipped inside the release

The pool size dropped from 50 to 5. Checkout exhausted it within minutes under normal traffic.

→

No canary

The release went to all regions in one step.

Config in the image

The setting was not reviewed as a config change.

Slow rollback

The rollback needed a manual approval.

The takeawayAny change that reaches every region at once can take checkout down.

Postmortem PM-112 / blameless05 / 07

Follow-up actions

Five actions, each with an owner

Tracked in the PM-112 board.

Five actions, each with an owner
DoneOwnerDue
Roll out releases region by region with a canaryPlatform team15 Oct
Move pool settings to reviewed configurationCheckout team30 Sep
Allow on-call to roll back without approvalSRE lead25 Sep
Alert on connection-pool saturationObservability30 Sep
Replay this incident in the next game daySRE lead20 Nov
Postmortem PM-112 / blameless06 / 07

Sources

  1. 02One in five checkouts failed for 52 minutesIncident record PM-112
  2. 03Detected in 4 minutes, resolved in 52Incident channel log
  3. Deployment history, 14 September
  4. 04Every region failed, the busiest the mostCheckout logs, 09:12-10:04 UTC