Learning on Web Dev Open is free for all.

System Design & Performance > Seeing it, and paying for itError budgets, and permission to ship
Phase 06Seeing it, and paying for it320 of 434

Error budgets, and permission to ship

Pick a target below a hundred per cent on purpose, measure against it, and let the remainder decide how bold you are allowed to be this month.

Concept14 minAI pair

An SLO is an indicator, a target and a window: the proportion of requests to this endpoint served successfully in under 300ms, at least 99.9 per cent, measured over 28 days. The point of choosing less than a hundred is that the remainder is a budget you are permitted to spend. Perfect reliability is not a goal, it is a way of never shipping anything.

The budget arbitrates the argument that otherwise recurs weekly. Budget remaining means you can take risks: ship the migration, try the new runtime. Budget exhausted means reliability work takes priority until it recovers. This turns "are we going too fast?" from a matter of temperament into a number both sides already agreed to, which is most of its value.

Alert on burn rate, not on errors. A handful of failures at three in the morning is noise; consuming a fortnight of budget in an hour is an emergency. Two thresholds, a fast burn that pages and a slow burn that opens a ticket, covers nearly everything and produces dramatically fewer false pages than alerting on any error at all, which is how alert fatigue starts and how real incidents get missed.

You should now be able to

  • Write an SLO with an indicator, a target and a window
  • Use a budget to arbitrate between reliability and feature work
  • Alert on burn rate rather than on individual errors
Ask the community

Loading…