The cleanest indicator that an SLO is a slide and not a contract is to ask the on-call engineer what it is. If they shrug, the SLO doesn’t exist. It exists in the deck where it was approved. It does not exist in the system that is meant to be defending it.

Most cloud estates we read have SLOs. Few of those SLOs are operational. The dashboard says 99.95%, but no alert burns when error budget burns. No release is held when the budget is exhausted. No retrospective revisits whether the threshold was right. The SLO is a number that goes in a quarterly review, alongside other numbers that also go in a quarterly review.

This post is about the other kind. The SLOs the rota actually defends. They look different on the page, but the bigger difference is what they cause to happen.

## Rule 1: write the SLO for the worst night, not the best month

The standard mistake is to set the threshold against the system’s median behaviour. 99.95% looks great when the trailing thirty days have been quiet. It looks vicious when a region wobbles, a noisy neighbour deploys a bad release, or the database does a long maintenance window.

A defendable SLO has been pressure-tested against the historical worst-case month. If the system can’t hit 99.95% on the worst week of the last year _with the team you have today_, you have not promised 99.95%. You have promised a heroic recovery — and the people who do the recovering get to leave, eventually, for a system that doesn’t ask that of them.

The negotiation we usually run is the inverse of the one buyers expect. We start with the historical telemetry and back out a threshold the system has empirically defended. Then we explain what work would be needed to defend a tighter one. Then the business decides whether that work is worth doing.

## Rule 2: latency SLOs need a percentile _and_ a tenancy

A p99 latency SLO across an entire service tells you almost nothing. It tells you something about an “average user” who does not exist. The user who matters — the enterprise tenant whose checkout volume is 40% of revenue — is invisible inside it.

> The user who matters is invisible inside a global p99.

Useful latency SLOs are scoped. By route. By tenant tier. By region. By workload class. _Checkout p99 latency, premium tenants, EU region, business hours, < 600ms._ That’s something a runbook can act on. “p99 < 800ms” globally is something a slide can show.

The objection is usually that scoped SLOs are operationally expensive. They are. They are also the only kind that survive a real incident, because they are the only kind whose alert tells you which user-visible thing is broken. Globally averaged SLOs trigger globally averaged confusion.

## Rule 3: error budget burn rate, not error budget remaining

Every SLO doc explains error budget. Most of them stop there, with a static threshold: “alert when 50% of monthly budget is consumed”. That alert fires too late to matter on a fast outage and too early to matter on a slow degradation.

Burn-rate alerts are the version that pages a human at the right time. The pattern, attributed to Google’s SRE workbook, has held up: a fast burn alert (e.g. 14.4× the sustainable rate, on a 1-hour window) catches incidents in progress; a slow burn alert (e.g. 3× on a 6-hour window) catches degradations the team can fix before lunch.

```
# Prometheus / OpenTelemetry alerting rule sketch
- alert: CheckoutFastBurn
  expr: |
    (
      sum(rate(http_requests_total{job="checkout",status=~"5.."}[1h]))
      /
      sum(rate(http_requests_total{job="checkout"}[1h]))
    ) > (14.4 * 0.001)   # 14.4x burn against a 99.9% SLO
  for: 2m
  labels:
    severity: page
    slo: checkout-availability
  annotations:
    summary: "Checkout burning error budget at 14.4x sustainable rate"
    runbook: "https://runbooks.example/checkout-fast-burn"

- alert: CheckoutSlowBurn
  expr: |
    (
      sum(rate(http_requests_total{job="checkout",status=~"5.."}[6h]))
      /
      sum(rate(http_requests_total{job="checkout"}[6h]))
    ) > (3 * 0.001)
  for: 15m
  labels:
    severity: ticket
    slo: checkout-availability
```

The fast-burn alert pages. The slow-burn alert opens a ticket and routes to working hours. Both reference the same SLO. Neither mentions a percentage of monthly budget — by the time that number changes meaningfully, the page has already gone out.

## Rule 4: the runbook is part of the SLO

An SLO without a runbook is a wish. The runbook is what makes the SLO actionable at 02:00, by the person who didn’t write any of it. If you can’t link to the runbook from the alert payload, the alert is incomplete.

Runbooks for SLO-burn alerts have a specific shape. They start with the question the SLO is answering (“is checkout degraded?”), the three or four queries that confirm or deny it, the two or three remediations the on-call has authority to execute, and the escalation path if those don’t move the burn rate.

What runbooks should _not_ contain is the entire architecture of the system. The on-call doesn’t need to learn the system at 02:00. They need to act on it. Architecture goes in the wiki. The runbook is for the next ten minutes.

## Rule 5: SLO violation gates releases

This is the rule that separates real SLO practice from theatre. If error budget is exhausted, releases stop. Not stop in spirit — stop in CI. The deployment pipeline reads the SLO, checks the budget, and refuses the deploy with a message that names the SLO and the path to override.

The override exists. It is a documented action with named approvers. It is also rare, because the team that built the SLO sized the budget against the team’s actual release cadence, and the team understands that exhausting it has a consequence.

The pushback is always the same: “we can’t stop releases, we have a roadmap”. The answer is also always the same: a system whose roadmap is incompatible with its reliability target has chosen its reliability target without thinking. Either the target was wrong (revisit it, with the business) or the engineering investment to defend it has been deferred (revisit _that_, with the business). The SLO is doing its job: forcing the conversation.

We’ve never met a team that regretted gating releases on error budget. We’ve met several that regretted not doing it sooner.

## Rule 6: review the SLO every quarter, like a real document

SLOs drift. Traffic patterns change. The system gets new dependencies. A tenant gets large enough to skew the percentile. The threshold that was right twelve months ago is now either too tight (causing alert fatigue) or too loose (covering up a real degradation).

The quarterly review is short. Pull the four numbers: SLI median, SLI worst week, alert count, alert true-positive rate. If the alert true-positive rate is below 70%, the threshold or the query is wrong. If the worst-week SLI is well within budget, the threshold is too loose. If the alert count is in single digits and the budget was burned, the alerting is wrong, not the SLO.

The output of the review is at most three actions: adjust the threshold, fix the query, or revise the runbook. If nothing changes for two consecutive quarters, the SLO has graduated from “operational” to “stable” and you can review it less often. Most don’t.

## What “the rota defends” means

The shorthand we use internally is “would the rota defend this number”. It’s a useful test. Take the SLO. Imagine the on-call engineer looking at the alert. Imagine them looking at the runbook. Imagine them paging a teammate, or holding a release, or rolling back a config change.

If, at every step, the SLO is helping them act — not slowing them down, not asking them to interpret a slide — then the rota will defend it. They will defend it because the document is on their side. They will revisit it when it stops being on their side. And the number on the dashboard will, finally, mean something.

The version that doesn’t survive contact with the rota dies a quiet death. It stays on the dashboard. It even gets reported up. But it isn’t doing any work. It is decoration in the shape of an engineering artefact.

If the SLOs in your estate are dashboards rather than contracts, the fix is rarely the SLO itself — it’s the rota, the alert routing, and the release gate. We do this work alongside an existing platform team; we don’t take the pager away.
