Which service failures should page an operator?
Page when a user-facing failure requires urgent human action—not simply when an infrastructure metric crosses a threshold. Use service-level objectives to define acceptable outcomes, then alert on rapid error-budget consumption. Route slower deterioration to a ticket. Keep separate alerts for imminent risks, such as capacity exhaustion, when waiting for user-visible failure would be too late. (sre.google)
This guide uses an illustrative API availability objective, not measured production results. The design is provider-independent; the configuration examples use Prometheus. All targets, traffic counts, metric names, and routing labels below are hypothetical and must be adapted and tested before production use.
1. Choose a user journey before choosing a metric
Start with something a user needs to accomplish: retrieve an order, submit a payment, or load a dashboard. “The servers are running” is not the same outcome as “the user can retrieve an order.”
Three terms organize the design:
- Service-level indicator (SLI): the measurement of service behavior.
- Service-level objective (SLO): the target for that measurement over a defined period.
- Error budget: the permitted amount of behavior that does not meet the objective. (sre.google)
For the example, choose this journey:
An authenticated customer retrieves an existing order through the production API.
Write down the scope before selecting a target:
- Which operation and environment are covered?
- Who owns the service and its response procedures?
- Which requests are eligible?
- Where will outcomes be observed?
- What user-impacting failures can that observation point miss?
Google’s SRE guidance distinguishes the desired outcome—the SLI specification—from its measurement implementation. Backend logs, external probes, and client instrumentation have different coverage and cost trade-offs. Backend-only measurements can miss requests that never reach the backend. (sre.google)
Practical check: Ask a product owner to describe a failed journey without using infrastructure terminology. If your proposed SLI cannot detect that failure, revise the measurement or document the gap.
2. Define good events and eligible events explicitly
A request-based availability SLI can be written as:
Availability SLI = good eligible requests / all eligible requests
For this hypothetical API:
- Eligible: production requests for the chosen operation that satisfy its documented authentication and input requirements.
- Good: an eligible request returns the expected successful response before the operation’s agreed timeout.
- Bad: an eligible request fails, times out, or returns an invalid result.
This is a proposed contract, not a universal HTTP classification. Google’s SRE guidance notes that an HTTP success response can still contain incorrect content, while application-specific failures require application-specific definitions. (sre.google)
Create a classification table with the service owner:
| Outcome | Proposed treatment |
|---|---|
| Correct successful response before timeout | Good |
| Server-side failure or dependency timeout | Bad |
| Invalid credentials supplied by the caller | Exclude if outside the defined journey |
| Service incorrectly rejects valid credentials | Bad |
| Caller exceeds a documented quota | Decide explicitly |
| Service throttles otherwise eligible traffic | Decide explicitly; do not silently exclude |
Do not blanket-exclude every 4xx response. Equally, do not count every invalid request as a service failure.
Keep availability and latency objectives distinct unless you intentionally define a combined outcome. An availability target alone does not establish whether responses are fast enough; latency needs its own measurement and threshold. (sre.google)
Practical check: Sample outcomes classified as good, bad, and excluded. Confirm that the implementation matches the written contract.
3. Select a measurement window and calculate the budget
Use this illustrative objective:
At least 99.9% of eligible order-retrieval requests are good over a rolling 30-day window.
The allowed bad-event fraction is:
1 − 0.999 = 0.001 = 0.1%
For a request-based objective, the budget depends on eligible request volume:
Allowed bad events = eligible events × (1 − SLO target)
This follows Google Cloud’s documented error-budget calculation. (docs.cloud.google.com)
Suppose a completed 30-day window contains:
- 2,000,000 eligible requests;
- 1,200 bad requests.
Then:
Allowed bad requests = 2,000,000 × 0.001 = 2,000
Budget consumed = 1,200 / 2,000 = 60%
Remaining allowance = 2,000 − 1,200 = 800 requests
Observed availability = 99.94%
These are hypothetical calculations, not operating results. In a rolling window, both eligible and bad events change as older events leave the window.
Document whether reporting uses a rolling period or a calendar period. Do not display one while silently evaluating another.
Also resist converting this request budget directly into “allowed downtime.” A request-based objective counts affected requests, so an outage during peak traffic has a different impact from an equally long outage during a quiet period.
Practical check: Put the numerator, denominator, target, and window on the same dashboard. A percentage without its request count lacks important context.
4. Use paired windows to detect fast and slow budget burn
Burn rate expresses the observed bad-event fraction relative to the allowed fraction:
Burn rate = observed bad-event fraction / allowed bad-event fraction
For the 99.9% example, an observed 0.8% bad-event fraction gives:
0.008 / 0.001 = 8× burn
A sustained burn rate of 1 consumes the allowance over the objective’s period; higher values consume it faster. Splunk’s current documentation also describes paired long and short windows for burn-rate alerting. (help.splunk.com)
Use the long window to establish significance and the short window to check whether the problem is still active. Require both windows in a pair to exceed the threshold.
Google’s SRE Workbook supplies these starting points for a 30-day objective: (sre.google)
| Route | Long window | Short window | Burn threshold | Bad-event threshold at 99.9% |
|---|---|---|---|---|
| Page | 1 hour | 5 minutes | 14.4× | 1.44% |
| Page | 6 hours | 30 minutes | 6× | 0.6% |
| Ticket | 3 days | 6 hours | 1× | 0.1% |
The fast-page threshold comes from choosing a nominal 2% budget expenditure over one hour:
Burn threshold = 0.02 × 720 hours / 1 hour = 14.4
These time-based budget-spend interpretations assume a representative traffic rate. For request-based objectives with uneven traffic, calculate actual budget consumption from event counts rather than treating the nominal percentages as exact. Red Hat’s explanation makes the steady-traffic assumption explicit. (redhat.com)
In a second hypothetical example, 100,000 hourly requests with 800 bad outcomes produce 8× burn. If the preceding 30 minutes also show 8× burn, the medium-burn pair qualifies for a page. The fast pair does not.
Practical check: Use AND within each pair and OR between paging pairs. Recalculate thresholds when changing the objective period.
5. Implement the ratios without hiding missing data
Assume the example instrumentation exports two counters:
slo_eligible_requests_total;slo_bad_requests_total.
Both must cover the same operation and eligibility definition. Bad requests must be a subset of eligible requests. Initialize both series, including a zero-valued bad counter, rather than exposing the bad counter only after a failure.
Here is an untested illustrative fast-burn rule group, not a complete production configuration:
groups:
- name: illustrative-api-slo
rules:
- record: service:slo_bad_ratio:5m
expr: |
sum by (service) (rate(slo_bad_requests_total[5m]))
/
sum by (service) (rate(slo_eligible_requests_total[5m]))
- record: service:slo_bad_ratio:1h
expr: |
sum by (service) (rate(slo_bad_requests_total[1h]))
/
sum by (service) (rate(slo_eligible_requests_total[1h]))
- alert: ApiAvailabilityFastBurn
expr: |
(service:slo_bad_ratio:1h > 0.0144)
and
(service:slo_bad_ratio:5m > 0.0144)
labels:
severity: page
annotations:
summary: "API availability budget is burning rapidly"
Prometheus recommends applying rate() before aggregation so that counter resets remain detectable. Summing bad-request rates and eligible-request rates before dividing also weights the result by traffic instead of averaging instance percentages. (prometheus.io)
Add corresponding recording rules and alerts for the medium-burn and ticket pairs. Preserve any additional identity labels needed to distinguish environments or objectives consistently across every window.
The example deliberately omits for. That field delays firing while an expression remains continuously active; it does not create another averaging window. (prometheus.io)
Practical check: Treat zero traffic, absent telemetry, and measured zero failures as different states. Do not force all three to display “healthy.”
6. Adapt the design for low-traffic services
Small denominators make percentages volatile. With 20 eligible requests in an hour, one bad request produces a 5% bad-event fraction:
0.05 / 0.001 = 50× burn
That result is mathematically correct, but its operational meaning depends on the journey. One failed, high-value transaction may deserve immediate investigation; one safely retried request may not.
Google’s low-traffic guidance discusses synthetic traffic, combining related services, reducing failure impact, and negotiating objectives that match user needs. Each approach has limitations: synthetic successes can obscure real-user failures, while aggregation can hide a complete failure in one small service. (sre.google)
Choose an explicit policy:
- Keep external probes visible separately from real-user outcomes.
- Aggregate only journeys with a defensible shared purpose or failure domain.
- Review whether a single failed event warrants paging.
- If adding a minimum-event gate, document which outages it can suppress and provide alternative detection.
Do not silently reduce the target to make noisy alerts disappear.
Practical check: Test one failed request, a complete outage, and zero traffic as separate scenarios. A rule that behaves well during busy periods may behave differently overnight.
7. Route alerts to an owner and validate the complete path
Prometheus evaluates alert conditions; Alertmanager handles grouping, deduplication, routing, silences, and inhibition. A severity: page label alone does not configure a paging receiver. (prometheus.io)
For the illustrative service, propose:
- Paging alerts go to the responsible service on-call.
- Ticket alerts go to the owning team’s work queue.
- A page inhibits overlapping ticket notifications for the same service and objective.
- Related notifications are grouped without combining unrelated environments.
Attach enough information to act: the affected journey, objective, current window ratios, request counts, dashboard, runbook, and responsible team. Prometheus annotations can carry descriptive information and runbook references. (prometheus.io)
Validate before enabling production paging:
- Check rule syntax.
- Unit-test healthy traffic, fast burn, slow burn, recovery, counter resets, and missing series.
- Replay representative historical behavior if available.
- Send controlled test alerts through a non-production receiver.
- Confirm grouping, inhibition, escalation, and resolution delivery.
Prometheus supports rule tests using supplied time series and expected alert outcomes through promtool test rules. These tests validate rule behavior, not the external notification delivery path. (prometheus.io)
Avoid three common mistakes: paging directly on every resource threshold, averaging instance percentages without traffic weighting, and treating alert delivery as proof that the measurement is correct. Check user impact, arithmetic, and routing independently. (sre.google)
Finally, review each page: Did it identify a meaningful threat? Could the responder act? Did another alert duplicate it? Adjust the design using that evidence—not a desire for either a perfectly quiet pager or maximum alert coverage. (sre.google)