Quick answer: Where should you look first?
Start with user-facing latency, traffic and errors at the API boundary, not individual server CPU charts. Identify affected routes and regions, then inspect representative slow traces and their correlated logs. Check recent changes, dependency delays and resource queues before selecting a reversible mitigation. Google’s SRE guidance groups latency, traffic, errors and saturation as the four golden signals. (sre.google)
This runbook uses Prometheus metrics and Grafana Loki logs, with OpenTelemetry-instrumented traces viewed in your existing trace backend. The queries are illustrative, read-only examples—not tested against your environment. Metric names, labels and log fields must be mapped to your actual telemetry.
1. Confirm user impact and establish the incident window
Record the first known impact time in UTC. Open the same time range across dashboards, logs, traces and deployment history, including a healthy period before the increase.
Check:
- Latency: Which routes have degraded? Inspect tail latency as well as the median.
- Traffic: Has demand increased, decreased or shifted between routes?
- Errors: Are requests failing, timing out or being rejected?
- Scope: Is the problem isolated to a region, release, instance group or dependency?
Separate successful-request latency from failed-request latency where instrumentation permits. A fast error response is not evidence of healthy service. Compare application measurements with gateway, synthetic or client telemetry when available; these measure different parts of the request path. (sre.google)
Write an initial impact statement:
“Requests to the affected route are exceeding its established latency objective in the affected region. Other routes remain under investigation.”
Use your service’s actual objective and severity policy. If no latency objective exists, report the observed change against a comparable healthy period without inventing a threshold.
Practical check: Confirm telemetry is arriving. Treat an empty graph as “unknown” until collection and query selection have been checked.
2. Use metrics to identify the affected request population
Start with route-level latency, then narrow by region or release. Avoid starting with every instance separately.
Assume, for illustration:
http_request_duration_secondsis a classic histogram measured in seconds.service,environment,routeandregionare existing labels.http_requests_totalis a counter with astatus_codelabel.routecontains normalized route templates, not individual request URLs.
Estimate p95 latency by route and region:
histogram_quantile(
0.95,
sum by (le, route, region) (
rate(http_request_duration_seconds_bucket{
service="checkout-api",
environment="prod"
}[5m])
)
)
For classic histograms, retain le when aggregating buckets. Apply rate() before aggregation so counter resets remain detectable. Native histograms require a different expression; do not assume _bucket series exist. (prometheus.io)
Check traffic alongside latency:
sum by (route, region) (
rate(http_requests_total{
service="checkout-api",
environment="prod"
}[5m])
)
Inspect server-error traffic:
sum by (route, region) (
rate(http_requests_total{
service="checkout-api",
environment="prod",
status_code=~"5.."
}[5m])
)
A histogram-derived percentile is an estimate whose precision depends on bucket resolution. Do not average instance-level p95 values to obtain a service-wide p95; aggregate histogram observations first. (prometheus.io)
Practical checks: Inspect the available labels before using these examples. Examine affected-route volume before interpreting percentile movement. An empty result may indicate a schema mismatch, not healthy latency.

3. Follow slow traces into correlated logs
Choose several slow requests from the affected route and incident window, plus healthy requests for comparison.
For each trace, inspect:
- The API server span’s total duration.
- Database and outbound HTTP spans.
- Repeated calls or retry attempts.
- Time gaps without explanatory child spans.
- Differences between affected and healthy instances.
A span represents an operation with timing and contextual attributes. Follow the sequence of operations that determines request completion; do not simply add all child-span durations because operations can overlap. (opentelemetry.io)
Search associated logs using a trace ID. OpenTelemetry SDKs can attach trace and span context to log records, but confirm that your logging integration actually emits it. (opentelemetry.io)
For JSON logs containing a trace_id field:
{service_name="checkout-api", environment="prod"}
| json
| __error__=""
| trace_id="REPLACE_WITH_TRACE_ID"
Use your actual stream labels and field names. Loki’s JSON parser extracts fields for subsequent filtering; __error__="" excludes parsing-error entries from this view. Review parsing failures separately if expected logs disappear. (grafana.com)
Set a narrow time range and a bounded result limit. Look for pool-acquisition errors, outbound timeouts, cancellation and retry messages—not just generic “error” entries.
Common mistake: Treating a missing span as proof that an operation did not occur. Trace sampling and instrumentation coverage limit what you can observe. Check both before ruling out a dependency. (opentelemetry.io)
4. Compare the onset with recent changes
Build a short change timeline covering:
- Application releases and feature flags.
- Pool, timeout and concurrency configuration.
- Gateway or routing changes.
- Database migrations.
- Scheduled jobs and traffic changes.
For each candidate, ask:
- Did it precede the first observed impact?
- Does its deployment scope match the affected population?
- Do unaffected instances provide a useful comparison?
- Is there a plausible mechanism connecting it to the evidence?
A nearby deployment is a lead, not proof. Prefer a comparison between changed and unchanged populations over a conclusion based only on timestamps.
Before rollback, check whether schema or configuration changes make the previous version unsafe. AWS operational-excellence guidance emphasizes small, reversible changes and planning for unsuccessful changes; apply that principle rather than assuming every release can be reversed immediately.
Practical check: Preserve the release identifier, configuration difference and relevant evidence before changing anything. State the hypothesis explicitly: “This release may increase connection-holding time,” rather than “The deployment caused it.”
5. Distinguish database-pool waiting from slow database work
Check application pool usage, pending acquisition requests, acquisition timeouts and wait duration together. OpenTelemetry defines connection-pool instruments for these measurements, but the pool conventions are marked Development, and available names depend on instrumentation. Verify exported metrics rather than guessing their Prometheus translations. (opentelemetry.io)
Use this diagnostic distinction:
| Evidence pattern | Working hypothesis | Next check |
|---|---|---|
| Used connections near the configured maximum; pending requests rising | Requests may be waiting for pool capacity | Connection-holding time and recent pool changes |
| Database operations themselves are slower | Database work may be the bottleneck | Database latency, locks and resource headroom |
| A long unexplained interval before a database span | Acquisition waiting or missing instrumentation is possible | Pool metrics and acquisition logs |
| Only one application instance is affected | A local pool or process problem is possible | Instance-specific telemetry and configuration |
These patterns guide investigation; none proves a cause alone.
Illustrative worked example: An affected route slows after a release. Slow traces show a long interval before database work begins. Pool telemetry shows rising acquisition waiting, while observed database-operation duration remains close to its healthy baseline.
The working hypothesis is pool waiting, not “the database is slow.” Investigate whether the release holds connections longer or changed pool configuration. Select an approved rollback or workload reduction if evidence supports it.
Do not enlarge pools or add API replicas blindly. Additional application capacity can increase pressure on a constrained database or other backend. (sre.google)
After mitigation, check that acquisition waiting declines and user-facing requests recover.
6. Investigate downstream timeouts and local saturation
For outbound calls, compare dependency duration, timeout counts and attempt volume with incoming API demand.
Illustrative scenario: Slow traces contain repeated calls to the same downstream service. Logs report timeouts, and outbound attempts increase without a comparable rise in incoming requests.
The downstream delay is one hypothesis; retry amplification is another. Google’s SRE guidance explains that retries can worsen overload, especially when multiple layers retry. Recommended safeguards include bounded retries, randomized exponential backoff and retry budgets. (sre.google)
Use existing, approved controls to consider:
- Disabling a nonessential dependency-backed feature.
- Reducing optional background work.
- Limiting retry amplification.
- Applying a tested degraded mode.
- Routing elsewhere only after checking destination health and capacity.
Also inspect local CPU saturation or throttling, memory pressure, garbage collection, worker queues, in-flight requests and connection limits. Queueing and resource exhaustion can contribute to rising latency; do not restrict the investigation to CPU utilization alone. (sre.google)
Common mistake: Increasing timeouts to make timeout errors disappear. Longer waits can retain resources, while overly short deadlines can reject legitimate work. Evaluate deadlines against the request’s end-to-end budget and known service behavior. (sre.google)
7. Choose one reversible mitigation and escalate early
Before acting, record:
| Decision field | What to write |
|---|---|
| Hypothesis | The bottleneck you intend to relieve |
| Evidence | Relevant metrics, trace IDs and log excerpts |
| Action | The approved mitigation |
| Scope | Affected route, feature or deployment population |
| Expected signal | What should improve |
| Reversal condition | What would make you undo the action |
Prefer one controlled change at a time when the incident allows. If restarting is justified, canary it and preserve evidence first; restarting broadly can shift load or amplify an existing failure. (sre.google)
Escalate under your organization’s policy, particularly when:
- User impact meets the incident-declaration threshold.
- Multiple services or regions are affected.
- A dependency owner must authorize the next action.
- Safe mitigation is unavailable.
- Recovery attempts fail or impact expands.
- Data correctness or security may be involved.
Assign an incident lead, an operations owner and a communications owner as needed. Explicit roles reduce coordination ambiguity during response. (sre.google)
8. Verify recovery and preserve the timeline
Do not declare recovery because one latency chart falls.
Use this acceptance checklist:
- Affected routes meet their established latency and error criteria.
- Successful request throughput recovers.
- Gateway or client timeout evidence improves.
- Pool waiting, queues or retries move toward healthy behavior.
- Comparable request traffic is still reaching the service.
- Fresh traces support the expected improvement.
- No other region or dependency has inherited the problem.
Observe recovery across your predefined validation window, accounting for telemetry delay and query windows. There is no universal duration suitable for every API.
Record detection, impact confirmation, escalation, hypotheses, actions and validation in UTC. For each action, include its owner and result; separate facts from assumptions. Maintain a shared incident record so responders can coordinate without reconstructing the investigation. (sre.google)
Finally, assign owners for temporary-control removal, instrumentation gaps and follow-up work. A successful mitigation restores service; it does not, by itself, establish the root cause.