Quick answer
When cloud spending rises unexpectedly, first verify the billing window and data freshness. Then isolate the account, service, and usage type responsible; compare usage and cost on the same basis; and assign an accountable technical owner. Treat shutdowns as production changes—not automatic responses to a billing alert.
Investigation output: a documented cost driver, an owner who has accepted the investigation, a safe next action, and a time to verify the outcome.
This runbook uses AWS Cost Anomaly Detection, Cost Explorer, and Cost and Usage Report 2.0 (CUR 2.0) queried through Amazon Athena. It assumes an existing billing export and authorized access to billing data and workload telemetry.
The SQL examples are illustrative and have not been tested against your environment. Replace table names and dates, check the exported schema, and apply your normal query-cost controls.
1. Verify the billing window before treating the increase as real
Record the anomaly’s start date, expected spend, actual spend, account, service, Region, and usage type. AWS provides these investigation dimensions, but its ranked potential root causes identify billing contributors—not necessarily the application change that caused them. (docs.aws.amazon.com)
Next, establish how current the evidence is:
- Record when you received the alert.
- Check the latest successful export delivery.
- Identify the latest usage interval represented.
- Mark recent intervals as potentially incomplete.
- Compare equivalent periods: complete days against complete days, or matching elapsed hours.
AWS Cost Anomaly Detection evaluates net unblended cost and runs approximately three times daily after billing data is processed. Its underlying Cost Explorer data can lag usage by up to 24 hours. A newly created monitor can also take 24 hours to begin detection. It is not a real-time spending brake. (docs.aws.amazon.com)
A new Data Export can take up to 24 hours to start delivery; billing exports refresh at least daily thereafter. Export granularity and delivery frequency are different concepts: hourly rows do not imply hourly delivery. (docs.aws.amazon.com)
Practical check: if a deployment happened this morning, inspect operational telemetry now, but do not interpret an unchanged billing chart as proof that spending stopped.
2. Rank the billing contributors using a consistent cost basis
Start with the anomaly details, then use Cost Explorer to narrow the account, service, Region, and usage type. Export queries are useful when you need repeatable grouping or line-item detail. AWS supports querying delivered CUR 2.0 data through Athena; aggregation belongs there, not in the Data Exports query field. (docs.aws.amazon.com)
For an initial breakdown:
SELECT
CAST(line_item_usage_start_date AS date) AS usage_day,
line_item_usage_account_id AS account_id,
line_item_product_code AS service,
line_item_usage_type AS usage_type,
line_item_line_item_type AS charge_type,
SUM(line_item_unblended_cost) AS unblended_cost
FROM billing.cur2
WHERE line_item_usage_start_date >= TIMESTAMP '2026-09-21 00:00:00'
AND line_item_usage_start_date < TIMESTAMP '2026-09-24 00:00:00'
GROUP BY 1, 2, 3, 4, 5
ORDER BY unblended_cost DESC;
Replace the illustrative dates with the investigation window and add your table’s actual billing-period partition predicate where available.
This query deliberately reports unblended cost, not the detector’s net unblended cost. Keep that distinction visible. AWS defines unblended cost from usage and unblended rate; its net field represents after-discount cost. (docs.aws.amazon.com)
Run the same breakdown for a comparable baseline. Rank increases, not just the largest current charges.
Before trusting the result, verify that the table reads one current export version per billing period. In “create new” delivery mode, AWS retains successive refreshes; the latest manifest identifies the current files. Summing every retained refresh would count repeated data. (docs.aws.amazon.com)
3. Route the investigation to a technical owner, including shared costs
Use an explicit ownership order:
- A maintained workload or resource-to-team registry.
- Activated allocation metadata that agrees with that registry.
- The account’s accountable team.
- The shared-platform owner.
- A central CloudOps or FinOps queue until someone accepts responsibility.
Suggested resource tags are owner_team, application, and environment. These are proposed conventions, not AWS-required names. Use stable team identifiers rather than personal contact details.
Applying resource tags alone is insufficient for billing allocation: activate the relevant cost allocation tags. AWS notes that tags can take up to 24 hours to appear in the billing console. Do not include sensitive information in them. (docs.aws.amazon.com)
For exports containing the CUR 2.0 tags map, resource tags use keys such as resourceTags/owner_team; account tags have a separate accountTag/ prefix. Check the actual map before writing routing logic. (docs.aws.amazon.com)
Keep two roles separate:
- Technical owner: investigates and approves operational changes.
- Financial allocation recipient: receives an agreed share of the expense.
A shared networking charge might require the platform team to investigate traffic, even when several application teams bear its cost.
AWS Cost Categories supports proportional, fixed, and even splits. However, those split-charge results appear on the Cost Categories details page; they do not alter CUR or Cost Explorer data. Do not expect the query above to contain those allocated totals. (docs.aws.amazon.com)
4. Separate additional usage from a changed effective rate
After isolating a contributor, compare a homogeneous billing slice: the same account, service, usage type, operation, and unit. Narrow further by Region and product characteristics when relevant.
For ordinary Usage rows:
SELECT
CAST(line_item_usage_start_date AS date) AS usage_day,
line_item_usage_type AS usage_type,
line_item_operation AS operation,
pricing_unit AS unit,
SUM(line_item_usage_amount) AS usage_quantity,
SUM(line_item_unblended_cost) AS unblended_cost,
SUM(line_item_unblended_cost)
/ NULLIF(SUM(line_item_usage_amount), 0) AS cost_per_unit
FROM billing.cur2
WHERE line_item_usage_start_date >= TIMESTAMP '2026-09-21 00:00:00'
AND line_item_usage_start_date < TIMESTAMP '2026-09-24 00:00:00'
AND line_item_line_item_type = 'Usage'
GROUP BY 1, 2, 3, 4
ORDER BY unblended_cost DESC;
These fields describe usage quantities, operations, costs, and pricing units. This deliberately narrow query is a diagnostic view—not a complete invoice or commitment-cost calculation. (docs.aws.amazon.com)
Interpret the comparison as follows:
| Observation | Next investigation |
|---|---|
| Quantity rises; cost per unit is stable | Traffic, runtime, retries, scaling, or scheduled jobs |
| Quantity is stable; cost per unit rises | Product mix, pricing terms, discounts, or coverage |
| Both change | Investigate quantity and rate separately |
| Quantity is zero | Inspect the charge type; division cannot explain it |
Do not call a higher calculated average a provider price increase without checking the underlying terms. A changed workload mix can change that average.
For a single comparable slice, let quantity be Q and calculated cost per unit be R. The exact arithmetic decomposition is:
- Usage contribution:
(Q₁ − Q₀) × R₀ - Rate contribution:
Q₁ × (R₁ − R₀)
Together, these equal Q₁R₁ − Q₀R₀. This explains a comparable slice, not an entire mixed-service bill.
5. Work through egress and apparently idle compute
The following quantities are invented teaching examples—not measured results, AWS prices, or savings estimates.
Example A: increased egress
Suppose one comparable transfer slice rises from 200 GB to 500 GB per complete day, while its calculated cost per GB remains unchanged.
That is a 150% quantity increase. Under those assumptions, the slice’s cost also rises 150%. The investigation should focus on additional traffic rather than a changed rate.
Ask the owner to compare:
- Deployment and configuration timestamps.
- Export-job schedules.
- Request volume and response sizes.
- Retry counts and repeated downloads.
- Network-path or destination changes.
Do not assume every networking charge is internet egress. Inspect the actual usage type before combining transfer-related rows. Resource IDs can also be blank for data-transfer line items, so resource-level billing alone may not identify the originating application. (docs.aws.amazon.com)
Example B: additional compute runtime
Suppose equivalent instances increase from 48 to 96 aggregate instance-hours per complete day at the same calculated rate. That slice’s quantity and cost double.
Now inspect instance inventory, scaling configuration, application activity, and scheduled work. If monitoring suggests inactivity, record that as a hypothesis requiring owner confirmation—not permission to stop anything.
AWS bills running EC2 instances even when idle. Stopping an instance does not eliminate its EBS storage charges. (docs.aws.amazon.com)
Use this worksheet for either case:
| Field | What to record |
|---|---|
| Scope | Account, service, Region, usage type, operation |
| Window | Baseline and anomaly intervals; freshness limitations |
| Cost basis | Unblended, net unblended, or another explicit measure |
| Change | Quantity, cost, and calculated rate differences |
| Ownership | Accepted technical owner and financial recipient |
| Evidence | Query results, telemetry, and change records |
| Decision | Explanation, action, approval, rollback, verification time |
6. Require a reliability review before reducing or shutting down resources
Treat cost containment as a controlled operational change. Before proceeding, require answers to these questions:
- Dependencies: What users, jobs, or services depend on this resource or network path?
- Capacity: Does apparently spare capacity support failover, recovery, or a scheduled peak?
- Data: What persistent and ephemeral state must be protected?
- Approval: Has the accountable owner approved the change?
- Rollback: How will service be restored, and who will do it?
- Validation: Which user-facing indicators must remain healthy?
Prefer a small, reversible action with a defined observation period over deleting several suspected resources at once. This follows the supplied AWS operational-excellence guidance on safe automation and reversible changes.
For EC2 specifically, stop/start can erase instance-store data, and a stopped instance still incurs EBS storage charges. A stop is therefore neither risk-free nor a guarantee of zero remaining cost. (docs.aws.amazon.com)
If no owner is known, keep the case in the escalation queue. “Untagged” is an ownership defect, not a shutdown criterion.
Record immediate operational validation separately from later billing validation. A healthy service after the change does not yet establish the financial outcome.

7. Tune thresholds and close with evidence—not an estimated saving
AWS anomaly subscriptions support absolute and percentage impact thresholds, combined with AND or OR. These thresholds control notifications; detected anomalies below the notification threshold remain visible. (docs.aws.amazon.com)
An illustrative policy might notify when impact reaches $100 AND 30%. Those are example thresholds, not recommended defaults.
- AND requires both conditions, reducing notifications for small-dollar percentage spikes.
- OR catches either condition, broadening coverage.
- For a near-zero baseline, emphasize absolute impact and investigate new usage directly.
Set thresholds according to the team’s spend profile and ability to respond. Route alerts to an owned queue rather than an unattended mailbox.
Close the investigation only when:
- The owner has accepted the explanation.
- The increase is classified as expected, wasteful, a billing adjustment, or unresolved.
- Any mitigation has passed operational validation.
- Refreshed billing data supports the outcome—or a follow-up is scheduled.
- Missing ownership or allocation metadata has a tracked correction.
Avoid three common mistakes: comparing different cost bases, mistaking retained export refreshes for new usage, and declaring savings before comparable refreshed data exists.
A useful closeout is: “Additional transfer followed a configuration change; the owner approved a rollback; operational checks passed; billing verification is scheduled after the next refresh.” That records evidence without promising an unmeasured financial result.