Answer box
Prove a backup works by restoring it into an isolated environment, checking the recovered data, and exercising essential application workflows. Record the recoverable data point and the time required to reach an agreed service-ready state—not just the database restore time. Compare that evidence with your recovery objectives. AWS explicitly recommends recovery testing rather than assuming that completed backups are usable. (docs.aws.amazon.com)
This runbook uses an illustrative cloud application with Amazon RDS for PostgreSQL and native RDS automated backups. It covers a same-account, same-Region restoration into a separate drill VPC. It does not cover Aurora, AWS Backup restore jobs, or production failover. The commands are templates, not a tested deployment.
1. Define what recovery must accomplish
Agree on two business requirements before selecting a backup:
- Recovery point objective (RPO): the maximum acceptable data-loss interval.
- Recovery time objective (RTO): the maximum acceptable delay between service interruption and service restoration. (aws.amazon.com)
Translate these into a drill contract. Specify the simulated incident, critical workflows, required dependencies, acceptable performance, and who can declare recovery complete.
For this exercise, define:
- Start: the simulated service interruption, before access checks and restoration.
- Finish: successful completion of the agreed application acceptance tests.
- Exclusions: live traffic switching, external-provider recovery, and any infrastructure prepared before the timer started.
Record database availability as an intermediate milestone. Do not stop the recovery clock there.
A same-Region exercise cannot demonstrate recovery from Region loss. Likewise, prebuilding the drill network does not measure emergency network provisioning. Report these limitations alongside the outcome rather than silently treating a partial drill as full disaster recovery. AWS distinguishes recovery strategies by disaster scope and the infrastructure needed to restore service. (aws.amazon.com)
2. Map dependencies and prove recovery access
Build a recovery inventory around the application—not just its database.
| Dependency | Evidence or decision to prepare |
|---|---|
| Application artifact | Recoverable image or package, exact version, deployment configuration |
| PostgreSQL | Source identifier, engine version, backup retention, recovery point |
| Database configuration | Parameter group, extensions, authentication requirements |
| Secrets and encryption | Recovery-role access, secret references, KMS key identifier |
| Object storage | Required files and their consistency relationship with database records |
| Queues and scheduled work | Whether to recreate, restore, suppress, or reconcile |
| Network and DNS | Drill-only endpoints and permitted connectivity |
| External integrations | Disabled connection, test account, or local stub |
For each dependency, name an owner and explain how it will be recovered. Treat an unavailable artifact repository or inaccessible secret as a recovery blocker.
Encryption deserves its own check. RDS encryption covers database storage, logs, automated backups, and snapshots. These resources depend on AWS KMS keys; RDS checks the creating principal’s key access when creating an encrypted instance. Existing production access is therefore not sufficient evidence that the intended recovery role can restore successfully. (docs.aws.amazon.com)
Before the drill, verify the recovery role’s RDS permissions, relevant KMS authorization, and ability to retrieve required application secrets. Do not disable a production key to test failure handling.
Record the actual PostgreSQL version and application artifact version. This guide deliberately does not prescribe an engine upgrade during recovery.

3. Establish isolation before starting the application
Prepare an isolated drill network with no route to production systems. AWS’s disaster-recovery drill guidance recommends this boundary, reinforced by security groups and network access controls. Although that guidance concerns Elastic Disaster Recovery, the isolation principle is applicable to this proposed application drill. (docs.aws.amazon.com)
Use the following pre-start checklist:
- Restore into a dedicated DB subnet group in the drill VPC.
- Use a drill database security group that permits connections only from approved drill clients.
- Review application outbound access, allowing only required drill dependencies.
- Give the application a drill-only database endpoint and credentials.
- Replace email, SMS, payment, and webhook integrations with test destinations or stubs.
- Keep schedulers, queue consumers, and background workers stopped until their configuration is reviewed.
- Keep production DNS records and load-balancer targets unchanged.
Security groups control inbound and outbound traffic, but they are only one part of the isolation design. Review the effective network configuration rather than relying on the word “private” in a subnet’s name. (docs.aws.amazon.com)
For this runbook, also use a drill application identity without permission to modify production cloud resources. Network isolation alone is not a substitute for reviewing cloud-service permissions.
Before launching the restored application, check configuration for production database hosts, bucket names, queue URLs, and notification destinations. Require both application-level suppression and infrastructure-level restrictions for potentially harmful integrations.
4. Select the recovery point and restore a new database
RDS point-in-time recovery creates a new DB instance without modifying the source instance. RDS uploads transaction logs every five minutes; inspect LatestRestorableTime rather than interpreting that interval as a guaranteed RPO. The selected point must be within the available recovery window. (docs.aws.amazon.com)
A snapshot restore also creates a new instance. It restores the entire DB instance, not an individual database, and cannot restore over an existing instance. Choose that route when the exercise is specifically testing recovery from a named snapshot. (docs.aws.amazon.com)
The following read-only inspection template assumes you have set AWS_PROFILE, AWS_REGION, and SOURCE_DB:
aws rds describe-db-instances \
--profile "$AWS_PROFILE" \
--region "$AWS_REGION" \
--db-instance-identifier "$SOURCE_DB" \
--query 'DBInstances[0].{
ARN:DBInstanceArn,
Engine:Engine,
Version:EngineVersion,
RetentionDays:BackupRetentionPeriod,
LatestRestorableTime:LatestRestorableTime,
Encrypted:StorageEncrypted,
KmsKeyId:KmsKeyId
}' \
--output json
Check the account, Region, engine, retention, and recovery timestamp before proceeding. The official CLI reference documents these inspection fields. (docs.aws.amazon.com)
For PITR, explicitly supply the drill network and parameter group:
aws rds restore-db-instance-to-point-in-time \
--profile "$AWS_PROFILE" \
--region "$AWS_REGION" \
--source-db-instance-identifier "$SOURCE_DB" \
--target-db-instance-identifier "$DRILL_DB" \
--restore-time "$RESTORE_TIME_UTC" \
--db-subnet-group-name "$DRILL_SUBNET_GROUP" \
--vpc-security-group-ids "$DRILL_DB_SECURITY_GROUP" \
--db-parameter-group-name "$DRILL_PARAMETER_GROUP" \
--no-publicly-accessible \
--deletion-protection
This command creates a billable resource. All variables must be reviewed, existing resources must be appropriate for the selected engine, and DRILL_DB must be a new identifier distinct from the source. Use a UTC restore timestamp. The CLI supports these settings; this template is intentionally not a complete production configuration. (docs.aws.amazon.com)
Review instance class, storage, authentication, and Availability Zone topology separately. PITR defaults include a single-AZ deployment for PostgreSQL unless configured otherwise; omitted parameter-group settings can also lead to defaults. Do not assume the restored topology matches production. (docs.aws.amazon.com)
When the instance becomes available, verify its actual endpoint, VPC, security groups, encryption, and applied configuration before permitting application connections.
5. Validate recovered data and application behavior
Start with database checks, then move to business workflows. AWS identifies restoring without querying or retrieving data as a recovery-testing anti-pattern. (docs.aws.amazon.com)
Prepare expected results before the drill. Useful checks include:
- Known records that should exist at the selected recovery point.
- Important relationships and application-specific invariants.
- Required schemas and extensions.
- Consistency between database references and recovered object-storage files.
- Evidence that recent acknowledged transactions were recovered.
Do not compare a restored database with production’s current row count: production may have continued changing. Compare with evidence appropriate to the chosen recovery point.
For a fictional schema containing orders and customers, an illustrative relationship check is:
BEGIN READ ONLY;
SELECT COUNT(*) AS orphan_orders
FROM orders AS o
LEFT JOIN customers AS c ON c.id = o.customer_id
WHERE o.customer_id IS NOT NULL
AND c.id IS NULL;
ROLLBACK;
Adapt this to your actual schema and deletion rules. Zero is an expected result only if your application forbids such orphan records. PostgreSQL read-only transactions prohibit ordinary writes to non-temporary tables; they do not replace endpoint verification or isolation controls. (postgresql.org)
Next, deploy the selected application artifact and test representative workflows:
- Authenticate using a controlled drill account.
- Read known recovered records.
- Create a synthetic record in the restored database only.
- Retrieve it through the application.
- Exercise one approved background workflow.
- Confirm any notification reaches only the test destination.
Retain logs and test outcomes without exporting unnecessary sensitive data. Mark stubbed integrations as untested, not passed.
Finally, exercise representative reads and record latency against the agreed acceptance threshold. RDS can become available while data blocks are still loading in the background, so availability alone does not establish normal performance. (docs.aws.amazon.com)
6. Calculate recovery evidence without inventing a result
Consider a planning example, not a measured drill:
- Simulated interruption: 12:00 UTC
- Proposed restore target: 11:50 UTC
- Agreed RPO: 15 minutes
- Agreed RTO: 60 minutes
The selected recovery-point gap is:
12:00 − 11:50 = 10 minutes
That selection fits inside the example RPO, but it does not prove that the recovered data passes. Confirm that the target was available and that the restored records support the claimed recovery point.
For stronger evidence, use recovery markers whose successful commits were independently recorded. Document their frequency and timing uncertainty. Do not treat an application’s created_at value as proof of database commit time.
For recovery time, calculate:
Observed recovery duration =
application acceptance completed − simulated interruption
Record intermediate timestamps for authorization, restore submission, database availability, application deployment, and validation.
Until the drill actually runs, the recovery duration is unmeasured. Provisioning estimates and previous exercises are useful planning inputs, but neither is the current exercise’s result.
The final report should distinguish:
- Measured facts.
- Estimates and assumptions.
- Tested dependencies.
- Excluded or stubbed dependencies.
- Failed checks and unresolved uncertainty.
7. Clean up safely and turn findings into a repeatable runbook
Before deletion, preserve the drill report, restore identifiers, selected timestamp, validation evidence, and unresolved issues.
Then perform a scoped cleanup:
- Stop drill application instances, workers, and schedules.
- Verify the account, Region, drill database ARN, and identifier.
- Remove only drill-specific secrets, endpoints, and supporting resources.
- Disable deletion protection on the drill database only when approved.
- Delete the drill database using the agreed snapshot and backup-retention choices.
- Check for remaining drill snapshots, retained backups, and other chargeable resources.
Deleting an RDS instance does not delete its manual snapshots. Retained automated backups and manual snapshots can continue incurring charges. Do not use a broad cleanup operation that could touch production recovery points or shared KMS keys. (docs.aws.amazon.com)
Common failures should become explicit runbook checks:
| Failure | Practical correction |
|---|---|
| Restore role cannot use the encryption key | Validate recovery authorization before timing the exercise |
| Application still targets production | Gate startup on endpoint and configuration review |
| Database works but application fails | Test artifacts, secrets, schema compatibility, and dependencies |
| “Available” is mistaken for service-ready | Require data and application acceptance tests |
| Workers send real notifications | Keep them stopped until test destinations and restrictions are verified |
Repeat the exercise periodically and after significant workload changes. A successful drill is evidence for its tested scope and conditions—not a guarantee of every future recovery. (docs.aws.amazon.com)