Uncategorized

Proving Your Backups Work: An Isolated Restore Drill for a Cloud Application

Aisle between black server cabinets with illuminated status lights and overhead cabling.

Answer box

Prove a backup works by restoring it into an isolated environment, checking the recovered data, and exercising essential application workflows. Record the recoverable data point and the time required to reach an agreed service-ready state—not just the database restore time. Compare that evidence with your recovery objectives. AWS explicitly recommends recovery testing rather than assuming that completed backups are usable. (docs.aws.amazon.com)

This runbook uses an illustrative cloud application with Amazon RDS for PostgreSQL and native RDS automated backups. It covers a same-account, same-Region restoration into a separate drill VPC. It does not cover Aurora, AWS Backup restore jobs, or production failover. The commands are templates, not a tested deployment.

1. Define what recovery must accomplish

Agree on two business requirements before selecting a backup:

  • Recovery point objective (RPO): the maximum acceptable data-loss interval.
  • Recovery time objective (RTO): the maximum acceptable delay between service interruption and service restoration. (aws.amazon.com)

Translate these into a drill contract. Specify the simulated incident, critical workflows, required dependencies, acceptable performance, and who can declare recovery complete.

For this exercise, define:

  • Start: the simulated service interruption, before access checks and restoration.
  • Finish: successful completion of the agreed application acceptance tests.
  • Exclusions: live traffic switching, external-provider recovery, and any infrastructure prepared before the timer started.

Record database availability as an intermediate milestone. Do not stop the recovery clock there.

A same-Region exercise cannot demonstrate recovery from Region loss. Likewise, prebuilding the drill network does not measure emergency network provisioning. Report these limitations alongside the outcome rather than silently treating a partial drill as full disaster recovery. AWS distinguishes recovery strategies by disaster scope and the infrastructure needed to restore service. (aws.amazon.com)

2. Map dependencies and prove recovery access

Build a recovery inventory around the application—not just its database.

Dependency Evidence or decision to prepare
Application artifact Recoverable image or package, exact version, deployment configuration
PostgreSQL Source identifier, engine version, backup retention, recovery point
Database configuration Parameter group, extensions, authentication requirements
Secrets and encryption Recovery-role access, secret references, KMS key identifier
Object storage Required files and their consistency relationship with database records
Queues and scheduled work Whether to recreate, restore, suppress, or reconcile
Network and DNS Drill-only endpoints and permitted connectivity
External integrations Disabled connection, test account, or local stub

For each dependency, name an owner and explain how it will be recovered. Treat an unavailable artifact repository or inaccessible secret as a recovery blocker.

Encryption deserves its own check. RDS encryption covers database storage, logs, automated backups, and snapshots. These resources depend on AWS KMS keys; RDS checks the creating principal’s key access when creating an encrypted instance. Existing production access is therefore not sufficient evidence that the intended recovery role can restore successfully. (docs.aws.amazon.com)

Before the drill, verify the recovery role’s RDS permissions, relevant KMS authorization, and ability to retrieve required application secrets. Do not disable a production key to test failure handling.

Record the actual PostgreSQL version and application artifact version. This guide deliberately does not prescribe an engine upgrade during recovery.

Aisle between black server cabinets with illuminated status lights and overhead cabling.
Server room at The National Archives, UK. Illustrative infrastructure photograph; not an AWS facility. — The National Archives (UK). The National Archives (UK) Source CC BY 3.0

3. Establish isolation before starting the application

Prepare an isolated drill network with no route to production systems. AWS’s disaster-recovery drill guidance recommends this boundary, reinforced by security groups and network access controls. Although that guidance concerns Elastic Disaster Recovery, the isolation principle is applicable to this proposed application drill. (docs.aws.amazon.com)

Use the following pre-start checklist:

  1. Restore into a dedicated DB subnet group in the drill VPC.
  2. Use a drill database security group that permits connections only from approved drill clients.
  3. Review application outbound access, allowing only required drill dependencies.
  4. Give the application a drill-only database endpoint and credentials.
  5. Replace email, SMS, payment, and webhook integrations with test destinations or stubs.
  6. Keep schedulers, queue consumers, and background workers stopped until their configuration is reviewed.
  7. Keep production DNS records and load-balancer targets unchanged.

Security groups control inbound and outbound traffic, but they are only one part of the isolation design. Review the effective network configuration rather than relying on the word “private” in a subnet’s name. (docs.aws.amazon.com)

For this runbook, also use a drill application identity without permission to modify production cloud resources. Network isolation alone is not a substitute for reviewing cloud-service permissions.

Before launching the restored application, check configuration for production database hosts, bucket names, queue URLs, and notification destinations. Require both application-level suppression and infrastructure-level restrictions for potentially harmful integrations.

4. Select the recovery point and restore a new database

RDS point-in-time recovery creates a new DB instance without modifying the source instance. RDS uploads transaction logs every five minutes; inspect LatestRestorableTime rather than interpreting that interval as a guaranteed RPO. The selected point must be within the available recovery window. (docs.aws.amazon.com)

A snapshot restore also creates a new instance. It restores the entire DB instance, not an individual database, and cannot restore over an existing instance. Choose that route when the exercise is specifically testing recovery from a named snapshot. (docs.aws.amazon.com)

The following read-only inspection template assumes you have set AWS_PROFILE, AWS_REGION, and SOURCE_DB:

aws rds describe-db-instances \
  --profile "$AWS_PROFILE" \
  --region "$AWS_REGION" \
  --db-instance-identifier "$SOURCE_DB" \
  --query 'DBInstances[0].{
    ARN:DBInstanceArn,
    Engine:Engine,
    Version:EngineVersion,
    RetentionDays:BackupRetentionPeriod,
    LatestRestorableTime:LatestRestorableTime,
    Encrypted:StorageEncrypted,
    KmsKeyId:KmsKeyId
  }' \
  --output json

Check the account, Region, engine, retention, and recovery timestamp before proceeding. The official CLI reference documents these inspection fields. (docs.aws.amazon.com)

For PITR, explicitly supply the drill network and parameter group:

aws rds restore-db-instance-to-point-in-time \
  --profile "$AWS_PROFILE" \
  --region "$AWS_REGION" \
  --source-db-instance-identifier "$SOURCE_DB" \
  --target-db-instance-identifier "$DRILL_DB" \
  --restore-time "$RESTORE_TIME_UTC" \
  --db-subnet-group-name "$DRILL_SUBNET_GROUP" \
  --vpc-security-group-ids "$DRILL_DB_SECURITY_GROUP" \
  --db-parameter-group-name "$DRILL_PARAMETER_GROUP" \
  --no-publicly-accessible \
  --deletion-protection

This command creates a billable resource. All variables must be reviewed, existing resources must be appropriate for the selected engine, and DRILL_DB must be a new identifier distinct from the source. Use a UTC restore timestamp. The CLI supports these settings; this template is intentionally not a complete production configuration. (docs.aws.amazon.com)

Review instance class, storage, authentication, and Availability Zone topology separately. PITR defaults include a single-AZ deployment for PostgreSQL unless configured otherwise; omitted parameter-group settings can also lead to defaults. Do not assume the restored topology matches production. (docs.aws.amazon.com)

When the instance becomes available, verify its actual endpoint, VPC, security groups, encryption, and applied configuration before permitting application connections.

5. Validate recovered data and application behavior

Start with database checks, then move to business workflows. AWS identifies restoring without querying or retrieving data as a recovery-testing anti-pattern. (docs.aws.amazon.com)

Prepare expected results before the drill. Useful checks include:

  • Known records that should exist at the selected recovery point.
  • Important relationships and application-specific invariants.
  • Required schemas and extensions.
  • Consistency between database references and recovered object-storage files.
  • Evidence that recent acknowledged transactions were recovered.

Do not compare a restored database with production’s current row count: production may have continued changing. Compare with evidence appropriate to the chosen recovery point.

For a fictional schema containing orders and customers, an illustrative relationship check is:

BEGIN READ ONLY;

SELECT COUNT(*) AS orphan_orders
FROM orders AS o
LEFT JOIN customers AS c ON c.id = o.customer_id
WHERE o.customer_id IS NOT NULL
  AND c.id IS NULL;

ROLLBACK;

Adapt this to your actual schema and deletion rules. Zero is an expected result only if your application forbids such orphan records. PostgreSQL read-only transactions prohibit ordinary writes to non-temporary tables; they do not replace endpoint verification or isolation controls. (postgresql.org)

Next, deploy the selected application artifact and test representative workflows:

  1. Authenticate using a controlled drill account.
  2. Read known recovered records.
  3. Create a synthetic record in the restored database only.
  4. Retrieve it through the application.
  5. Exercise one approved background workflow.
  6. Confirm any notification reaches only the test destination.

Retain logs and test outcomes without exporting unnecessary sensitive data. Mark stubbed integrations as untested, not passed.

Finally, exercise representative reads and record latency against the agreed acceptance threshold. RDS can become available while data blocks are still loading in the background, so availability alone does not establish normal performance. (docs.aws.amazon.com)

6. Calculate recovery evidence without inventing a result

Consider a planning example, not a measured drill:

  • Simulated interruption: 12:00 UTC
  • Proposed restore target: 11:50 UTC
  • Agreed RPO: 15 minutes
  • Agreed RTO: 60 minutes

The selected recovery-point gap is:

12:00 − 11:50 = 10 minutes

That selection fits inside the example RPO, but it does not prove that the recovered data passes. Confirm that the target was available and that the restored records support the claimed recovery point.

For stronger evidence, use recovery markers whose successful commits were independently recorded. Document their frequency and timing uncertainty. Do not treat an application’s created_at value as proof of database commit time.

For recovery time, calculate:

Observed recovery duration =
application acceptance completed − simulated interruption

Record intermediate timestamps for authorization, restore submission, database availability, application deployment, and validation.

Until the drill actually runs, the recovery duration is unmeasured. Provisioning estimates and previous exercises are useful planning inputs, but neither is the current exercise’s result.

The final report should distinguish:

  • Measured facts.
  • Estimates and assumptions.
  • Tested dependencies.
  • Excluded or stubbed dependencies.
  • Failed checks and unresolved uncertainty.

7. Clean up safely and turn findings into a repeatable runbook

Before deletion, preserve the drill report, restore identifiers, selected timestamp, validation evidence, and unresolved issues.

Then perform a scoped cleanup:

  1. Stop drill application instances, workers, and schedules.
  2. Verify the account, Region, drill database ARN, and identifier.
  3. Remove only drill-specific secrets, endpoints, and supporting resources.
  4. Disable deletion protection on the drill database only when approved.
  5. Delete the drill database using the agreed snapshot and backup-retention choices.
  6. Check for remaining drill snapshots, retained backups, and other chargeable resources.

Deleting an RDS instance does not delete its manual snapshots. Retained automated backups and manual snapshots can continue incurring charges. Do not use a broad cleanup operation that could touch production recovery points or shared KMS keys. (docs.aws.amazon.com)

Common failures should become explicit runbook checks:

Failure Practical correction
Restore role cannot use the encryption key Validate recovery authorization before timing the exercise
Application still targets production Gate startup on endpoint and configuration review
Database works but application fails Test artifacts, secrets, schema compatibility, and dependencies
“Available” is mistaken for service-ready Require data and application acceptance tests
Workers send real notifications Keep them stopped until test destinations and restrictions are verified

Repeat the exercise periodically and after significant workload changes. A successful drill is evidence for its tested scope and conditions—not a guarantee of every future recovery. (docs.aws.amazon.com)