‹ All study guides

AWS disaster recovery strategies, RTO and RPO

Updated 12 October 2026 · 7 min read

A company runs everything in one AWS Region and the business asks a simple question: if that Region goes dark, how long until we are back, and how much data do we lose? The answer decides how much the company pays every month for a second Region that, on a good day, does nothing. Disaster recovery is that trade-off between cost and speed of recovery, and AWS describes it as four named strategies.

The Solutions Architect exams love this topic because the question almost writes itself: give two numbers, add a budget constraint, and ask which strategy fits. If you know what each strategy keeps running in the recovery Region, and what has to happen before it can take traffic, you can answer these in under a minute.

RTO and RPO

Two targets drive every DR decision.

  • Recovery time objective (RTO) is how long the workload may be unavailable. It is measured from the moment of the disaster to the moment customers are served again.
  • Recovery point objective (RPO) is how much data you may lose, expressed as time. An RPO of one hour means losing the last hour of writes is acceptable; an RPO of seconds means data must already be copied somewhere else almost as it is written.

Keep them separate in your head. RPO is about how data gets to the recovery site (backups taken on a schedule, or continuous replication). RTO is about how much of the application is already running there. A design can have an excellent RPO and a poor RTO: continuously replicated data sitting next to an application that still needs to be deployed from scratch.

The four strategies

AWS orders the strategies from cheapest and slowest to most expensive and fastest. The first three are active/passive: one Region serves traffic and the other waits. The fourth serves traffic from every Region all the time.

Backup and restore

Data is backed up and the backups are copied to the recovery Region. Nothing else runs there. On a disaster you deploy the infrastructure, restore the data and redeploy the application. RPO is the time since the last backup, and RTO is however long that rebuild takes, typically hours.

This only works in a reasonable time if the infrastructure is defined as code (CloudFormation or the CDK), AMIs are copied to the recovery Region, and the restore has been tested. AWS Backup is the service exams expect here: one backup plan can schedule backups across EC2, EBS, RDS and Aurora, DynamoDB, EFS and FSx, and a copy rule in the same plan sends each recovery point to a vault in another Region or another account. Backups also protect against something replication cannot: corrupted or deleted data, which replication would faithfully copy.

Pilot light

The data layer is live in the recovery Region and kept current by continuous replication, for example an Aurora global database secondary, an RDS cross-Region read replica, DynamoDB global tables or S3 replication. The application tier is defined and ready (AMIs copied, templates written) but is not running, or is not even deployed. On failover you promote the database, deploy or start the application servers and scale them up.

RPO is now seconds to minutes because data is replicated continuously. RTO is shorter than backup and restore, but still includes launching and scaling the compute, so it is typically tens of minutes. The key fact: a pilot light environment cannot serve a single request until someone takes action.

AWS Elastic Disaster Recovery (DRS) is the managed way to run pilot light for servers. An agent on each source server (on premises, in another cloud, or on EC2) replicates block-level changes continuously to a staging area of small, low-cost instances and volumes in your account. Nothing full-sized runs until you launch recovery instances for a drill or a real event, and they can come up within minutes from the latest state or an earlier point in time. It suits server-hosted applications and databases; it is not how you protect a managed service like RDS.

Warm standby

A complete, working copy of production runs in the recovery Region, just smaller: a couple of instances behind a load balancer instead of twenty, a database replica instead of a full cluster. Because it is already running, it can take real traffic the moment you send it there, at reduced capacity, and then scale out with Auto Scaling.

This is the distinction exams test most: pilot light needs servers turned on before it can answer; warm standby only needs to scale up. Warm standby costs more than pilot light because compute runs all the time, but RTO falls to minutes, and you can test it continuously by sending it a trickle of traffic. If the recovery Region runs enough capacity for the full load without scaling, AWS calls that hot standby.

Multi-site active/active

The full workload runs in two or more Regions and all of them serve users. There is no failover in the usual sense; when one Region fails, routing stops sending users there and the others absorb the load. RTO is near zero. RPO depends on how writes are handled: DynamoDB global tables accept writes in every Region, while an Aurora global database still writes in one primary Region and promotes another if that Region fails. This is the most expensive and complex option, and it still needs point-in-time backups, because a bad write replicates everywhere.

The services that make it work

  • Aurora global database replicates storage from the primary Region to up to 10 secondary Regions, typically with under a second of lag. A switchover moves the primary to a healthy secondary with no data loss, for planned drills. A failover is for a real outage of the primary Region and can lose whatever had not replicated yet.
  • Route 53 failover routing gives a primary record and a secondary record. While the primary's health check passes, Route 53 answers with the primary; when it fails, it answers with the secondary. That is active/passive. Active/active uses other policies (latency, weighted, geolocation) with health checks, so unhealthy endpoints simply drop out of answers.
  • AWS Backup schedules backups and copies them across Regions and accounts, and with AWS Organizations backup policies it can enforce one plan across many accounts.
  • Elastic Disaster Recovery covers lift-and-shift servers, including on-premises ones, with continuous replication and a cheap staging area.

How to choose

  1. Read the RPO first. Hours means scheduled backups are fine. Seconds or minutes means continuous replication, which rules out plain backup and restore.
  2. Then read the RTO. Hours fits backup and restore. Tens of minutes fits pilot light. A few minutes with "reduced capacity is acceptable" points to warm standby. Near zero, or "users in both Regions", points to active/active.
  3. Check the cost constraint. "Must not run full production capacity in the second Region" rules out active/active and hot standby. "Lowest possible ongoing cost" with a tight RPO points to pilot light or Elastic Disaster Recovery.
  4. Look for "no deployment or instance launch in the critical path". That phrase eliminates pilot light and backup and restore, leaving warm standby or better.
  5. For servers outside AWS, or EC2 fleets you want to protect without re-architecting, think Elastic Disaster Recovery.

Common exam traps

  • Confusing pilot light with warm standby. If the recovery Region's compute is at zero or not deployed, it is pilot light, however fast the scripts are. Warm standby is already serving.
  • Treating Multi-AZ as disaster recovery. RDS Multi-AZ and a multi-AZ Auto Scaling group protect against losing a data center, not a Region. A Region-wide requirement needs a second Region.
  • Using replication as a backup. Replication copies deletions and corruption too. An answer that drops point-in-time backups because "data is already replicated" is wrong.
  • Picking failover when the scenario is a planned drill. For a rehearsal with both Regions healthy and no data loss allowed, the Aurora operation is a switchover, not a failover.
  • Over-buying. Active/active meets every RTO, but if the question caps cost or forbids full capacity in the second Region, it is the wrong answer.
  • Forgetting the data layer in pilot light designs. An option that copies AMIs but restores the database from a nightly snapshot cannot meet an RPO of minutes.

A worked example

An online booking site runs in one Region on EC2 behind an Application Load Balancer, with Aurora MySQL. After an outage the board sets an RTO of 10 minutes and an RPO of 1 minute, and says the second Region must be able to take traffic as soon as failover is declared, but must not run at full size while the primary is healthy.

The RPO of one minute rules out scheduled snapshots, so data must replicate continuously: add an Aurora global database secondary cluster in the recovery Region. "Take traffic as soon as failover is declared" rules out pilot light, because there would be no running application to send users to. "Not full size" rules out active/active and hot standby. That leaves warm standby: a small Auto Scaling group behind its own load balancer in the second Region, Route 53 failover records with a health check on the primary, and a runbook that promotes the Aurora secondary and raises the Auto Scaling group's capacity. AWS Backup copies to the second Region still belong in the design, for the day someone deletes a table.

Practice questions

Further reading

Practice questions