Cloud Disaster Recovery Interview Questions and Answers, With a DR Drill Walk-Through
Explore essential interview questions and answers on cloud disaster recovery. This guide covers best practices, documentation maintenance, integrating third-party services, selecting a recovery vendor, and the role of data encryption. Enhance your understanding of disaster recovery and ensure business continuity with practical insights and expert advice.
Quick answer: Cloud disaster recovery interviews ask how you restore a whole service, not just data. Expect questions on RPO and RTO, DR strategies (backup and restore, pilot light, warm standby, active-active), failover and failback, replication, DNS, runbooks and testing. Explain trade-offs between cost and recovery speed, and describe how you would test.
Key takeaways
- RPO and RTO define the design; strategy choice trades cost against recovery speed.
- High availability handles small failures; disaster recovery handles loss of a site or region.
- Replication is not backup, and failover that was never tested often fails.
- Runbooks, automation and ownership matter as much as technology.
- Describe a drill: scope, trigger, failover, validation, failback and lessons.
Core concepts
1. What is disaster recovery?
The planning and technology used to restore services after a major disruption such as a regional outage, cyberattack or data centre loss, within agreed targets.
2. What are RPO and RTO?
RPO is the maximum acceptable data loss, measured in time. RTO is the maximum acceptable downtime. Shorter targets need more expensive designs.
3. What is the difference between high availability and disaster recovery?
High availability keeps a service running through component failures, usually within a region. Disaster recovery restores service after a larger failure, usually in another region or site.
4. What is the difference between disaster recovery and backup?
Backup protects data. Disaster recovery includes infrastructure, network, access, applications and the process to bring them back.
5. What is a business impact analysis?
A review of which services matter most, what their downtime costs and what dependencies they have. It sets priorities and RPO and RTO targets.
Strategies
6. Explain backup and restore as a DR strategy.
Data is backed up and the environment is rebuilt after a disaster. It is the cheapest option with the longest recovery time.
7. What is pilot light?
A minimal core of the environment, such as replicated data, runs continuously, and the rest is created during a disaster. It balances cost and speed.
8. What is warm standby?
A scaled-down but running copy of the full environment that is scaled up at failover. Recovery is faster, and cost is higher than pilot light.
9. What is active-active (multi-site)?
Several regions serve traffic at once, so loss of one has little impact. It gives the shortest recovery and the highest cost and complexity.
10. How do you choose a strategy?
Match it to RPO, RTO, budget, compliance and the skills of the team, per workload, not one choice for everything.
Design
11. How do you replicate data across regions?
Use database replication, storage replication or backup copies, and consider consistency, lag, bandwidth costs and data residency rules.
12. What is the risk of asynchronous replication?
Recent changes may be missing at failover, so RPO is above zero. Synchronous replication reduces this but adds latency.
13. How is traffic redirected during failover?
Through DNS changes with a suitable TTL, global load balancers or traffic managers, with health checks to trigger or guide the switch.
14. How do you handle infrastructure in the DR region?
Define it as code, so it can be created consistently, and keep images, secrets and configuration available there.
15. What about identity and access in a disaster?
Make sure identity services work in the DR site and that emergency access accounts exist, protected and tested.
16. How do you plan for ransomware as a disaster?
Keep immutable offline backups, isolate and investigate before restoring, restore from a known-clean point and rotate credentials.
Operations and testing
17. What is a runbook?
A step-by-step document, ideally partly automated, that tells responders how to detect, decide, fail over, validate and fail back.
18. What are failover and failback?
Failover moves service to the DR site. Failback returns it to the primary when it is healthy, and it needs the same care, including data synchronisation.
19. How often should you test DR?
On a defined schedule, and after major changes. Start with tabletop exercises, then partial failovers, then full drills.
20. What can go wrong in a DR test?
Missing permissions, stale images, DNS delays, forgotten dependencies, expired certificates and data that cannot be read in the new region.
21. How do you measure success?
Compare actual recovery time and data loss with RTO and RPO, and record gaps and fixes.
22. Who should be involved in DR?
Application owners, infrastructure, security, networking, communications and management, with clear roles and decision authority.
23. How do you handle communication during a disaster?
Prepare contact lists, templates and channels outside the affected systems, and give regular updates to stakeholders.
24. How do you control DR cost?
Use tiers, scale down standby systems, automate environment creation, and review storage and data transfer charges.
25. What is chaos engineering and how does it relate to DR?
Deliberately injecting controlled failures to see how systems respond. It builds confidence in resilience, but only in agreed, safe scopes.
A sample DR drill, step by step
- Choose one workload and agree scope, timing, success criteria and a rollback plan.
- Declare the scenario, such as loss of the primary region.
- Follow the runbook to start or scale the DR environment and restore the latest data.
- Redirect traffic and run functional and data checks.
- Record timings against RTO and the data position against RPO.
- Fail back, close the drill and write up the gaps and owners.
Common mistakes
- Having a plan but never testing it.
- Treating replication as backup.
- Forgetting dependencies such as DNS, identity and secrets.
- Choosing one DR strategy for every workload.
Next steps
Practise in the AWS training or the Azure administrator course. Related reading: cloud backup and recovery questions, cloud solutions architect questions and cloud operations questions.
Frequently Asked Questions
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0