Everyone has a disaster recovery plan, on paper

TL;DR

Many organizations have disaster recovery plans, but these often fail in real situations because critical dependencies are overlooked and not tested end to end. Successful drills can give a false sense of security if they do not include all essential systems, like authentication services. To ensure true resilience, regularly test your recovery process under real conditions and verify that all dependencies are restored and functional.

14 May 2025
Written by Martin Hinshelwood
1 minute read
Comments
Subscribe

Everyone has a disaster recovery plan, on paper.

But when the lights go out, very few organisations have systems that actually work. Spain, Portugal, Oracle, Heathrow… these weren’t random outages. They were textbook examples of systems that failed exactly as they were designed.

At Merrill Lynch, I saw it firsthand. Twice. We ran full disaster recovery drills. Everything looked like a success, until we tried to use the restored systems. Nothing worked. Why? Because the one system that everything else relied on, Active Directory, was never restored.

So yes, my app was “restored” successfully. But without authentication, it was useless.

This is the cost of pretending you’re resilient: you run the drills, tick the boxes, declare victory… and miss the critical dependencies that actually matter.

If you’ve never validated that your system works end to end, under real pressure, then you’re not resilient. You’re wishful.

Real resilience doesn’t live in plans. It lives in pain.

Smart Classifications

Each classification [Concepts, Categories, & Tags] was assigned using AI-powered semantic analysis and scored across relevance, depth, and alignment. Final decisions? Still human. Always traceable. Hover to see how it applies.

Comments
Subscribe

What to read next

Signal Leadership

During a massive flood in London, nearly every datacentre went down."

A London flood shut down most datacentres, but Rackspace stayed online by regularly live-testing failures, proving true resilience comes …

Read article
Signal DevOps

Resilience is not a department

Resilience must be built into products from the start, ensuring they withstand failures like outages or network loss, rather than being …

Read article
Signal Engineering Excellence

When Heathrow went down, they blamed the power supplier

Heathrow’s outage was caused by an over-sensitive disaster recovery system, not a power loss, highlighting the risks of untested resilience …

Read article
Article Engineering Excellence Product Development

Fragile by Design: The Cost of Pretending to Be Resilient

Explores how poor engineering, shallow product thinking, and organisational denial lead to fragile systems, stressing that true resilience …

Read article
Article DevOps Engineering Excellence

How to Build for Business Resilience and Continuity

Learn key strategies for building business resilience and continuity, including observability, system decoupling, routine deployments, team …

Read article
Article Engineering Excellence Product Development

Resilience is Part of the Product, Not an Afterthought

Resilience must be designed into products from the start, not added later. Build systems to detect, contain, and recover from failures, …

Read article