During a massive flood in London, nearly every datacentre went down."

TL;DR

Rackspace was the only London datacentre to stay online during a major flood because they regularly tested their backup systems by simulating real failures. Practising failure recovery under controlled but challenging conditions built true resilience in their operations. Development managers should regularly test their systems and teams under realistic failure scenarios to ensure they can recover when it matters most.

16 May 2025
Written by Martin Hinshelwood
1 minute read
Comments
Subscribe

During a massive flood in London, nearly every datacentre went down.

Except one: Rackspace.

Their backup systems worked. Their power held. They stayed online. When asked why, the CEO didn’t present a whitepaper. He held up a key.

Every month, he walked into the power room and pulled the main breaker. On purpose. Shut the whole thing off. Live test. No excuses.

It was painful. Risky. Uncomfortable. And that’s exactly why it worked.

You don’t build resilience by talking about it. You build it by living through the failure, on your terms, not the disaster’s.

If your systems only work when conditions are perfect, you’ve already lost. If your team has never practised failing safely, they won’t know how to recover when it’s real.

Real resilience is earned through pain. If it’s hard, do it more often.

Smart Classifications

Each classification [Concepts, Categories, & Tags] was assigned using AI-powered semantic analysis and scored across relevance, depth, and alignment. Final decisions? Still human. Always traceable. Hover to see how it applies.

Comments
Subscribe

What to read next

Signal Engineering Excellence

Everyone has a disaster recovery plan, on paper

Most disaster recovery plans fail in practice due to overlooked dependencies and lack of real-world testing, leaving organisations …

Read article
Signal Engineering Excellence

When Heathrow went down, they blamed the power supplier

Heathrow’s outage was caused by an over-sensitive disaster recovery system, not a power loss, highlighting the risks of untested resilience …

Read article
Signal DevOps

Resilience is not a department

Resilience must be built into products from the start, ensuring they withstand failures like outages or network loss, rather than being …

Read article
Article Engineering Excellence Product Development

Fragile by Design: The Cost of Pretending to Be Resilient

Explores how poor engineering, shallow product thinking, and organisational denial lead to fragile systems, stressing that true resilience …

Read article
Article DevOps Engineering Excellence

How to Build for Business Resilience and Continuity

Learn key strategies for building business resilience and continuity, including observability, system decoupling, routine deployments, team …

Read article
Article Engineering Excellence Product Development

Resilience is Part of the Product, Not an Afterthought

Resilience must be designed into products from the start, not added later. Build systems to detect, contain, and recover from failures, …

Read article