When Heathrow went down, they blamed the power supplier

TL;DR

Heathrow’s outage was not caused by a power loss but by an overly sensitive internal system that shut everything down in response to a minor fluctuation, revealing that their disaster recovery measures were untested for real-world chaos. Investing in infrastructure does not guarantee true resilience; resilience is proven only when systems are tested against unexpected failures. Development managers should regularly test and challenge their recovery processes to ensure they work under real conditions, not just ideal scenarios.

15 May 2025
Written by Martin Hinshelwood
1 minute read
Comments
Subscribe

When Heathrow went down, they blamed the power supplier.

A fire at one substation, they said, caused the disruption. Convenient story. But it wasn’t true.

Heathrow gets power from three independent substations. Any one of them could run the airport solo. The real failure? Their internal “resilience” system kicked in when it saw a fluctuation, not a loss. And in its attempt to “protect” the infrastructure, it shut the entire thing down.

It took all day to reboot. Not because power was unavailable, but because their disaster recovery system was too sensitive to survive a real incident.

That’s what happens when resilience is theatre. Flashy systems. Fancy architecture. And no one asking the hard question: what actually happens when things get messy?

If you haven’t tested for chaos, you’ve only prepared for comfort.

Don’t confuse infrastructure spend with resilience. Resilience is what happens when your assumptions fail.

Comments
Subscribe

What to read next

Signal Engineering Excellence

Everyone has a disaster recovery plan, on paper

Most disaster recovery plans fail in practice due to overlooked dependencies and lack of real-world testing, leaving organisations …

Read article
Signal Leadership

During a massive flood in London, nearly every datacentre went down."

A London flood shut down most datacentres, but Rackspace stayed online by regularly live-testing failures, proving true resilience comes …

Read article
Article Engineering Excellence Product Development

Fragile by Design: The Cost of Pretending to Be Resilient

Explores how poor engineering, shallow product thinking, and organisational denial lead to fragile systems, stressing that true resilience …

Read article
Signal DevOps

Resilience is not a department

Resilience must be built into products from the start, ensuring they withstand failures like outages or network loss, rather than being …

Read article
Article DevOps Engineering Excellence

How to Build for Business Resilience and Continuity

Learn key strategies for building business resilience and continuity, including observability, system decoupling, routine deployments, team …

Read article
Article Engineering Excellence Product Development

Resilience is Part of the Product, Not an Afterthought

Resilience must be designed into products from the start, not added later. Build systems to detect, contain, and recover from failures, …

Read article