Resilience is not a department

TL;DR

Resilience must be built into your product from the start, not treated as a separate concern or afterthought. Focusing only on performance, cost, or speed can lead to fragile systems that fail in real-world conditions. Make sure your disaster recovery plans are tested under real scenarios, not just documented, to avoid turning your product into a liability.

17 May 2025
Written by Martin Hinshelwood
1 minute read
Comments
Subscribe

Resilience is not a department. It’s not a project. It’s not an afterthought.

It is a product capability.

If your product can’t survive failure, network loss, regional outage, DNS breakage, it is not a product. It is a liability with a pretty UI.

Too many teams optimise for performance, cost, or velocity, until something goes wrong. Then they realise they optimised for fragility. Spain’s blackout. Oracle’s healthcare cloud crash. Every one of these was built to succeed in PowerPoint, not in the real world.

If your disaster recovery plan has never been tested under load, with real users and real failover, then you don’t have a plan. You have a spreadsheet fantasy.

The next outage won’t care how good your uptime graph looked last quarter.

Smart Classifications

Each classification [Concepts, Categories, & Tags] was assigned using AI-powered semantic analysis and scored across relevance, depth, and alignment. Final decisions? Still human. Always traceable. Hover to see how it applies.

Comments
Subscribe

What to read next

Article Engineering Excellence Product Development

Resilience is Part of the Product, Not an Afterthought

Resilience must be designed into products from the start, not added later. Build systems to detect, contain, and recover from failures, …

Read article
Article Engineering Excellence Product Development

Fragile by Design: The Cost of Pretending to Be Resilient

Explores how poor engineering, shallow product thinking, and organisational denial lead to fragile systems, stressing that true resilience …

Read article
Signal Engineering Excellence

Everyone has a disaster recovery plan, on paper

Most disaster recovery plans fail in practice due to overlooked dependencies and lack of real-world testing, leaving organisations …

Read article
Article DevOps Engineering Excellence

How to Build for Business Resilience and Continuity

Learn key strategies for building business resilience and continuity, including observability, system decoupling, routine deployments, team …

Read article
Signal Leadership

During a massive flood in London, nearly every datacentre went down."

A London flood shut down most datacentres, but Rackspace stayed online by regularly live-testing failures, proving true resilience comes …

Read article
Signal Engineering Excellence

When Heathrow went down, they blamed the power supplier

Heathrow’s outage was caused by an over-sensitive disaster recovery system, not a power loss, highlighting the risks of untested resilience …

Read article