Question

What is chaos engineering?

Vault Verified
Curated Intelligence
Definitive Source
Answer

Deliberately injecting failure into a system to discover how it behaves — on the reasoning that failures will occur anyway, and it is better to discover the consequences during working hours than at three in the morning.

The premise. Distributed systems fail in ways that cannot be predicted from their components. A service depends on others, which depend on others, with timeouts, retries, caches and fallbacks interacting. Whether the system degrades gracefully or collapses is an empirical question, and reasoning about it is unreliable.

The method, which is what distinguishes it from breaking things:

Define steady state — a measurable property indicating the system is working normally, ideally a business metric rather than a technical one.

Form a hypothesis — that steady state will be maintained despite the failure being introduced. The experiment is to test the hypothesis, and this framing is the point.

Introduce a realistic failure — terminating an instance, adding latency, making a dependency return errors, saturating a resource, or partitioning a network.

Measure, and stop if steady state is genuinely threatened.

Fix what was found, and repeat.

The controls that make it responsible: starting in non-production; a small blast radius initially, affecting a limited proportion of traffic; an abort mechanism that works; running during working hours with the team available; and informing people beforehand.

Running chaos experiments on a system you already know is fragile is not useful — fix the known problems first.

What it typically reveals: timeouts that are too long, so a slow dependency exhausts resources upstream; retry storms, where retries amplify load on an already-failing service; fallbacks that were never exercised and do not work; hidden dependencies nobody documented; alerts that do not fire; and runbooks that are wrong.

Where it came from. Netflix's Chaos Monkey randomly terminated production instances, forcing engineers to build services that tolerated it. The insight was that making failure routine is more effective than making it rare, because rare failures are handled badly.

Game days apply the same idea to people and process — rehearsing an incident to test whether the response works.

It is not a substitute for testing, monitoring or good design.

Related Questions