Chaos Engineering
Injecting realistic faults as controlled experiments: defining steady state, forming a hypothesis, bounding blast radius, the fault classes that actually teach you something, game days, and the prerequisites that separate an experiment from a self-inflicted outage.
Chaos engineering is the practice of injecting realistic faults into a system as controlled experiments, in order to discover how it actually behaves under failure before failure chooses the timing. It exists because distributed systems have emergent behaviour: the way components fail together is not derivable from how each one fails alone, and it is not captured by any test suite that exercises components in isolation.
What Chaos Engineering Is
The name is unfortunate, because the practice is the opposite of chaotic. It is experimental method applied to a running system: you state what you believe, you create the conditions that would falsify it, you observe, and you stop if it goes wrong. Breaking things without a hypothesis is not chaos engineering; it is an outage you caused on purpose.
What it is for is unknown unknowns — the tolerances you believe you have but have never exercised. A retry policy that has never seen a slow dependency, a failover that has never been triggered in anger, a circuit breaker configured from a default, a runbook nobody has executed. Each of these is an untested assumption, and untested assumptions in a reliability mechanism are liabilities rather than defences.
The Experiment Structure
1. Steady state Define 'normal' as a measurable output
(e.g. success rate, orders/min, p99 latency)
2. Hypothesis 'Under fault X, steady state stays within band Y'
3. Blast radius Smallest scope that can still falsify it
4. Abort criteria Written BEFORE the run; automated where possible
5. Inject Apply the fault; keep it running long enough
for retries, timeouts and alerts to engage
6. Observe Compare steady state against the control
7. Stop + restore Verify the system returns to steady state
8. Act Confirmed tolerance -> document it
Deviation -> defect with an owner, then re-run
Steady state is the load-bearing part. You must define normal in terms of an output that matters — successful checkouts per minute, request success rate, end-to-end latency distribution — and be able to measure it continuously and at low latency. An internal metric like CPU is not a steady-state definition, because the whole point is to detect whether users are affected. If you cannot measure your steady state well enough to notice a deviation within the experiment window, you are not ready to run the experiment; you are ready to improve your observability.
The hypothesis must be falsifiable and specific: "when a single instance of the pricing service becomes unavailable, checkout success rate stays within its normal band and p99 latency rises by no more than the configured retry budget allows". Vague hypotheses ("the system will be fine") produce experiments whose result nobody can dispute or act on.
The abort criteria must be written before the run, not judged during it. Under pressure, people rationalise a deviation as noise. Deciding in advance — "abort if success rate drops below X for more than N seconds" — removes that judgement from the moment when it is least reliable.
Blast Radius and Where to Run
Blast radius is the set of users, requests and resources an experiment can affect. The rule is to start with the smallest radius that could still falsify the hypothesis, and to expand only after the small version passes. Concretely: one instance before one zone; a shadow or internal traffic segment before real users; one percent of requests before all of them; a single non-critical dependency before a critical one.
There is an honest tradeoff about where to run. Pre-production environments are safe but frequently differ from production in exactly the dimensions that produce interesting failures: data volume, traffic mix and skew, topology, dependency versions, real third parties, and the presence of other tenants. A pass in staging is weak evidence. That is the argument for running in production — and it is only defensible with a bounded radius, a working kill switch, an informed on-call, and someone watching the steady-state signal in real time.
A kill switch that has itself never been tested is not a kill switch. Verify that the experiment can be stopped, and that stopping it actually restores the previous state, before you rely on it.
Fault Classes Worth Injecting
| Fault | What it tests | Typically reveals |
|---|---|---|
| Instance or process termination | Redundancy, rebalancing, session handling | Sticky state, unbalanced shards, slow re-election |
| Latency injection into a dependency | Timeouts, pools, backpressure, circuit breakers | Pool exhaustion and cascading slowness |
| Error injection into a dependency | Error handling, fallbacks, retry policy | Retry storms, missing fallbacks, error swallowing |
| Resource pressure (CPU, memory, disk) | Degradation behaviour, limits, eviction | Noisy-neighbour effects and abrupt kills |
| Network partition | Split-brain handling, quorum, consistency choices | Two sides both believing they are authoritative |
| Packet loss and jitter | Protocol-level resilience, keepalives | Hidden assumptions about a clean network |
| DNS or service-discovery failure | Resolution caching, startup paths | Services that cannot start, only continue |
| Zone or region evacuation | Capacity headroom, failover, data locality | Survivors without enough headroom to absorb the load |
| Clock skew | Token validity, ordering, scheduled work | Auth failures and out-of-order processing |
| Certificate or credential expiry | Rotation automation, alerting lead time | Renewal that was never actually automated |
| Corrupt or wrong-but-successful responses | Validation, poison-message handling | Bad data propagating quietly downstream |
One asymmetry is worth stating plainly: slow is more informative than dead. Most systems handle a dependency that fails fast, because the connection is refused and the error path runs. A dependency that responds slowly holds connections, occupies threads and pool slots, and pushes the caller past its own timeouts — which is how a single slow component cascades into an outage. If you inject only one fault class, inject latency.
The second asymmetry: partial is worse than total. A total failure is unambiguous and triggers failover. A partial failure — some requests succeeding, one replica lagging, a node that responds to health checks but not to work — leaves the system unsure whether to fail over, and is where most surprising behaviour lives.
Game Days
A game day is a scheduled exercise where a team injects a scenario and works the response end to end. The technical result matters, but the organisational result is usually the larger finding: did the alert fire, did it reach the right person, was the runbook correct, could the incident commander be identified, did anyone know how to reach the dependency's owner, did the status update go out.
Run them with the people who would actually respond, not with the architects who designed the system. Announce them in advance until the response process is reliable; unannounced exercises are a later-maturity practice and require organisational trust that has been earned. Treat every finding as you would a real incident finding — same review, same owners, same tracking.
Prerequisites and When Not To
Chaos engineering is worth starting when you have a measurable steady state, a way to stop an experiment, an on-call process that works, and no large backlog of already-known reliability defects. That last condition is the one teams skip. If you already know your database has no tested failover and your retries have no backoff, you do not need an experiment to discover it — you need to fix it. Chaos is for finding what you do not already know.
Do not run experiments against systems where the failure has irreversible external consequences — moving money, sending messages to customers, actuating physical equipment, deleting data — unless the affected path is genuinely simulated. Do not run them without the on-call rotation knowing, or you will convert an experiment into an incident response with a false cause. Do not run them during a change freeze, a launch, or while another incident is open. And do not run them without a named person who has both the authority and the ability to stop the experiment.
Failure Modes
| Failure mode | Why it fails | Counter |
|---|---|---|
| No hypothesis | Breaking things produces a story, not a finding | Write the falsifiable claim first |
| Unmeasurable steady state | You cannot tell whether the fault mattered | Fix observability before running experiments |
| Blast radius too large on the first run | An experiment becomes an incident | Smallest scope that can falsify; expand only after a pass |
| Untested kill switch | You cannot stop what you started | Exercise the abort path before relying on it |
| Injecting a fault the system never claimed to tolerate | Produces a predictable failure, not information | Test stated tolerances; unclaimed ones are design work |
| Chaos theatre | Experiments run, results filed, nothing changes | Findings enter the same backlog as incident actions |
| No re-run after the fix | The repair is unverified | Close a finding only when the original fault stops causing deviation |
| Staging-only forever | Results do not transfer to production topology | Plan the path to bounded production experiments |
The outcome that makes the practice worth its cost is narrow and concrete: every experiment either confirms a tolerance — which is now a documented, exercised property rather than an assumption — or produces a specific defect with a specific fix. And a fix is not proven by a code review; it is proven by re-running the same experiment and observing that the steady state no longer deviates. Verifying the repair with the original fault is what separates this from a testing exercise that generates reports.