ESC
Type to search guides, tutorials, and reference documentation.
Verified by Garnet Grid

Chaos Engineering

Injecting realistic faults as controlled experiments: defining steady state, forming a hypothesis, bounding blast radius, the fault classes that actually teach you something, game days, and the prerequisites that separate an experiment from a self-inflicted outage.

Chaos engineering is the practice of injecting realistic faults into a system as controlled experiments, in order to discover how it actually behaves under failure before failure chooses the timing. It exists because distributed systems have emergent behaviour: the way components fail together is not derivable from how each one fails alone, and it is not captured by any test suite that exercises components in isolation.


What Chaos Engineering Is

The name is unfortunate, because the practice is the opposite of chaotic. It is experimental method applied to a running system: you state what you believe, you create the conditions that would falsify it, you observe, and you stop if it goes wrong. Breaking things without a hypothesis is not chaos engineering; it is an outage you caused on purpose.

What it is for is unknown unknowns — the tolerances you believe you have but have never exercised. A retry policy that has never seen a slow dependency, a failover that has never been triggered in anger, a circuit breaker configured from a default, a runbook nobody has executed. Each of these is an untested assumption, and untested assumptions in a reliability mechanism are liabilities rather than defences.


The Experiment Structure

1. Steady state      Define 'normal' as a measurable output
                     (e.g. success rate, orders/min, p99 latency)
2. Hypothesis        'Under fault X, steady state stays within band Y'
3. Blast radius      Smallest scope that can still falsify it
4. Abort criteria    Written BEFORE the run; automated where possible
5. Inject            Apply the fault; keep it running long enough
                     for retries, timeouts and alerts to engage
6. Observe           Compare steady state against the control
7. Stop + restore    Verify the system returns to steady state
8. Act               Confirmed tolerance -> document it
                     Deviation -> defect with an owner, then re-run

Steady state is the load-bearing part. You must define normal in terms of an output that matters — successful checkouts per minute, request success rate, end-to-end latency distribution — and be able to measure it continuously and at low latency. An internal metric like CPU is not a steady-state definition, because the whole point is to detect whether users are affected. If you cannot measure your steady state well enough to notice a deviation within the experiment window, you are not ready to run the experiment; you are ready to improve your observability.

The hypothesis must be falsifiable and specific: "when a single instance of the pricing service becomes unavailable, checkout success rate stays within its normal band and p99 latency rises by no more than the configured retry budget allows". Vague hypotheses ("the system will be fine") produce experiments whose result nobody can dispute or act on.

The abort criteria must be written before the run, not judged during it. Under pressure, people rationalise a deviation as noise. Deciding in advance — "abort if success rate drops below X for more than N seconds" — removes that judgement from the moment when it is least reliable.


Blast Radius and Where to Run

Blast radius is the set of users, requests and resources an experiment can affect. The rule is to start with the smallest radius that could still falsify the hypothesis, and to expand only after the small version passes. Concretely: one instance before one zone; a shadow or internal traffic segment before real users; one percent of requests before all of them; a single non-critical dependency before a critical one.

There is an honest tradeoff about where to run. Pre-production environments are safe but frequently differ from production in exactly the dimensions that produce interesting failures: data volume, traffic mix and skew, topology, dependency versions, real third parties, and the presence of other tenants. A pass in staging is weak evidence. That is the argument for running in production — and it is only defensible with a bounded radius, a working kill switch, an informed on-call, and someone watching the steady-state signal in real time.

A kill switch that has itself never been tested is not a kill switch. Verify that the experiment can be stopped, and that stopping it actually restores the previous state, before you rely on it.


Fault Classes Worth Injecting

FaultWhat it testsTypically reveals
Instance or process terminationRedundancy, rebalancing, session handlingSticky state, unbalanced shards, slow re-election
Latency injection into a dependencyTimeouts, pools, backpressure, circuit breakersPool exhaustion and cascading slowness
Error injection into a dependencyError handling, fallbacks, retry policyRetry storms, missing fallbacks, error swallowing
Resource pressure (CPU, memory, disk)Degradation behaviour, limits, evictionNoisy-neighbour effects and abrupt kills
Network partitionSplit-brain handling, quorum, consistency choicesTwo sides both believing they are authoritative
Packet loss and jitterProtocol-level resilience, keepalivesHidden assumptions about a clean network
DNS or service-discovery failureResolution caching, startup pathsServices that cannot start, only continue
Zone or region evacuationCapacity headroom, failover, data localitySurvivors without enough headroom to absorb the load
Clock skewToken validity, ordering, scheduled workAuth failures and out-of-order processing
Certificate or credential expiryRotation automation, alerting lead timeRenewal that was never actually automated
Corrupt or wrong-but-successful responsesValidation, poison-message handlingBad data propagating quietly downstream

One asymmetry is worth stating plainly: slow is more informative than dead. Most systems handle a dependency that fails fast, because the connection is refused and the error path runs. A dependency that responds slowly holds connections, occupies threads and pool slots, and pushes the caller past its own timeouts — which is how a single slow component cascades into an outage. If you inject only one fault class, inject latency.

The second asymmetry: partial is worse than total. A total failure is unambiguous and triggers failover. A partial failure — some requests succeeding, one replica lagging, a node that responds to health checks but not to work — leaves the system unsure whether to fail over, and is where most surprising behaviour lives.


Game Days

A game day is a scheduled exercise where a team injects a scenario and works the response end to end. The technical result matters, but the organisational result is usually the larger finding: did the alert fire, did it reach the right person, was the runbook correct, could the incident commander be identified, did anyone know how to reach the dependency's owner, did the status update go out.

Run them with the people who would actually respond, not with the architects who designed the system. Announce them in advance until the response process is reliable; unannounced exercises are a later-maturity practice and require organisational trust that has been earned. Treat every finding as you would a real incident finding — same review, same owners, same tracking.


Prerequisites and When Not To

Chaos engineering is worth starting when you have a measurable steady state, a way to stop an experiment, an on-call process that works, and no large backlog of already-known reliability defects. That last condition is the one teams skip. If you already know your database has no tested failover and your retries have no backoff, you do not need an experiment to discover it — you need to fix it. Chaos is for finding what you do not already know.

Do not run experiments against systems where the failure has irreversible external consequences — moving money, sending messages to customers, actuating physical equipment, deleting data — unless the affected path is genuinely simulated. Do not run them without the on-call rotation knowing, or you will convert an experiment into an incident response with a false cause. Do not run them during a change freeze, a launch, or while another incident is open. And do not run them without a named person who has both the authority and the ability to stop the experiment.


Failure Modes

Failure modeWhy it failsCounter
No hypothesisBreaking things produces a story, not a findingWrite the falsifiable claim first
Unmeasurable steady stateYou cannot tell whether the fault matteredFix observability before running experiments
Blast radius too large on the first runAn experiment becomes an incidentSmallest scope that can falsify; expand only after a pass
Untested kill switchYou cannot stop what you startedExercise the abort path before relying on it
Injecting a fault the system never claimed to tolerateProduces a predictable failure, not informationTest stated tolerances; unclaimed ones are design work
Chaos theatreExperiments run, results filed, nothing changesFindings enter the same backlog as incident actions
No re-run after the fixThe repair is unverifiedClose a finding only when the original fault stops causing deviation
Staging-only foreverResults do not transfer to production topologyPlan the path to bounded production experiments

The outcome that makes the practice worth its cost is narrow and concrete: every experiment either confirms a tolerance — which is now a documented, exercised property rather than an assumption — or produces a specific defect with a specific fix. And a fix is not proven by a code review; it is proven by re-running the same experiment and observing that the steady state no longer deviates. Verifying the repair with the original fault is what separates this from a testing exercise that generates reports.

Jakub Dimitri Rezayev
Jakub Dimitri Rezayev
Founder & Chief Architect • Garnet Grid Consulting

Jakub holds an M.S. in Customer Intelligence & Analytics and a B.S. in Finance & Computer Science from Pace University. With deep expertise spanning D365 F&O, Azure, Power BI, and AI/ML systems, he architects enterprise solutions that bridge legacy systems and modern technology — and has led multi-million dollar ERP implementations for Fortune 500 supply chains.

View Full Profile →