ESC
Type to search guides, tutorials, and reference documentation.

Infrastructure as Code: State, Drift, and Blast Radius

How declarative infrastructure tools work, why a state file exists, how drift happens, how to split stacks so one mistake cannot take out everything, and the failure modes that show up only at scale.

Infrastructure as code is the practice of describing servers, networks, databases, permissions, and everything else that supports an application in machine-readable files, and letting a tool make the real environment match that description. The point is not automation for its own sake — scripts automate too. The point is that the description is reviewable, versioned, and reproducible, so an environment becomes something you can rebuild rather than something you hope keeps running.


What Infrastructure as Code Is

The defining test is whether you could delete an environment and recreate it from the repository. If the answer requires "and then someone clicks these five things in the console", the environment is not actually described in code, and the parts that are missing are exactly the parts that will fail during a recovery.

Two Different Tool Families

Provisioning / orchestration
  Creates and destroys cloud resources: networks, clusters, databases, IAM.
  Declarative, works against provider APIs, keeps a record of what it created.
  Examples: Terraform, OpenTofu, Pulumi, AWS CloudFormation, Bicep, Crossplane.

Configuration management
  Brings an existing machine to a desired configuration: packages, files, services.
  Convergent, agent or SSH driven, historically aimed at long-lived servers.
  Examples: Ansible, Chef, Puppet, Salt.

Image building
  Bakes a machine or container image once so instances start already correct.
  Examples: Packer, container image builds.

These solve different problems and are frequently combined incorrectly. Using a provisioning tool to run shell commands on a machine at create time produces resources whose real configuration is invisible to the tool that owns them. If a machine needs ongoing configuration, either bake it into an image and treat instances as immutable, or hand it to a configuration tool that actually models it.


The Declarative Model

A declarative tool works in three phases: read the desired state from your files, read the actual state of the world, and compute a plan that closes the gap. The plan is the product. Reviewing a plan before applying it is the single highest-value habit in the whole practice, because the plan is where an innocuous-looking edit reveals that it will destroy and recreate a database.

Why There Is a State File

Your configuration says "a network named app-vpc should exist". The cloud says "there is a network with id vpc-0a3f". Something has to remember that those are the same object, and that it was this configuration that created it. That mapping is the state. Without it the tool cannot tell the difference between "this resource needs creating" and "this resource exists and I own it", and it cannot know what to delete when you remove a block from your files.

config  ──▶  desired: resource "network" "app" { cidr = "10.0.0.0/16" }
state   ──▶  mapping:  network.app  ->  vpc-0a3f  (last known attributes)
provider──▶  actual:   vpc-0a3f exists, cidr 10.0.0.0/16, plus a tag added by hand

plan = diff(desired, actual, via state)
  ~ update in place        safe
  + create                 usually safe
  - destroy                pay attention
  -/+ replace              pay very close attention

Three consequences follow. State must be stored remotely and shared, or two engineers will hold divergent views of the same environment. It must be locked during an apply, or two concurrent runs will corrupt it. And it contains resource attributes, which frequently include sensitive values — so the state backend deserves the same protection as a secret store, not a public bucket.

Drift

Drift is the gap that opens when reality changes without the code changing: a console edit during an incident, an autoscaler adjusting a capacity field, a provider changing a default. Drift is not automatically a defect. What is dangerous is undetected drift, because the next apply will either revert someone's emergency fix or refuse to run at all, and you will find out at the worst moment. Detect it by running plan on a schedule and alerting on non-empty output; decide deliberately whether each difference should be adopted into code or corrected away.


Blast Radius and Stack Decomposition

The most consequential design decision in infrastructure as code is how much lives in one state. A single state holding an entire organisation makes every change a risk to everything, makes plans slow, and serialises every team behind one lock. Splitting too finely produces a web of cross-stack references nobody can reason about.

The useful split is by rate of change and by ownership. Foundational, rarely-changing, hard-to-replace things — accounts, networks, DNS zones, identity — belong in their own stacks with their own review discipline. Fast-moving application infrastructure belongs in stacks per service or per environment. Dependencies should point one way, from fast-moving to slow-moving, and cross-stack values should be read through a stable published interface rather than by reaching into another stack's internals.


Modules and Composition

Modules package a pattern so it can be reused with parameters. They earn their keep when a pattern is genuinely repeated and when the module encodes decisions you want made consistently — encryption on, logging on, sane defaults. They become a liability when they are created in advance of a second use case, because a module with one caller is indirection with no benefit, and because every parameter you add to accommodate a new caller makes the module harder to read than the resources it wraps.

Two practical rules keep modules healthy. Pin module and provider versions explicitly so an upstream change cannot alter a plan you did not ask to change. And keep a module's interface small: if callers must understand the module's internals to use it correctly, the abstraction has failed and inlining the resources is the honest fix.


Testing and Policy

Infrastructure code is testable at several levels, each cheaper than the one below. Static analysis and policy engines evaluate the configuration or the proposed plan against rules — no public storage buckets, encryption required, mandatory cost-allocation tags — and fail the pipeline before anything is created. Plan review catches destructive changes a rule cannot anticipate. And for modules that matter, an integration test applies into a throwaway environment, asserts against the real API, and destroys it.

Policy as code is particularly valuable because it converts review knowledge into an automated, non-negotiable gate. A reviewer who checks for a missing encryption flag will eventually miss one; a policy that refuses the plan will not. Wire the same checks into the delivery pipeline that runs the apply.


Failure Modes

Failure modeMechanismCorrection
One monolithic stateEvery change plans against everything; one lock for everyoneSplit by ownership and rate of change
Silent replacementChanging an immutable attribute destroys and recreates the resourceRead the plan; look for replace, not just update
Unpinned providers or modulesAn upstream release changes behaviour between two identical runsPin versions and upgrade deliberately
Secrets in state or in codeSensitive attributes recorded in plaintext state, or committedUse a secret manager; encrypt and access-control the state backend
Console editsOut-of-band change becomes drift nobody is watchingRestrict write access to the automation identity; scheduled drift detection
Apply without plan reviewAutomation applies whatever the diff turned out to bePlan on the pull request, apply the saved plan after approval
Copy-pasted environmentsStaging and production diverge silentlyOne definition, environment-specific inputs
No import pathExisting resources cannot be adopted, so code covers only the newImport deliberately; treat unmanaged resources as a tracked gap

When Not to Codify

Exploration is the clear exception: clicking through a console to understand a service is faster than writing a module for something you may not keep, provided the experiment is thrown away rather than promoted. Resources the provider's API does not model well, one-off manual approvals, and anything where the codified version would be a thin wrapper around a single API call that nobody will ever run twice are all reasonable to leave alone — as long as they are recorded as deliberately unmanaged rather than quietly forgotten.

Everything that must survive an account rebuild, a region failure, or the departure of the person who set it up should be in code. From there, the natural next steps are reconciling the applied state continuously with GitOps, and packaging the good defaults so teams get them without becoming infrastructure experts — the subject of platform engineering.