Infrastructure as Code: State, Drift, and Blast Radius
How declarative infrastructure tools work, why a state file exists, how drift happens, how to split stacks so one mistake cannot take out everything, and the failure modes that show up only at scale.
Infrastructure as code is the practice of describing servers, networks, databases, permissions, and everything else that supports an application in machine-readable files, and letting a tool make the real environment match that description. The point is not automation for its own sake — scripts automate too. The point is that the description is reviewable, versioned, and reproducible, so an environment becomes something you can rebuild rather than something you hope keeps running.
What Infrastructure as Code Is
The defining test is whether you could delete an environment and recreate it from the repository. If the answer requires "and then someone clicks these five things in the console", the environment is not actually described in code, and the parts that are missing are exactly the parts that will fail during a recovery.
Two Different Tool Families
Provisioning / orchestration
Creates and destroys cloud resources: networks, clusters, databases, IAM.
Declarative, works against provider APIs, keeps a record of what it created.
Examples: Terraform, OpenTofu, Pulumi, AWS CloudFormation, Bicep, Crossplane.
Configuration management
Brings an existing machine to a desired configuration: packages, files, services.
Convergent, agent or SSH driven, historically aimed at long-lived servers.
Examples: Ansible, Chef, Puppet, Salt.
Image building
Bakes a machine or container image once so instances start already correct.
Examples: Packer, container image builds.
These solve different problems and are frequently combined incorrectly. Using a provisioning tool to run shell commands on a machine at create time produces resources whose real configuration is invisible to the tool that owns them. If a machine needs ongoing configuration, either bake it into an image and treat instances as immutable, or hand it to a configuration tool that actually models it.
The Declarative Model
A declarative tool works in three phases: read the desired state from your files, read the actual state of the world, and compute a plan that closes the gap. The plan is the product. Reviewing a plan before applying it is the single highest-value habit in the whole practice, because the plan is where an innocuous-looking edit reveals that it will destroy and recreate a database.
Why There Is a State File
Your configuration says "a network named app-vpc should exist". The cloud says "there is a network with id vpc-0a3f". Something has to remember that those are the same object, and that it was this configuration that created it. That mapping is the state. Without it the tool cannot tell the difference between "this resource needs creating" and "this resource exists and I own it", and it cannot know what to delete when you remove a block from your files.
config ──▶ desired: resource "network" "app" { cidr = "10.0.0.0/16" }
state ──▶ mapping: network.app -> vpc-0a3f (last known attributes)
provider──▶ actual: vpc-0a3f exists, cidr 10.0.0.0/16, plus a tag added by hand
plan = diff(desired, actual, via state)
~ update in place safe
+ create usually safe
- destroy pay attention
-/+ replace pay very close attention
Three consequences follow. State must be stored remotely and shared, or two engineers will hold divergent views of the same environment. It must be locked during an apply, or two concurrent runs will corrupt it. And it contains resource attributes, which frequently include sensitive values — so the state backend deserves the same protection as a secret store, not a public bucket.
Drift
Drift is the gap that opens when reality changes without the code changing: a console edit during an incident, an autoscaler adjusting a capacity field, a provider changing a default. Drift is not automatically a defect. What is dangerous is undetected drift, because the next apply will either revert someone's emergency fix or refuse to run at all, and you will find out at the worst moment. Detect it by running plan on a schedule and alerting on non-empty output; decide deliberately whether each difference should be adopted into code or corrected away.
Blast Radius and Stack Decomposition
The most consequential design decision in infrastructure as code is how much lives in one state. A single state holding an entire organisation makes every change a risk to everything, makes plans slow, and serialises every team behind one lock. Splitting too finely produces a web of cross-stack references nobody can reason about.
The useful split is by rate of change and by ownership. Foundational, rarely-changing, hard-to-replace things — accounts, networks, DNS zones, identity — belong in their own stacks with their own review discipline. Fast-moving application infrastructure belongs in stacks per service or per environment. Dependencies should point one way, from fast-moving to slow-moving, and cross-stack values should be read through a stable published interface rather than by reaching into another stack's internals.
Modules and Composition
Modules package a pattern so it can be reused with parameters. They earn their keep when a pattern is genuinely repeated and when the module encodes decisions you want made consistently — encryption on, logging on, sane defaults. They become a liability when they are created in advance of a second use case, because a module with one caller is indirection with no benefit, and because every parameter you add to accommodate a new caller makes the module harder to read than the resources it wraps.
Two practical rules keep modules healthy. Pin module and provider versions explicitly so an upstream change cannot alter a plan you did not ask to change. And keep a module's interface small: if callers must understand the module's internals to use it correctly, the abstraction has failed and inlining the resources is the honest fix.
Testing and Policy
Infrastructure code is testable at several levels, each cheaper than the one below. Static analysis and policy engines evaluate the configuration or the proposed plan against rules — no public storage buckets, encryption required, mandatory cost-allocation tags — and fail the pipeline before anything is created. Plan review catches destructive changes a rule cannot anticipate. And for modules that matter, an integration test applies into a throwaway environment, asserts against the real API, and destroys it.
Policy as code is particularly valuable because it converts review knowledge into an automated, non-negotiable gate. A reviewer who checks for a missing encryption flag will eventually miss one; a policy that refuses the plan will not. Wire the same checks into the delivery pipeline that runs the apply.
Failure Modes
| Failure mode | Mechanism | Correction |
|---|---|---|
| One monolithic state | Every change plans against everything; one lock for everyone | Split by ownership and rate of change |
| Silent replacement | Changing an immutable attribute destroys and recreates the resource | Read the plan; look for replace, not just update |
| Unpinned providers or modules | An upstream release changes behaviour between two identical runs | Pin versions and upgrade deliberately |
| Secrets in state or in code | Sensitive attributes recorded in plaintext state, or committed | Use a secret manager; encrypt and access-control the state backend |
| Console edits | Out-of-band change becomes drift nobody is watching | Restrict write access to the automation identity; scheduled drift detection |
| Apply without plan review | Automation applies whatever the diff turned out to be | Plan on the pull request, apply the saved plan after approval |
| Copy-pasted environments | Staging and production diverge silently | One definition, environment-specific inputs |
| No import path | Existing resources cannot be adopted, so code covers only the new | Import deliberately; treat unmanaged resources as a tracked gap |
When Not to Codify
Exploration is the clear exception: clicking through a console to understand a service is faster than writing a module for something you may not keep, provided the experiment is thrown away rather than promoted. Resources the provider's API does not model well, one-off manual approvals, and anything where the codified version would be a thin wrapper around a single API call that nobody will ever run twice are all reasonable to leave alone — as long as they are recorded as deliberately unmanaged rather than quietly forgotten.
Everything that must survive an account rebuild, a region failure, or the departure of the person who set it up should be in code. From there, the natural next steps are reconciling the applied state continuously with GitOps, and packaging the good defaults so teams get them without becoming infrastructure experts — the subject of platform engineering.