CI/CD Pipelines: Continuous Integration and Continuous Delivery
How a change travels from a commit to production: build reproducibility, artifact promotion, test ordering, deployment strategies, and the failure modes that quietly make a pipeline untrustworthy.
A CI/CD pipeline is the automated path a change takes from a developer's commit to a running system. Continuous integration is the discipline of merging every change into a shared mainline frequently and proving, automatically, that the mainline still builds and passes its tests. Continuous delivery is the discipline of keeping every one of those verified builds in a state where it could be released on demand. The pipeline is the machinery that makes both claims credible without a human having to remember the steps.
What CI/CD Actually Is
The value of CI is not the build server. It is the short feedback loop and the shared mainline. If branches live long enough to diverge substantially, the integration pain returns no matter how much automation sits on top of it — the merge conflict is deferred, not prevented. CI is a working agreement (small changes, merged often, trunk always green) that a build server enforces.
Three Terms That Get Confused
Continuous Integration
Every change merges to mainline frequently.
Every merge triggers build + automated verification.
A red mainline is the team's top priority.
Continuous Delivery
Every verified build is a release candidate.
Deploying is a business decision, not an engineering project.
The release path itself is exercised constantly, so it is boring.
Continuous Deployment
Every verified build goes to production automatically.
No human gate. Requires very high confidence in the tests,
strong observability, and a fast, automatic rollback path.
Most organizations want continuous delivery. Continuous deployment is a further step that only pays off when detection and rollback are fast enough that a bad change is a minor, self-correcting incident rather than an outage.
Anatomy of a Pipeline
A pipeline is a sequence of gates, each one cheaper to fail than the next. The ordering principle is simple: run the check that is fastest and most likely to fail as early as possible, so the expensive stages only ever run on changes that have already earned them.
Build Once, Promote the Artifact
The single most important structural rule is that the thing you tested must be the thing you ship. Build the artifact exactly once, give it an immutable identity, and then promote that same artifact through environments by changing configuration only. Rebuilding per environment reintroduces the class of bug the pipeline exists to eliminate: staging and production differ, and nobody knows how.
commit ──▶ build ──▶ artifact:sha-9f2c1b (immutable, content-addressed)
│
├──▶ deploy to dev (same artifact, dev config)
├──▶ deploy to staging (same artifact, staging config)
└──▶ deploy to prod (same artifact, prod config)
Config comes from the environment. Secrets come from a secret store.
Neither is baked into the artifact.
Reproducibility is what makes that identity meaningful. Pin dependency versions with a lockfile, pin base images by digest rather than by a moving tag, and keep the build agent's toolchain declared rather than installed by hand. A build that produces different output from the same input is a build you cannot reason about during an incident.
Ordering Tests by Cost
Stage 1 lint, type check, unit tests seconds runs on every push
Stage 2 build + component/contract tests a few min runs on every push
Stage 3 integration tests against real deps longer runs pre-merge
Stage 4 deploy to staging + smoke/e2e longest runs post-merge
Stage 5 progressive rollout to production continuous watched by alerts
Contract tests deserve particular attention in a service architecture. They let each service verify its side of an interface independently, which keeps the slow, brittle, end-to-end suite small instead of letting it grow into the only thing anyone trusts.
Deployment Strategies
How the new version replaces the old is a separate decision from how it was built and tested. The strategies differ mainly in what they cost and how quickly they can be undone.
| Strategy | Mechanism | Cost | Rollback |
|---|---|---|---|
| Recreate | Stop old, start new | Cheapest | Redeploy old; downtime both ways |
| Rolling | Replace instances in batches | Low | Roll forward or roll back through the same slow loop |
| Blue/green | Two full environments, switch traffic at once | Two environments at once | Switch traffic back; near-instant |
| Canary | Small traffic share to the new version, widen on healthy signals | Needs routing plus good metrics | Shift traffic away; blast radius already small |
| Feature flag | Ship code dark, enable behaviour separately | Flag lifecycle debt | Flip the flag; no redeploy |
Two constraints cut across all of them. First, a rollback is only real if the new version's database changes are backward compatible with the old code — which is why schema changes are normally expand/migrate/contract across several releases rather than one destructive step. Second, canary and blue/green are only as good as the signal you promote on; automating promotion against a metric nobody trusts just automates a bad decision. Decoupling deployment from release with feature flags is often the cheapest way to make production changes reversible.
The Pipeline Is Production Code
Pipeline definitions belong in the repository they build, reviewed like any other change. Treating them as configuration that lives in a web console produces infrastructure nobody can diff, review, or restore. The same reasoning extends to the runners: an agent that accumulated its toolchain by hand is a snowflake, and the day it dies you discover which undocumented package the build depended on.
A pipeline also holds credentials and can write to production, which makes it a high-value target. Scope its credentials to the narrowest action it needs, prefer short-lived federated identity over long-lived static keys, and be deliberate about which events can run privileged jobs — a workflow that runs untrusted code from a fork with access to release secrets is a supply-chain incident waiting to be discovered.
Common Failure Modes
| Failure mode | What it looks like | Correction |
|---|---|---|
| Flaky tests | Re-run until green becomes normal | Quarantine and fix or delete; a test nobody trusts is worse than no test |
| Slow feedback | Developers context-switch away and batch changes | Parallelize, cache dependencies, move slow checks off the pre-merge path |
| Rebuild per environment | Works in staging, fails in production | Build once, promote the artifact |
| Approval theatre | A required sign-off nobody has context to refuse | Replace with automated policy checks plus a fast rollback |
| Shared mutable staging | Queueing, contention, and results nobody believes | Ephemeral per-change environments |
| Untested rollback path | The first real rollback is attempted during an outage | Exercise rollback regularly, including the data layer |
| Green build, unshippable | Passing tests that never exercise startup or config | Smoke-test the deployed artifact, not just the code |
Measuring a Pipeline
Four widely used delivery measures are deployment frequency, lead time for changes (commit to running in production), change failure rate, and time to restore service. They are useful because they resist gaming in opposite directions: speed measures alone reward recklessness, stability measures alone reward never shipping, and the pair together describes throughput and safety at once.
Measure them on your own system and watch the trend rather than comparing against an external benchmark. The diagnostic value comes from decomposing lead time into its stages — waiting for review, waiting for CI, waiting for a release window — because the longest stage is almost always the one worth fixing, and it is frequently not the one people assume.
When Not to Invest Further
Pipeline work has diminishing returns like anything else. If the constraint is a long manual QA cycle, a regulator-mandated release window, a shared environment nobody can provision, or a coupled monolith where every change requires a full-system regression, faster builds will not help; the queue simply moves. Firmware, embedded systems, and shipped client software have genuinely different economics, and a release train can be the correct answer there.
The honest check is to look at where time actually goes between commit and production. If most of it is spent waiting on a human decision, invest in decoupling and reversibility — GitOps for a reconciled delivery half, infrastructure as code for reproducible environments, and platform engineering for the templates that make a good pipeline the default rather than a per-team achievement.