ESC
Type to search guides, tutorials, and reference documentation.

CI/CD Pipelines: Continuous Integration and Continuous Delivery

How a change travels from a commit to production: build reproducibility, artifact promotion, test ordering, deployment strategies, and the failure modes that quietly make a pipeline untrustworthy.

A CI/CD pipeline is the automated path a change takes from a developer's commit to a running system. Continuous integration is the discipline of merging every change into a shared mainline frequently and proving, automatically, that the mainline still builds and passes its tests. Continuous delivery is the discipline of keeping every one of those verified builds in a state where it could be released on demand. The pipeline is the machinery that makes both claims credible without a human having to remember the steps.


What CI/CD Actually Is

The value of CI is not the build server. It is the short feedback loop and the shared mainline. If branches live long enough to diverge substantially, the integration pain returns no matter how much automation sits on top of it — the merge conflict is deferred, not prevented. CI is a working agreement (small changes, merged often, trunk always green) that a build server enforces.

Three Terms That Get Confused

Continuous Integration
  Every change merges to mainline frequently.
  Every merge triggers build + automated verification.
  A red mainline is the team's top priority.

Continuous Delivery
  Every verified build is a release candidate.
  Deploying is a business decision, not an engineering project.
  The release path itself is exercised constantly, so it is boring.

Continuous Deployment
  Every verified build goes to production automatically.
  No human gate. Requires very high confidence in the tests,
  strong observability, and a fast, automatic rollback path.

Most organizations want continuous delivery. Continuous deployment is a further step that only pays off when detection and rollback are fast enough that a bad change is a minor, self-correcting incident rather than an outage.


Anatomy of a Pipeline

A pipeline is a sequence of gates, each one cheaper to fail than the next. The ordering principle is simple: run the check that is fastest and most likely to fail as early as possible, so the expensive stages only ever run on changes that have already earned them.

Build Once, Promote the Artifact

The single most important structural rule is that the thing you tested must be the thing you ship. Build the artifact exactly once, give it an immutable identity, and then promote that same artifact through environments by changing configuration only. Rebuilding per environment reintroduces the class of bug the pipeline exists to eliminate: staging and production differ, and nobody knows how.

commit ──▶ build ──▶ artifact:sha-9f2c1b  (immutable, content-addressed)
                          │
                          ├──▶ deploy to dev      (same artifact, dev config)
                          ├──▶ deploy to staging  (same artifact, staging config)
                          └──▶ deploy to prod     (same artifact, prod config)

Config comes from the environment. Secrets come from a secret store.
Neither is baked into the artifact.

Reproducibility is what makes that identity meaningful. Pin dependency versions with a lockfile, pin base images by digest rather than by a moving tag, and keep the build agent's toolchain declared rather than installed by hand. A build that produces different output from the same input is a build you cannot reason about during an incident.

Ordering Tests by Cost

Stage 1  lint, type check, unit tests          seconds      runs on every push
Stage 2  build + component/contract tests       a few min    runs on every push
Stage 3  integration tests against real deps    longer       runs pre-merge
Stage 4  deploy to staging + smoke/e2e          longest      runs post-merge
Stage 5  progressive rollout to production      continuous   watched by alerts

Contract tests deserve particular attention in a service architecture. They let each service verify its side of an interface independently, which keeps the slow, brittle, end-to-end suite small instead of letting it grow into the only thing anyone trusts.


Deployment Strategies

How the new version replaces the old is a separate decision from how it was built and tested. The strategies differ mainly in what they cost and how quickly they can be undone.

StrategyMechanismCostRollback
RecreateStop old, start newCheapestRedeploy old; downtime both ways
RollingReplace instances in batchesLowRoll forward or roll back through the same slow loop
Blue/greenTwo full environments, switch traffic at onceTwo environments at onceSwitch traffic back; near-instant
CanarySmall traffic share to the new version, widen on healthy signalsNeeds routing plus good metricsShift traffic away; blast radius already small
Feature flagShip code dark, enable behaviour separatelyFlag lifecycle debtFlip the flag; no redeploy

Two constraints cut across all of them. First, a rollback is only real if the new version's database changes are backward compatible with the old code — which is why schema changes are normally expand/migrate/contract across several releases rather than one destructive step. Second, canary and blue/green are only as good as the signal you promote on; automating promotion against a metric nobody trusts just automates a bad decision. Decoupling deployment from release with feature flags is often the cheapest way to make production changes reversible.


The Pipeline Is Production Code

Pipeline definitions belong in the repository they build, reviewed like any other change. Treating them as configuration that lives in a web console produces infrastructure nobody can diff, review, or restore. The same reasoning extends to the runners: an agent that accumulated its toolchain by hand is a snowflake, and the day it dies you discover which undocumented package the build depended on.

A pipeline also holds credentials and can write to production, which makes it a high-value target. Scope its credentials to the narrowest action it needs, prefer short-lived federated identity over long-lived static keys, and be deliberate about which events can run privileged jobs — a workflow that runs untrusted code from a fork with access to release secrets is a supply-chain incident waiting to be discovered.


Common Failure Modes

Failure modeWhat it looks likeCorrection
Flaky testsRe-run until green becomes normalQuarantine and fix or delete; a test nobody trusts is worse than no test
Slow feedbackDevelopers context-switch away and batch changesParallelize, cache dependencies, move slow checks off the pre-merge path
Rebuild per environmentWorks in staging, fails in productionBuild once, promote the artifact
Approval theatreA required sign-off nobody has context to refuseReplace with automated policy checks plus a fast rollback
Shared mutable stagingQueueing, contention, and results nobody believesEphemeral per-change environments
Untested rollback pathThe first real rollback is attempted during an outageExercise rollback regularly, including the data layer
Green build, unshippablePassing tests that never exercise startup or configSmoke-test the deployed artifact, not just the code

Measuring a Pipeline

Four widely used delivery measures are deployment frequency, lead time for changes (commit to running in production), change failure rate, and time to restore service. They are useful because they resist gaming in opposite directions: speed measures alone reward recklessness, stability measures alone reward never shipping, and the pair together describes throughput and safety at once.

Measure them on your own system and watch the trend rather than comparing against an external benchmark. The diagnostic value comes from decomposing lead time into its stages — waiting for review, waiting for CI, waiting for a release window — because the longest stage is almost always the one worth fixing, and it is frequently not the one people assume.


When Not to Invest Further

Pipeline work has diminishing returns like anything else. If the constraint is a long manual QA cycle, a regulator-mandated release window, a shared environment nobody can provision, or a coupled monolith where every change requires a full-system regression, faster builds will not help; the queue simply moves. Firmware, embedded systems, and shipped client software have genuinely different economics, and a release train can be the correct answer there.

The honest check is to look at where time actually goes between commit and production. If most of it is spent waiting on a human decision, invest in decoupling and reversibility — GitOps for a reconciled delivery half, infrastructure as code for reproducible environments, and platform engineering for the templates that make a good pipeline the default rather than a per-team achievement.