Agile & Operations
How iterative delivery actually works: feedback economics, work-in-progress limits and flow, honest delivery metrics, the link to operations, and the failure modes that turn agile practice into ceremony.
Agile is a set of working agreements for delivering software in small, frequent increments so that feedback arrives while there is still time and budget to act on it. It is not a methodology, a certification, or a particular schedule of meetings — those are implementations of the idea, and they are frequently mistaken for it. This guide covers the mechanics that make iterative delivery work, the measurements that tell you whether it is working, and the failure modes that turn the practice into ceremony.
What Agile Actually Is
The central claim of iterative delivery is economic rather than moral: in software, the cost of discovering that you built the wrong thing rises steeply with the time taken to discover it. Requirements drift, markets move, and the people who wrote the specification learn things after they write it. Shipping a small usable slice and observing what happens converts an expensive late correction into a cheap early one.
Nearly every agile practice follows from that one claim. Short iterations exist to force a feedback event onto a schedule rather than leaving it to chance. Working software is the unit of progress because a document cannot be wrong in a way a user will notice. Cross-functional teams exist so that a slice can be finished without passing through a queue of handoffs. Retrospectives exist because the process itself is one of the things being iterated on.
Agile Versus a Framework
Scrum, Kanban, Extreme Programming and the various scaled frameworks are implementations of that idea with different opinions about structure. Scrum is prescriptive about cadence, accountabilities and a fixed set of events. Kanban is prescriptive about visualising flow and limiting work in progress, and deliberately silent about roles and iteration length. Extreme Programming is the most prescriptive about engineering practice itself — pairing, test-first development, continuous integration, collective code ownership.
Treating the framework as the goal is the most common and most expensive mistake in this area. A team can hold every event on schedule and still ship once a year. A team running no named framework at all can deliver weekly with a tight feedback loop. The useful question to ask of any practice is: what feedback does this generate, and which decision does that feedback change? A practice that answers neither question is overhead, regardless of which framework recommends it.
The Core Mechanics
Iteration and Feedback
An iteration is a timebox that ends in something a stakeholder can look at and react to. Its critical property is that the box is fixed and the scope flexes. Reversing that — fixing the scope and letting the date move — reintroduces exactly the late-discovery problem iteration exists to prevent, while keeping the meetings.
The length of the iteration should be set by how long you can afford to be wrong. Shorter iterations cost more in fixed overhead per cycle and give faster correction; longer ones amortise the overhead and let error accumulate. Teams with expensive, slow release paths often pick long iterations to reduce the pain, which is the wrong lever: the fix is to make releasing cheap.
Work in Progress and Flow
Little's Law states that for a stable system, the average number of items in the system equals the arrival rate multiplied by the average time an item spends in it. This is a queueing identity, not a productivity heuristic, and its consequence for delivery is unavoidable: holding throughput constant, the more work you have in progress, the longer each item takes to finish.
Starting more work therefore does not finish more work. It lengthens every cycle time, delays every feedback event, and multiplies the context-switching cost paid by the people doing it. Explicit limits on work in progress are among the cheapest interventions available, because they require no new tooling and no reorganisation — only the discipline to finish something before starting the next thing.
The Backlog as a Forecast
A backlog is an ordered set of options, not a set of promises. The ordering carries almost all of the value; the estimates carry much less than teams assume. Summing optimistic per-item estimates produces a plan that is confidently precise and routinely wrong, because it ignores the variance that dominates knowledge work.
The more honest approach is to forecast from the team's own observed history — how many items of comparable size this team actually completed per unit of time, expressed as a range rather than a number. That range is uncomfortable to present and far more useful than a single date, because it makes the uncertainty a property of the plan rather than a surprise discovered later.
Measuring Delivery Honestly
The four delivery measures popularised by the DORA research programme — deployment frequency, lead time for changes, change failure rate, and time to restore service — are useful largely because they are hard to improve in isolation. Deploy more often while cutting corners and the change failure rate moves. Stabilise by deploying rarely and frequency and lead time move. Watching the set is more informative than optimising any member of it.
Velocity deserves a warning of its own. It is a capacity-planning artefact for a single team using a single estimation convention. It is not a productivity measure, and comparing it across teams is meaningless because the unit is locally defined. The moment velocity becomes a target for individuals or a comparison between teams, estimates inflate, the number rises, and delivery does not change — the standard behaviour of any measure converted into a target.
Metrics should open conversations, not close them. A dashboard that nobody uses to ask a question is reporting, not measurement.
Where Agile Meets Operations
Iterative delivery generates a continuous stream of change, and change is the dominant source of production incidents in most systems. This is why delivery practice and operational practice are the same conversation rather than adjacent ones. A team that increases its release frequency without also investing in automated testing, progressive rollout, and a fast rollback path has simply increased the rate at which it can break production.
The enabling practices are well established: integrate continuously rather than merging long-lived branches, keep the deployment path automated and identical across environments, decouple deploy from release using feature flags so that shipping code and exposing behaviour are separate decisions, and make rollback cheap enough that it is the boring first response to a bad deploy rather than a last resort. These are covered in more depth under DevOps and SRE.
Operational load is also real capacity. A team carrying an on-call rotation cannot commit its full nominal capacity to planned work, and a plan that assumes otherwise is fiction by midweek. Reserve capacity for interrupt-driven work explicitly and make the reservation visible, so that the trade-off between operational burden and feature delivery is a decision someone makes rather than a debt the team silently absorbs.
Common Failure Modes
| Failure mode | What it looks like in practice | What to do instead |
|---|---|---|
| Ceremony without feedback | The stand-up is status reporting to a manager; nothing is decided and nothing changes as a result | Point every recurring event at a specific decision; delete the ones that do not have one |
| Fixed scope, fixed date, fixed team | Iterative delivery of a commitment that was made as a waterfall plan | Let one of the three flex. If none can, say so in writing before the work starts |
| Velocity as a performance target | Estimates inflate, the chart improves, delivery is unchanged | Use throughput for forecasting only; never compare it across teams |
| Proxy product ownership | The team receives requirements secondhand and every question takes days to answer | Put someone with real decision authority close to the team |
| Unlimited work in progress | Everything is nearly done and nothing has shipped | Set an explicit WIP limit and finish before starting |
| No definition of done | Completed work returns weeks later needing tests, documentation or monitoring | Write the checklist down, including tests, docs, observability and rollback |
| Retrospectives without follow-through | The same complaints recur unchanged for months | Cap output at one or two actions, give each an owner, review them first next time |
When Not to Use It
Iterative delivery is a good default, not a universal one. It fits poorly in several identifiable situations, and recognising them early avoids a great deal of wasted argument.
- Irreversible, high-consequence increments. Where a wrong increment reaching production causes physical harm or unrecoverable loss — safety-critical control systems, certain medical and avionics work, anything operating against a design baseline approved by a regulator — up-front specification and formal verification are not bureaucratic overhead. They are the correct economics for that risk profile.
- Feedback that cannot arrive within the timebox. If the real feedback loop is gated by hardware lead times, a single annual customer event, or an experiment that takes months to read out, shortening the iteration does not shorten the loop. It only adds meetings.
- Genuinely well-understood, repetitive work. A migration with a known procedure does not benefit from being re-planned every two weeks. A checklist and a runbook outperform a sprint board.
- Contractual structures that forbid it. Fixed-price, fixed-scope engagements price the supplier's risk on a defined deliverable. Running an adaptive process inside that envelope without renegotiating the envelope produces conflict, not agility.
Hybrid arrangements are normal and should not be treated as a compromise or a failure of conviction. Running iterative delivery inside a stage-gated programme, or applying it to the parts of a system where feedback is cheap while specifying the parts where it is not, is usually a more honest fit than forcing one model across an organisation with genuinely different risk profiles.
Key Takeaways
- The justification for iterating is the cost of late discovery. Any practice that does not shorten the discovery loop is overhead.
- Fix the timebox and flex the scope. Fixing both reintroduces the problem the timebox was meant to solve.
- Limiting work in progress is the cheapest available improvement to cycle time, and it requires no tooling.
- Forecast from observed throughput as a range; treat estimates as options rather than commitments.
- Watch delivery measures as a set, because individually they are trivially gamed.
- Release frequency and operational investment must rise together, or you have only sped up the failure rate.
- Name the situations where iterative delivery is the wrong tool rather than applying it by default.