ESC
Type to search guides, tutorials, and reference documentation.
Verified by Garnet Grid

Capacity Planning

How to decide how much of a resource to provision: demand modelling, the queueing behaviour that makes high utilisation fragile, Little's Law, headroom sized to the failure you must absorb, autoscaling traps, and the limits nobody remembers until they bind.

Capacity planning is the work of deciding how much of a resource to provision so that demand is served within its objective, at a cost and a risk you have chosen deliberately. It answers two questions that need different tools: does it fit now, which is a measurement problem, and will it fit later, which is a forecasting problem.


What Capacity Planning Is

The output of capacity planning is not a number of servers. It is a defensible chain: an expected demand curve, a measured per-unit capacity, a target amount of headroom with a stated reason, and the lead time required to add more. Any link you cannot show is where the plan will break.

Note the objective in the definition. Capacity is meaningless without a service level to serve it against — "can handle the traffic" is a different statement from "can handle the traffic while keeping latency and error rate inside their objectives". A system usually saturates its latency objective well before it saturates its throughput.


Modelling Demand

Demand has structure, and modelling that structure is most of the work:

  • Organic growth — the slow trend, best fitted over a window long enough to survive a single anomalous month.
  • Seasonality — daily, weekly and annual cycles. A daily peak that is several times the daily mean is common, and it is the peak you must serve.
  • Events — launches, campaigns, migrations, partner integrations. These are known in advance and are not in the trend, so they must be added by hand.
  • Derived demand — load created by your own system: retries, batch jobs, backfills, replication, cache warming, and the client that polls.
  • Demand you do not control — a partner's retry policy, a crawler, a mobile client whose release you do not schedule.

Two habits protect the model. First, plan against peaks, not averages. The peak-to-mean ratio is what makes a peak-provisioned system expensive, and it is the number that tells you whether elasticity is worth its complexity. Second, account for failure-mode demand. During degradation, clients retry, queues drain in bursts, and caches are cold — the load arriving at a struggling system is higher than the load arriving at a healthy one. Capacity planned only for the happy path is capacity that disappears exactly when it is needed.


Knowing Which Resource Binds

The binding constraint is rarely CPU, and teams that plan only on CPU are surprised regularly. Enumerate the candidates explicitly:

DimensionTypical way it bindsSymptom when it saturates
CPUCompute-bound request handlingRising latency with high run-queue depth
MemoryWorking set, caches, per-connection buffersEviction, swapping, or the process being killed
Disk IOPS / throughputRandom reads, write amplification, compactionLatency grows while CPU looks idle
Network bandwidth / PPSLarge payloads, replication, chatty fan-outRetransmits and tail latency, often asymmetric by direction
Connections / file descriptorsPer-connection limits, ephemeral portsConnection refused or exhausted, unrelated to load average
Pool slots (threads, workers, DB connections)Concurrency ceiling before any hardware limitQueueing in front of an apparently idle system
Provider quotasAccount or region limits on instances, addresses, API callsScaling silently stops working at a round number
Licences and third-party rate limitsContractual ceilingsThrottling that no amount of your own hardware fixes

The way to find the real constraint is to load the system until something degrades and observe what saturated first, rather than to reason about it. A saturation test also gives you the more valuable artifact: per-instance capacity. Once you know what one unit can serve while meeting its objective, fleet sizing becomes arithmetic instead of argument.


Utilisation, Queueing and Why Headroom Exists

The durable result behind all capacity work comes from queueing theory: as utilisation of a resource approaches saturation, the time requests spend waiting grows non-linearly. Well below saturation, added load barely changes response time. Close to saturation, a small increase in load produces a large increase in queueing delay, because arrivals are not perfectly smooth and every burst has to be absorbed by a resource with nothing spare.

This is why teams run systems at a target utilisation rather than at full utilisation, and why "we still have spare capacity" and "latency is fine" are different claims. It is also why there is no universal correct target: the right level depends on how bursty arrivals are, how variable service times are, how much latency the objective allows, and how long it takes to add capacity. A system with smooth traffic and instant scaling can run far closer to saturation than one with spiky traffic and a long provisioning lead time.

Little's Law is the other tool worth internalising, because it is an identity rather than a model: the average number of items in a system equals the average arrival rate multiplied by the average time each item spends in the system.

L = λ × W

  L = average items in the system   (concurrency, queue depth)
  λ = average arrival rate          (requests per second)
  W = average time in the system    (seconds per request)

Rearranged for sizing a worker pool:
  workers needed  =  arrival rate  ×  time per request

It is an identity, not an approximation — it holds for any stable
system, which is why it is safe to use on pools, queues and
in-flight request counts alike.

It applies to thread pools, connection pools, in-flight requests and queue depth alike, and it is the fastest way to size a pool from a target throughput and a measured service time — or to spot that a pool is too small to ever reach the throughput you are asking of it.

Finally, scaling is not free. A portion of any workload is serialised — a lock, a leader, a shared index, a single writer — and that portion bounds the speedup you can get by adding workers. Worse, coordination between workers has a cost that grows with their number, so beyond some point adding capacity can reduce throughput rather than increase it. Plan on measuring where that point is rather than assuming linear scaling.


Sizing Headroom to the Failure You Must Absorb

Headroom is not a superstition percentage; it is derived from the largest loss you have committed to surviving. If you run across three failure domains and must keep serving when one is lost, the survivors must absorb the evacuated share on top of their own — which means the steady-state utilisation target follows directly from the redundancy model, not from habit.

Add to that the two variables people forget: the lead time to add capacity (seconds for a warm autoscaling group, weeks for hardware or a quota increase) and the ramp rate of demand. Headroom must cover the demand growth that can occur within your lead time. A system that can double capacity in two minutes needs far less standing headroom than one whose additional capacity requires a purchase order.


Autoscaling and Its Traps

Autoscaling converts a planning problem into a control problem, and control problems have their own failure modes:

  • Scaling on a lagging indicator. If the metric you scale on only moves after users are already queued, the scaler always acts late.
  • Provisioning slower than the ramp. If instances take minutes to become useful and traffic arrives in seconds, autoscaling is a recovery mechanism, not a protection mechanism.
  • Cold start. New instances with empty caches and unwarmed connection pools serve worse than existing ones, so adding capacity briefly makes aggregate latency worse.
  • Thrashing. Aggressive scale-in followed by immediate scale-out, driven by a noisy metric and no cooldown.
  • Scaling into a bottleneck. More stateless instances each opening connections to one database is a way to convert a compute shortage into a database outage.
  • Fighting another controller. An autoscaler adding capacity while a load shedder rejects traffic, or two scalers keyed on different metrics, can oscillate indefinitely.
  • No ceiling. A scaler with no upper bound turns a traffic anomaly or a retry storm into an unbounded bill.

None of these argue against autoscaling; they argue for treating the scaling policy as a system component with its own tests and its own failure analysis, and for keeping a manual override that an incident commander can reach.


Failure Modes

Failure modeMechanism
Planning in averagesPeak demand is what breaks; the mean never describes it
Ignoring the dependency's capacityYour fleet scales; the shared database, queue or third party does not
Forecasting from an uncontrolled metricThe trend is driven by someone else's release or retry policy
Recovery not plannedFleet can serve peak but cannot come back from cold — retry storms and empty caches exceed steady-state demand
Capacity measured on an idle test rigNo noisy neighbours, warm caches, unrepresentative data size
Quota discovered during an incidentProvider or licence ceilings are invisible until scaling silently stops
Headroom consumed silentlyGrowth eats the margin between reviews because nobody owns the number

When Not to Bother

Formal capacity planning costs analyst time, load-testing infrastructure and forecasting discipline. For a small service with elastic infrastructure, a cheap resource footprint and a forgiving objective, over-provisioning is simply cheaper than the analysis that would let you provision precisely — and the honest answer is to buy the headroom and spend the attention elsewhere.

The calculus flips when any of these are true: the resource is expensive or has a long lead time; the workload is stateful, so adding capacity means rebalancing data; the demand is spiky enough that a miss is user-visible; or a quota, a licence or a physical limit means you cannot buy your way out at short notice. Those are the systems where a plan is worth writing down.

Capacity work leans on two neighbours. The measurements that make it real come from observability, and the per-unit limits and saturation points come from the same load and profiling work described in performance tuning — an efficiency improvement and a capacity increase are the same thing seen from two directions.

Jakub Dimitri Rezayev
Jakub Dimitri Rezayev
Founder & Chief Architect • Garnet Grid Consulting

Jakub holds an M.S. in Customer Intelligence & Analytics and a B.S. in Finance & Computer Science from Pace University. With deep expertise spanning D365 F&O, Azure, Power BI, and AI/ML systems, he architects enterprise solutions that bridge legacy systems and modern technology — and has led multi-million dollar ERP implementations for Fortune 500 supply chains.

View Full Profile →