Performance Tuning
Measuring and improving latency and throughput honestly: why the mean lies, why percentiles do not average, coordinated omission, the USE and RED framings, profiling, the standard bottleneck classes, and the caching and retry behaviour that turns a small problem into an outage.
Performance tuning is the disciplined process of making a system meet a stated latency and throughput objective under realistic load. The discipline is mostly in the measurement: almost every wasted optimisation effort traces back to a number that was measured in a way that could not have revealed the real problem.
What Performance Tuning Is
Three quantities describe a system's performance, and conflating them causes most confusion. Latency is how long one operation takes. Throughput is how many operations complete per unit of time. Utilisation is how much of a resource is consumed. They are related but not interchangeable — a system can have excellent throughput and unacceptable latency, and a system at low utilisation can still be slow if the delay is spent waiting rather than working.
Performance work should always be anchored to an objective. Without one, there is no definition of "fast enough", so the work has no termination condition and no way to justify its cost against other work.
Measuring Latency Honestly
The mean latency of a request stream is close to useless on its own, because latency distributions are skewed: a long tail of slow requests barely moves the mean while dominating what users experience. Report distributions — high percentiles alongside the median — and look at the shape, because a bimodal distribution (a fast path and a slow path) tells you something a single number cannot.
Two properties of percentiles are worth stating precisely, because violating them is common and silently wrong:
- Percentiles do not average. The mean of the p99 values reported by ten servers is not the p99 of the combined traffic. Combining requires the underlying distributions (histograms or sketches), not the summary numbers.
- Percentiles do not add across stages. A request passing through three stages that each have a p99 does not have a p99 equal to their sum, because the slow cases in each stage are not the same requests.
- A percentile over a long window hides short, severe events. A five-minute total outage can leave a daily percentile looking acceptable. Match the window to the duration of the events you care about.
Tail latency dominates fan-out. If serving one user request requires many internal calls, and any one of them being slow makes the whole request slow, then the probability that a request avoids every slow call shrinks as the fan-out grows. The mechanism is multiplicative: the more backends a request must wait on, the more likely it is that at least one of them is in its own slow tail. This is why a service whose individual dependencies all look healthy can still be slow, and why reducing tail latency in a shared dependency helps far more than reducing its median.
Coordinated omission is the measurement error to know by name. A load generator that sends a request, waits for the response, and only then sends the next one cannot record the requests it failed to send while the system was stalled. During a pause, such a client simply issues fewer requests — so the stall is under-represented, and the resulting percentiles look far better than reality. The fix is to drive load at a fixed schedule and measure each request's latency from its intended start time, counting the time it spent waiting to be sent.
A Method That Terminates
Performance work without a method degenerates into changing things that seem slow. Two framings keep it bounded, and they are complementary:
| Framing | Applied to | Ask for each |
|---|---|---|
| USE | Every resource (CPU, memory, disks, network, pools) | Utilisation, saturation (queued work), and errors |
| RED | Every service or endpoint | Rate, errors, and duration distribution |
The loop itself is simple and should be followed strictly: establish the objective → measure the current distribution → find the largest contributor → change one thing → re-measure against the same workload → stop when the objective is met. Changing several things at once means you cannot attribute the improvement, and you will keep a change that did nothing while discarding the one that worked.
Profiling is how you find the contributor inside a process. A sampling profiler periodically captures stacks and aggregates them, typically rendered as a flame graph where width is time. The important distinction is on-CPU versus off-CPU time: on-CPU profiling shows which code burned cycles, off-CPU analysis shows where the process was blocked — on a lock, on I/O, on a network response, on garbage collection. Teams that profile only on-CPU time frequently conclude their code is efficient while the request spends most of its life waiting.
The Standard Bottleneck Classes
Most latency problems in server software fall into a small number of recurring shapes. Recognising them is faster than rediscovering them.
| Class | Signature | Usual fix |
|---|---|---|
| N+1 queries | Latency scales with result-set size; many tiny identical queries | Batch, join, or prefetch the related set |
| Missing or unusable index | One query dominates; plan shows a full scan | Index the predicate; or change the predicate so an index applies |
| Lock contention | Throughput flattens or falls as concurrency rises | Shorten the critical section; shard the lock; use optimistic concurrency |
| Pool exhaustion | Queueing in front of an idle-looking backend | Size the pool from arrival rate × service time; fix the slow holder |
| Chatty remote calls | Latency dominated by round trips, not work | Batch, colocate, or move the computation to the data |
| Head-of-line blocking | One slow item stalls everything behind it on a shared path | Separate queues by cost class; bound per-item work |
| Pause events (GC, compaction, checkpoint) | Periodic latency spikes uncorrelated with load | Tune the pause source; reduce allocation or write amplification |
| Serialisation and payload size | CPU and network both rise with response size | Return less; compress; use a cheaper encoding |
| Cold start | First requests after deploy or scale-out are slow | Warm caches and connections before taking traffic |
Queues, Timeouts and Retries
An unbounded queue converts a throughput deficit into a latency problem and hides it. If arrivals exceed service rate, the queue grows, and every request now waits behind the backlog — so instead of visibly rejecting some requests, the system accepts all of them and serves all of them too late. A bounded queue with explicit rejection is nearly always the better failure mode: it preserves the objective for the requests you do accept, and it makes the deficit visible immediately.
Timeouts must form a budget. If a caller's timeout is longer than the sum of the timeouts below it, the caller holds resources long after the work is doomed. A useful discipline is to pass a remaining-time budget down the call chain and have each layer refuse work it cannot finish within it, so the system stops spending on requests whose deadline has already passed.
Retries amplify load exactly when the system is weakest. Three properties keep them safe: a cap on attempts, exponential backoff with jitter so retries from many clients do not synchronise into a wave, and a retry budget that stops retrying globally once the retry rate exceeds a fraction of the base rate. Without them, a brief degradation becomes a self-sustaining overload — the system is slow, clients retry, the extra load makes it slower. Circuit breakers address the same dynamic from the caller's side by failing fast once a dependency is evidently unhealthy.
Caching and Its Debts
A cache is a bet that recomputation is more expensive than staleness. The bet is usually right, and the debts are always the same:
- Staleness. You have chosen to serve data that may be wrong; the question is only for how long, and whether any caller cannot tolerate it.
- Invalidation. Correct invalidation is the hard part, and the failure is silent — a stale entry looks exactly like a fresh one.
- Stampede on expiry. When a popular entry expires, many concurrent requests miss simultaneously and all recompute it. Single-flight coordination, early probabilistic refresh, and staggered expiry times each address this.
- The cold-start cliff. After a restart or a mass eviction, the origin briefly sees the traffic the cache was absorbing.
- A second consistency surface. Multiple cache layers with different lifetimes can disagree, producing behaviour that depends on which layer a request happened to hit.
The subtlest debt is that a cache hides a capacity deficit. A system whose origin cannot serve its full demand looks healthy for as long as the hit rate holds. The deficit is revealed only at the worst moment — a restart, an eviction storm, a key-space change, or a deploy that changes the cache key. Knowing whether your origin can survive a cold cache is a capacity question, and it is best answered deliberately rather than during an incident.
Benchmarking Traps
- Benchmarking an idle machine. No neighbours, no background jobs, no competing tenants — none of which resemble production.
- Warm everything. Caches full, connections open, JIT already optimised: valid for steady state, invalid for anything involving a restart or a scale-out.
- Unrepresentative data. A dataset small enough to fit in memory measures a different system from the one you run.
- Uniform synthetic load. Real traffic is skewed — hot keys, hot tenants, hot routes — and skew is what finds contention.
- Measuring the client. A load generator that saturates its own CPU, connection pool or event loop reports its own limits as the server's.
- One run. Without repetition and variance you cannot distinguish an improvement from noise.
- Ignoring the error rate. A configuration that is fast because it is rejecting or truncating work is not faster.
When to Stop
Stop when the objective is met. Performance work has strongly diminishing returns, and past the objective it competes with reliability, feature and cost work that may matter more. Two corollaries follow. First, do not optimise code that is not on the critical path — a profile, not intuition, decides what is. Second, do not accept an obviously wasteful design because optimisation is "premature"; the guidance against premature optimisation was never an argument for choosing an algorithm or a data access pattern you already know will not scale.
Resist optimisations that trade away clarity for an improvement that the objective did not require, and re-measure after every dependency upgrade and traffic change: a performance property that is not continuously verified is a property you used to have. The signals that let you verify it continuously are the subject of observability, and the behaviour of these mechanisms under injected fault is what chaos engineering is for.