ESC
Type to search guides, tutorials, and reference documentation.
← Back to all categories
🖥️

Backend Development

What backend engineering is really about: request lifecycle, where state lives, concurrency and connection pools, transactions, background work, caching, failure handling, and the operational habits that keep a service correct under load.

Backend development is the discipline of keeping shared mutable state correct while many callers touch it concurrently and parts of the system are failing. Frameworks, languages, and deployment targets change constantly; that problem statement does not. Almost every hard backend bug is a variation of two people writing at once, or of something that should have had a deadline not having one.

What backend work actually is

Strip away the framework and a backend service does four things: it accepts a request, authorises it, reads and writes durable state, and returns a result or an error. The interesting engineering lives in the guarantees around those steps — whether a half-completed operation can leave the system inconsistent, whether a repeat of the same request causes a second effect, and whether one slow dependency can take the whole service down.

A useful way to evaluate any backend design is to ask three questions of every write path: What happens if this runs twice? What happens if it stops halfway? What happens if the thing it depends on is slow rather than down? Slow is the one teams forget, and it is the most damaging of the three.

The request lifecycle

Keeping the layers distinct is what makes a service testable and changeable:

  • Transport / edge — TLS termination, routing, request size limits, coarse rate limiting.
  • Boundary validation — parse the input into a trusted internal type once, at the edge. Everything past that point should be working with values that are already known-valid, not re-checking strings.
  • Authorization — checked against the object being acted on, not only against the route.
  • Domain logic — ideally pure and free of I/O, which makes it the easiest part to test exhaustively.
  • Persistence and integrations — the only layer that knows about the database or an external vendor.

The single highest-leverage rule here is parse, do not validate repeatedly. Converting untrusted input into a typed value at the boundary eliminates an entire family of bugs where one code path checked a constraint and another did not.

Stateless services and where state really lives

"Stateless service" does not mean there is no state; it means no state that only exists on one instance. That property is what lets you add capacity, restart on deploy, and lose a node without losing correctness.

The common violations are easy to miss: an in-process cache that instances disagree about, a scheduled job that assumes it is the only instance running it, an in-memory rate limit counter, a WebSocket session that only one node knows about, or a file written to local disk. Each of these is fine until there is a second instance — which is exactly when you are under load and least able to debug it.

The fix is usually to move the state somewhere all instances share and to make the coordination explicit: a distributed lock or a leader election for singleton jobs, a shared store for counters and sessions, object storage instead of local disk.

Concurrency models and connection pools

Two dominant models: thread-per-request (blocking calls are fine; concurrency is bounded by threads, each with real memory cost) and event-loop / async (one or a few threads multiplex many in-flight operations; concurrency is cheap, but a single blocking call starves everything sharing that loop). Neither is faster in the abstract. The async model wins when the service is dominated by waiting on I/O; the threaded model is simpler to reason about and harder to catastrophically misuse.

The classic async failure is a CPU-bound operation or a synchronous library call on the loop: throughput collapses across every unrelated request, and the metrics show latency rising everywhere at once with no obvious culprit.

Connection pools are where concurrency meets the database, and they are a queue you are probably not measuring. A pool that is too small turns into a hidden serialisation point — requests wait to get a connection before they even start waiting on the query. A pool that is too large moves the contention into the database, where it is worse, because each connection consumes server-side memory and adds scheduling overhead. The number that matters is not the pool size in isolation but pool size multiplied by instance count, which is what the database actually sees; autoscaling can quietly multiply this past the database's limit. Always measure time-spent-waiting-for-a-connection separately from query time.

Data access and transactions

Wrap the writes that must succeed or fail together in one transaction, and keep that transaction short and free of network calls to anything else. A transaction held open across an HTTP call to a third party holds locks for the duration of someone else's outage.

Know which isolation level you are actually running under, because the default is frequently weaker than people assume. Under a weak level, the read-then-write pattern — read a balance, compute a new one, write it back — silently loses updates under concurrency. The fixes are to make the write conditional on the value you read (optimistic concurrency, typically via a version column), to let the database compute the new value atomically, or to take an explicit lock.

Schema migrations deserve the same expand-and-contract discipline as APIs: add the new column, write to both, backfill, move reads, then drop the old one. The reason is simply that during a rolling deploy, old and new code run at the same time against the same schema, so any migration that is only compatible with one of them breaks requests for the length of the deploy.

Background work and the dual-write problem

Anything slow, retryable, or not needed for the response should be moved off the request path. That introduces a queue, and queues bring their own contract.

Most messaging systems give you at-least-once delivery. That is not a defect to engineer around; it is the honest guarantee, and the consequence is that every consumer must be idempotent. Combine that with a limited number of retries and a dead-letter destination so poison messages stop blocking the queue instead of being reprocessed forever.

The deepest trap in this area is the dual write: committing to the database and then publishing an event as two separate operations. If the process dies between them, the state and the event disagree permanently, and no amount of retrying fixes it because the retry has no record that it is needed. The durable fix is the transactional outbox — write the event into the same database, in the same transaction as the state change, and have a separate relay publish committed outbox rows. Now there is exactly one commit point, and the relay can safely retry because the consumers are idempotent.

Caching

A cache trades freshness for latency and load. Deciding what staleness is acceptable is a product decision, not a technical one, and it should be made explicitly per dataset.

The mechanics that repeatedly bite: invalidation is harder than population, so prefer short expiries plus explicit invalidation on write rather than relying on either alone; a stampede occurs when a popular key expires and every concurrent request recomputes it simultaneously, which you prevent by having one caller recompute while others serve the stale value or wait; and a cache that becomes load-bearing is a new single point of failure, so know whether your service survives the cache being empty or unreachable. If it does not, you have not added a cache — you have added a dependency.

Failure handling: timeouts, retries, backpressure

Timeouts are the foundation. Every outbound call gets one; a call without a deadline can occupy a worker indefinitely. Deadlines should be propagated, so work whose caller has already given up can be abandoned rather than completed for nobody.

Retries need exponential backoff, jitter, a bounded attempt count, and a check that the operation is idempotent. Retries at several layers multiply — three layers each retrying three times is up to twenty-seven calls for one logical request — so decide deliberately which layer owns retrying.

Circuit breakers stop hammering a dependency that is clearly failing and let it recover, while giving your service a fast, predictable failure instead of a slow, resource-consuming one. Bulkheads — separate pools per dependency — stop one sick downstream from consuming all of your workers.

Backpressure is the one most often missing. When work arrives faster than it can be completed, an unbounded internal queue converts a throughput problem into a memory problem and then into a crash, after which the restarted instance faces the same backlog. Bounded queues and explicit load shedding — rejecting excess work quickly and visibly — keep the service in a state it can recover from. Systems that lack this exhibit metastable failure: they stay down after the triggering cause is gone, because the recovery work itself is now the overload.

Observability

You need to answer "what is happening right now" and "what happened to this one request" without redeploying. In practice that means structured logs (fields, not sentences), metrics for rates, errors, saturation and latency distributions rather than averages, and traces carrying a correlation identifier across service boundaries so one request can be reconstructed end to end.

Averages hide the failures users actually experience; a stable mean latency is compatible with a meaningful share of requests timing out. Look at tail percentiles, and prefer a histogram you can re-slice over a pre-computed average you cannot.

Health checks deserve a specific warning: a liveness check that verifies the database is reachable will fail every instance simultaneously during a database blip, and the orchestrator will restart a fleet that was not broken. Keep liveness ("is this process wedged?") separate from readiness ("should this instance receive traffic right now?").

Common failure modes

  • The missing timeout. One slow dependency exhausts the worker pool and the service stops serving requests that do not even touch it.
  • The N+1 query. Fine with test data, quadratic in production. Visible instantly in a trace, invisible in code review.
  • The unbounded query. A query with no limit that works until one customer's data grows.
  • The dual write. Database and message broker disagree, permanently.
  • The non-idempotent consumer. At-least-once delivery meets a handler that charges a card.
  • Retry storms. Every layer retrying independently, multiplying load exactly when the system is weakest.
  • The hidden singleton. A cron or cache assumption that silently breaks the moment you run two instances.

When to build a service, and when not to

Split a service out when a component has a genuinely different scaling profile, a different availability requirement, a different data ownership boundary, or a different release cadence that is being held back by the rest of the system.

Do not split along team lines alone, and do not split components that must change together. Two services that are always deployed in lockstep are a distributed monolith: you have taken on network failure, partial failure, serialisation cost, and version skew, and received none of the independence that was supposed to pay for them. A module boundary inside one deployable gives most of the design benefit at a fraction of the operational cost, and it is far easier to move a module out later than to merge two services back together.

01

API Development

How to design, evolve, secure, and operate an HTTP API: protocol styles, contract design, error and pagination models, idempotency, versioning, rate limiting, and the failure modes that break clients.

→
02

System Design

A practitioner's reference for designing distributed systems: requirements and constraints, scaling and partitioning, consistency models, synchronous versus event-driven communication, failure domains, backpressure, and knowing when not to distribute.

→
03

Mobile Engineering

What makes mobile different from web and server work: irreversible releases, old versions in the wild, offline and sync, constrained resources, permissions, and the release engineering that keeps a shipped app recoverable.

→
04

Backend Engineering

Rate limiting, caching strategies, health checks, graceful degradation, and server design.

→
05

Caching Strategies: Serving Data at the Speed of Memory

Implement caching that reduces latency from hundreds of milliseconds to single-digit milliseconds. Covers cache placement, invalidation strategies, cache-aside vs write-through patterns, distributed caching with Redis, CDN caching, and the pitfalls that turn your cache into a source of stale data and subtle bugs.

→
06

Distributed Locking

Production engineering guide for distributed locking covering patterns, implementation strategies, and operational best practices.

→
07

CDC Change Data Capture

Production engineering guide for cdc change data capture covering patterns, implementation strategies, and operational best practices.

→