ESC
Type to search guides, tutorials, and reference documentation.
← Back to all categories
🔌

API Development

How to design, evolve, secure, and operate an HTTP API: protocol styles, contract design, error and pagination models, idempotency, versioning, rate limiting, and the failure modes that break clients.

An API is a contract between two pieces of software that are deployed, versioned, and operated independently. Everything that makes API work hard follows from that one property: once another team, another company, or a shipped mobile app depends on your response shape, you no longer control when they upgrade. The code is the easy part. The contract is the product.

What an API actually is

A useful mental model is that an API has four surfaces, and you have to design all four deliberately:

  • The data model you expose. Not your database schema — a deliberately chosen projection of it.
  • The operations you allow and the preconditions under which they succeed.
  • The failure semantics — what a caller should do when something goes wrong, expressed in a form code can branch on.
  • The operational envelope — rate limits, timeouts, payload sizes, pagination limits, and availability expectations.

Teams routinely design the first two and leave the last two implicit. The implicit ones are where production incidents come from, because callers discover them by hitting them.

Choosing a protocol style

There is no universally correct style. There are workload shapes, and each style fits some of them.

StyleFits whenCosts you
REST over HTTPResource-shaped domains, many unknown clients, caching and intermediaries matterChatty for composite reads; over-fetching; relationships get awkward
RPC (including gRPC)Service-to-service calls inside one organisation, action-shaped operations, strict schemasHarder through browsers and generic proxies; tighter coupling to generated stubs
GraphQLMany differently-shaped clients reading one graph; UI teams need to iterate without backend round tripsQuery cost control, caching, and authorisation all become your problem
Events and webhooksThe consumer needs to know when something changed, not to ask repeatedlyDelivery guarantees, ordering, replay, and consumer-side idempotency

A practical heuristic: pick the style that makes the common call cheap and the rare call possible. If your clients' dominant access pattern requires three round trips in your chosen style, the style is wrong for that workload — or you need a purpose-built endpoint for it.

Designing the contract

Model resources, not screens. Endpoints shaped around the current UI calcify the moment the UI changes. Endpoints shaped around domain nouns survive redesigns.

Never expose an unbounded collection. Every list endpoint needs pagination from the first version, with a maximum page size the server enforces. Cursor-based pagination (an opaque token encoding the sort position) is more robust than offset pagination under concurrent writes, because offsets skip and duplicate rows when the underlying set shifts between pages.

Be explicit about time and identity. Timestamps go over the wire in an unambiguous absolute format; identifiers are opaque strings to the client even if they are integers to you. Both decisions are cheap on day one and very expensive to change later.

Separate the transport status from the domain outcome. HTTP status codes describe what happened to the request; the body describes what happened to the business operation. Returning an HTTP success for a failed domain operation forces every client to parse the body to know whether to retry — and some of them will not.

Errors a client can act on

An error response has one job: let the caller decide what to do next without reading English. That requires, at minimum, a stable machine-readable code that you treat as part of the contract, a human-readable message intended for logs rather than end users, and enough structure to identify which field or entity failed for validation errors.

The distinction that matters most is retryable versus terminal. A caller that cannot tell the difference will either retry things that will never succeed — turning a small failure into a sustained load problem — or give up on transient failures that a single retry would have cleared. If your error model only communicates this through the status code, document precisely which codes are retryable.

Idempotency and safe retries

Networks fail in a specific, annoying way: the caller cannot distinguish "the request never arrived" from "the request succeeded and the response was lost". Both look like a timeout. This means any operation that can be retried must be safe to execute twice, or you will double-charge, double-ship, and double-post.

Reads and deletes are naturally idempotent. Creates are not. The standard mechanism is an idempotency key: the client generates a unique key per logical operation and sends it with the request; the server records the key alongside the result of the first successful execution and, on seeing the key again, returns the stored result instead of re-executing. The subtleties are all in the details — the key must be recorded in the same transaction as the effect, it needs a retention window, and a replayed key with a different payload is a client bug you should reject loudly rather than silently honour.

On the client side, retries need exponential backoff with jitter and a bounded attempt count. Without jitter, a fleet of clients that failed together retries together, and the synchronised wave keeps the dependency down.

Evolving an API without breaking clients

Changes fall into two classes. Additive changes — new optional fields, new endpoints, new enum-tolerant values — are safe if and only if your clients ignore unknown fields. Say so explicitly in your documentation, because a strict-parsing client turns every additive change into a breaking one.

Breaking changes — removing a field, tightening validation, changing a type or a default, changing the meaning of an existing value — require a migration path. The durable pattern is expand and contract: add the new shape alongside the old, move clients across, verify nobody is using the old shape, then remove it. The verification step is the one teams skip, and it is the only one that makes the removal safe.

Whatever versioning scheme you choose — a path segment, a header, a media type — the important decisions are not syntactic. They are: how long a version is supported, how a deprecation is announced, how you measure who is still on it, and whether you can turn an old version off. An API with unlimited versions and no usage telemetry has no deprecation policy, only intentions.

Authentication and authorization

Keep the two separate in your head and in your code. Authentication establishes who is calling. Authorization decides whether that caller may perform this operation on this object. Authentication is a boundary concern you can centralise; authorization usually is not, because it depends on domain state the gateway cannot see.

Design notes that hold up over time: scope tokens to the narrowest capability that works; make machine credentials and user credentials structurally distinct so you can revoke and audit them separately; check authorization on the object being touched rather than only on the route, or you get the classic flaw where a valid token for one account can read another account's record by changing an identifier in the path.

Rate limits, quotas, and timeouts

A public API with no limit is a denial-of-service vector operated by your own customers, usually by accident. Limits should be visible (the response says what the limit is and when it resets), attributed to a meaningful principal rather than an IP address alone, and graduated so an expensive endpoint can be limited more tightly than a cheap one. Returning the retry timing tells well-behaved clients how to back off instead of guessing.

Timeouts are the other half. Every outbound call your API makes needs a deadline shorter than the deadline your caller gave you, or a slow dependency converts into exhausted connections and a queue of requests nobody is waiting for any more. Propagating the caller's remaining deadline downstream is what makes this composable.

Documentation and contract testing

A machine-readable schema generated from, or enforced against, the running service is worth more than prose, because prose drifts silently. The failure mode of hand-written documentation is not that it is wrong on day one — it is that nothing breaks when it becomes wrong.

Contract tests close that gap from the other side: the consumer declares the shape it depends on, and the provider's build fails if it stops satisfying that shape. This catches the specific and common case where a provider makes a change that is technically additive but violates an assumption a consumer actually relies on.

Common failure modes

  • The leaked internal model. Serialising database rows directly means every schema change is an API change, and every internal field is now a public commitment.
  • The chatty collection. A list endpoint that returns identifiers, forcing N follow-up calls. Either embed what callers always need or provide a bulk read.
  • Unbounded anything. No page cap, no payload cap, no query depth cap. The first heavy client finds it.
  • Retry storms. Aggressive client retries plus no server-side limiting turns a brief degradation into a sustained outage that persists after the original cause is gone.
  • Silent breaking changes. Tightening validation is a breaking change even though nothing was removed.
  • Errors as prose. Free-text messages with no stable code mean clients end up matching on message text, which you then cannot change.

When to build an API, and when not to

Build one when independent parties need programmatic access on their own release schedule, when you are deliberately decoupling two systems that must evolve separately, or when the integration surface is a product in its own right.

Do not build one when the "two sides" are the same team deploying the same binary — an internal function call has better failure semantics, better refactorability, and no versioning burden. Do not build a general-purpose API when you have exactly one known consumer with one known use case; a purpose-built endpoint is cheaper to build and vastly cheaper to change. Generality is a cost you should pay once you have evidence you need it.