ESC
Type to search guides, tutorials, and reference documentation.

Platform Engineering and Internal Developer Platforms

Treating internal infrastructure as a product: paved roads, self-service provisioning, guardrails instead of approval gates, the interfaces a platform exposes, and the reasons platform teams fail to get adopted.

Platform engineering is the practice of building and operating the internal capabilities that application teams use to ship software — and treating those capabilities as a product, with users, a roadmap, documentation, support, and adoption as the measure of success. The assembled result is usually called an internal developer platform: the paved road from an idea to a running, observable, secure service.


What Platform Engineering Is

The distinguishing word is product. A platform team's customers are internal engineers, its product is the set of workflows those engineers use, and its success measure is voluntary adoption. That framing has teeth: it means user research before building, it means a supported interface with a compatibility promise, and it means that a capability nobody uses is a failure regardless of how elegant it is.

The Problem It Responds To

The pattern that created the need:

  "You build it, you run it" moved operational responsibility to product teams.
  The cloud-native toolchain then grew: containers, orchestration, IaC,
  service meshes, observability stacks, policy engines, CI systems, registries.

  Each product team now needs working knowledge of all of it — or reinvents
  a worse version of it — just to deploy a web service.

  Cognitive load, not capability, becomes the constraint.

The platform response:
  Extract the parts that are the same for everyone into a supported product,
  so a product team makes application decisions and not infrastructure ones.

The Team Topologies vocabulary is useful here: a platform team exists to reduce the cognitive load of stream-aligned teams by providing a self-service capability, delivered as a service rather than through collaboration on every request. The measure of a healthy relationship is that product teams consume the platform without needing to talk to the platform team — conversation should be the exception, not the delivery mechanism.


What an Internal Developer Platform Provides

CapabilityWhat it removes from the product team
Service scaffoldingDeciding project structure, CI wiring, base image, and telemetry from scratch
Delivery pipeline templatesBuilding a bespoke CI/CD pipeline per repository
Environment provisioningFiling tickets and waiting for a namespace, database, or queue
Runtime abstractionHand-writing deployment manifests for every service
Secrets and identityInventing a credential distribution scheme per team
Observability defaultsWiring logs, metrics, traces, and dashboards individually
Service catalogGuessing who owns a service and how to reach them at 3am
Policy and complianceRemembering the security rules that a scan will fail on later

The capability that ties the rest together is the golden path: a documented, supported, opinionated route through those capabilities for the common case. A golden path is not a mandate. It is the option that is so obviously easier that choosing something else becomes a deliberate act with owned consequences.


The Interface Is the Product

Platforms are consumed through some combination of four interfaces, and the choice shapes adoption more than the underlying implementation does:

  • A manifest or spec file in the application repository. The developer declares what the service needs; the platform reconciles it. Fits existing review workflows and version control.
  • A command line tool. Fits the inner loop, scriptable, easy to adopt incrementally. The most common first interface, and often the most used one long after a portal exists.
  • A portal. Good for discovery, catalogs, ownership, and documentation. Weak as the only way to perform an action, because a click is not reviewable and not repeatable.
  • An API. The substrate the other three should be built on. If the portal can do something the API cannot, the platform has a manual step wearing a user interface.

Build the API first and treat every other surface as a client of it. That ordering is what makes the platform automatable by its users, and it is what allows the portal to be replaced later without re-implementing the platform.


Guardrails, Not Gates

Every platform has to enforce constraints: encryption, network boundaries, cost controls, data residency. There are two ways to do it and they produce very different organisations.

Gate        A human approval standing between a developer and an action.
            Scales linearly with request volume; becomes a queue;
            the approver usually lacks context to refuse, so it becomes
            a rubber stamp that costs time and provides no safety.

Guardrail   A policy evaluated automatically at the moment of change,
            with a clear message about what failed and how to fix it.
            Scales with automation rather than headcount; denies the
            unsafe action rather than delaying every action.

Guardrails should fail closed on the small set of rules that genuinely matter and stay silent otherwise. A platform that blocks a hundred things produces workarounds; a platform that blocks five things, explains each clearly, and offers a documented exception path produces compliance. Keep the escape hatch real: teams with genuinely different requirements must be able to step off the paved road, and should then own what they stepped off into.


Build, Buy, or Assemble

Very few organisations build a platform from nothing. The realistic decision is which layers to assemble from existing components and where to write the thin, opinionated glue that encodes your decisions. Buy or adopt the commodity layers — source control, CI runners, registries, orchestration, observability backends, secret stores. Build the parts that encode your organisation's specific choices: the scaffolding templates, the service manifest schema, the promotion workflow, the catalog's ownership model.

The failure on the build side is a platform team writing an orchestrator. The failure on the buy side is adopting a product whose model does not match how your teams work and then spending years bending teams to fit it. A useful test: if a component would be roughly the same at any company, do not write it; if it encodes a decision that is specifically yours, do not outsource it. The longer form of this analysis is in platform build versus buy.


Why Platform Teams Fail

Failure modeWhat happensCorrection
Mandate without easeAdoption is ordered rather than earned; teams build workaroundsWin on convenience; measure voluntary adoption
Built without usersMonths of work solving a problem nobody hadInterview teams; ship a thin path end to end early
Leaky abstraction, no escapeThe abstraction breaks and users cannot get underneath itExpose the layer below; document the escape hatch
Ticket queue in disguiseSelf-service in name; a human fulfils every requestAutomate the top request types; count tickets as defects
Unversioned breaking changesA platform change breaks consumers without warningVersion the interface; deprecate with notice and migration help
No support modelUsers get stuck, get no answer, and leaveA staffed channel with a response expectation
Scope sprawlThe platform owns everything and can maintain none of itAn explicit charter naming what is out of scope
Documentation as an afterthoughtThe golden path exists but is undiscoverableTreat docs as part of the deliverable, not a follow-up

Measuring a Platform Honestly

Activity metrics — tickets closed, features shipped, services onboarded by decree — tell you what the platform team did, not whether it helped. The measures worth tracking are about the users: what fraction of services use the golden path without being told to; how long it takes a new engineer to get a change to production for the first time; how much support load each capability generates; and what developers say in a periodic survey. Pair those with the delivery measures of the teams you serve, so you can tell whether platform investment is actually moving lead time and change failure rate.

Track each of these against your own baseline over time. External comparisons are close to meaningless here, because they depend entirely on organisation size, domain, and what the platform was scoped to do.


When You Do Not Need One

A platform is an investment that pays back through repetition. With one product, a handful of services, and a team small enough that everyone already knows how deployment works, a dedicated platform team is overhead: the coordination cost of an internal supplier exceeds the cognitive load it removes. The right move at that size is a paved path rather than a platform — a good template repository, a shared pipeline definition, infrastructure as code for the environments, and written documentation, all maintained by the same people who use them.

The signal that it is time to invest is repetition with divergence: several teams solving the same infrastructure problem differently, onboarding taking weeks because every service is bespoke, or security and cost controls that exist on paper and nowhere in the actual path to production. At that point a platform stops being overhead and starts being the cheapest way to make the right thing the default.