Zero Trust Architecture
Zero trust explained as an engineering change rather than a product: per-request authorization, policy decision and enforcement points, device posture, workload identity and mTLS, a staged migration off the VPN, and what zero trust does not fix.
Zero trust is an architecture in which being on a particular network grants nothing. Every request carries an identity, is authorized against policy at the moment it is made, and is logged. It replaces the perimeter assumption — inside is trusted, outside is not — with a per-request decision taken as close to the resource as possible.
It is not a product, and no vendor can sell it to you as one. It is a change to where authorization happens, which means it is mostly an identity project, a policy project, and a migration project. Teams that buy a gateway and declare the work done usually end up with a perimeter that costs more.
What Zero Trust Is
The perimeter model made one bet: authenticate hard at the boundary, then let traffic inside move freely. It failed on contact with reality — laptops leave the office, software-as-a-service lives outside the boundary, and an attacker with a foothold inherits the same free movement as a legitimate employee.
Zero trust makes a different bet. Trust is never implied by location; it is computed per request from identity, device state, and context, and it expires. The vocabulary most commonly used for the components — policy engine, policy administrator, policy enforcement point — comes from the NIST zero trust architecture publication (SP 800-207), and it is worth adopting because it separates the three jobs cleanly.
The Control Loop
signals
┌──────────────────────────────┐
│ identity + factor strength │
│ device posture │
│ resource sensitivity │
│ request context (time, geo) │
│ live risk / anomaly signals │
└──────────────┬───────────────┘
▼
POLICY ENGINE ── decides allow / deny / step-up
│
POLICY ADMIN ── issues, refreshes, REVOKES the session
│
request ─────▶ ENFORCEMENT POINT ─────▶ resource
│ (app, API, database,
▼ host, workload)
decision log
(who, what, when, why, allowed?)
The decision is RE-EVALUATED, not granted once.
Network position contributes NOTHING to it.
Three design consequences follow immediately. The enforcement point must sit in the data path, which makes it a latency and availability dependency. The policy engine must be reachable and fast, which usually means caching decisions for a bounded interval and accepting a small revocation delay. And the decision log is not optional telemetry — it is the only way to answer who reached what, which is the question asked during every incident and every audit.
Decide deliberately what happens when the policy plane is unreachable. Failing closed protects the resource and risks an outage; failing open protects availability and quietly disables the control at exactly the moment an attacker would most like it disabled. Whichever you choose, choose it explicitly, write it down, and test it — this behaviour is almost never exercised before the day it matters.
User Access
For human access, the change is from “the VPN grants a network route” to “the proxy grants one application, per request”. An identity-aware proxy terminates the connection, authenticates the user against the identity provider, evaluates policy, and only then forwards to the application. The user never gets a route to anything else, so a compromised laptop cannot scan a subnet it was never given.
Device posture becomes a first-class policy input: is this device managed, encrypted, patched, running the controls the organisation requires. This is what makes the difference between authenticating a person and authorizing a session, and it is the input most likely to cause friction — a posture rule that blocks legitimate work will be exempted, and the exemption will be permanent. Step-up authentication is the pressure valve: allow ordinary work on ordinary evidence, and require a stronger factor for the small set of sensitive operations.
Workload to Workload
Most traffic in a modern estate is service to service, and most segmentation projects stop at the north-south edge. The east-west requirement is the same three ingredients: each workload has a verifiable identity, calls are mutually authenticated, and authorization is expressed per caller and per operation rather than per IP address.
Identity for workloads is issued by the platform — short-lived certificates or tokens tied to what the workload is, not to a secret someone pasted into a configuration file. Mutual TLS then gives both authentication and transport protection in one mechanism. A service mesh is one way to deliver this; libraries and sidecars are others. The requirement is identity, policy, and telemetry — not any particular implementation.
Segmenting by identity rather than by address is the point. Addresses are recycled, ephemeral, and forged; a rule written against a workload identity survives a redeploy, and a rule written against a subnet decays into a permanent allow-any that nobody dares remove.
A Staged Migration
1. inventory who uses which application, from where, on what
2. identity first one directory · MFA · joiner-mover-leaver that works
3. one application put it behind an identity-aware proxy
4. log-only run the policy without enforcing; read the would-be denies
5. enforce only once the deny set is understood and owned
6. repeat by sensitivity, not by ease
7. east-west workload identity + mTLS for the crown jewels
8. retire the VPN LAST — once nothing depends on the route itself
The ordering is load-bearing. Attempting enforcement before identity is reliable produces a policy layer built on a directory that does not reflect who works here, which fails in both directions at once. Running enforcement before a log-only period produces an outage on day one and a permanent organisational allergy to the project. And retiring the VPN early, while some application still depends on raw network reachability, guarantees an emergency exception that reinstates the flat network you were trying to leave.
What Zero Trust Does Not Fix
- Application-layer defects. A correctly authenticated, correctly authorized session still reaches a broken endpoint. Missing object-level authorization inside the application is invisible to the proxy in front of it. That work stays with AppSec and threat modeling.
- Authorized misuse. Someone exfiltrating data they are entitled to read is, to the policy engine, a normal day. Detection here comes from behavioural analysis and data controls, not from access control.
- Credential and session theft where everything looks right. If both the identity and the device check out, the request is allowed. Continuous evaluation, short sessions, and phishing-resistant factors narrow the window; they do not close it.
- Availability. You have added a component to the critical path of every request. That is a real cost, and it must be designed for rather than discovered.
- Complexity. Policy is code without tests unless you give it tests. A large, unowned policy set is its own kind of outage waiting to happen.
Failure Modes
| Anti-pattern | What it produces | Fix |
|---|---|---|
| Buying a gateway and declaring victory | A perimeter with a new logo; internal traffic still implicitly trusted | Measure by what fraction of access decisions are per-request and logged |
| Enforcing before identity is trustworthy | Policy computed from a directory that is wrong in both directions | Fix lifecycle and factors first; see IAM |
| A legacy VLAN kept “temporarily” | The old perimeter survives as the path of least resistance | Track it as an exception with an owner and an end date |
| Device posture tuned too strictly | Blocked legitimate work, then broad permanent exemptions | Start with the strongest rule you can support, and use step-up rather than denial |
| Enforcement everywhere, logging nowhere | You cannot answer who accessed what during an incident | Decision logs are part of the enforcement point, not an add-on |
| Break-glass path is a flat network | The emergency path is the softest path, and attackers look for it | Break-glass is monitored, alarmed, time-boxed and rehearsed |
| Policy with no owner or tests | Rules accumulate, nobody removes any, effective access is unknowable | Policy as code: versioned, reviewed, tested, with an effective-access query |
| Undefined behaviour when the policy plane is down | Silent fail-open discovered during the incident | Choose the failure mode explicitly and test it |
When Not to Do This
A small estate with one directory, enforced multi-factor authentication, host firewalls, and a handful of applications already has most of the achievable benefit. Adding a policy plane there buys complexity and an extra outage surface for a marginal gain.
Operational and industrial networks are a different case: legacy protocols that cannot carry identity, devices that cannot run an agent, and equipment with a service life measured in decades. The correct answer there is usually hard segmentation and monitored gateways, not per-request authorization inside the protocol.
And if identity is not yet reliable — if offboarding takes weeks, if local accounts bypass single sign-on, if shared credentials are still normal — then zero trust is theatre. The policy engine will faithfully authorize the wrong people. Fix identity first; the rest of this is a sequence of manageable migrations once that holds.
Checklist
- One authoritative directory, with working joiner-mover-leaver
- Phishing-resistant factors on administrative and sensitive paths
- Applications reachable through an enforcement point, not a network route
- Device posture is an input to policy, with a step-up path instead of a hard block
- Every workload has a platform-issued identity; no shared static keys
- Service-to-service calls mutually authenticated and authorized per operation
- Policy stored as code, reviewed and tested like code
- Decision logs retained and queryable by resource and by principal
- Policy-plane failure behaviour chosen deliberately and exercised
- Break-glass and legacy exceptions have owners and expiry dates
- The VPN is retired only after nothing depends on the route
Related reading in this section: identity and access management is the prerequisite layer, application security and DevSecOps covers the defects a valid session still reaches, threat modeling tells you which trust boundary to enforce first, and compliance engineering turns decision logs into evidence. See Networking for the transport layer underneath and Security & Compliance for the section index.