ESC
Type to search guides, tutorials, and reference documentation.
Verified by Garnet Grid

Zero Trust Architecture

Zero trust explained as an engineering change rather than a product: per-request authorization, policy decision and enforcement points, device posture, workload identity and mTLS, a staged migration off the VPN, and what zero trust does not fix.

Zero trust is an architecture in which being on a particular network grants nothing. Every request carries an identity, is authorized against policy at the moment it is made, and is logged. It replaces the perimeter assumption — inside is trusted, outside is not — with a per-request decision taken as close to the resource as possible.

It is not a product, and no vendor can sell it to you as one. It is a change to where authorization happens, which means it is mostly an identity project, a policy project, and a migration project. Teams that buy a gateway and declare the work done usually end up with a perimeter that costs more.


What Zero Trust Is

The perimeter model made one bet: authenticate hard at the boundary, then let traffic inside move freely. It failed on contact with reality — laptops leave the office, software-as-a-service lives outside the boundary, and an attacker with a foothold inherits the same free movement as a legitimate employee.

Zero trust makes a different bet. Trust is never implied by location; it is computed per request from identity, device state, and context, and it expires. The vocabulary most commonly used for the components — policy engine, policy administrator, policy enforcement point — comes from the NIST zero trust architecture publication (SP 800-207), and it is worth adopting because it separates the three jobs cleanly.


The Control Loop

                    signals
        ┌──────────────────────────────┐
        │ identity + factor strength   │
        │ device posture               │
        │ resource sensitivity         │
        │ request context (time, geo)  │
        │ live risk / anomaly signals  │
        └──────────────┬───────────────┘
                       ▼
                 POLICY ENGINE  ── decides allow / deny / step-up
                       │
                 POLICY ADMIN   ── issues, refreshes, REVOKES the session
                       │
 request ─────▶  ENFORCEMENT POINT  ─────▶  resource
                       │                      (app, API, database,
                       ▼                       host, workload)
                 decision log
                 (who, what, when, why, allowed?)

The decision is RE-EVALUATED, not granted once.
Network position contributes NOTHING to it.

Three design consequences follow immediately. The enforcement point must sit in the data path, which makes it a latency and availability dependency. The policy engine must be reachable and fast, which usually means caching decisions for a bounded interval and accepting a small revocation delay. And the decision log is not optional telemetry — it is the only way to answer who reached what, which is the question asked during every incident and every audit.

Decide deliberately what happens when the policy plane is unreachable. Failing closed protects the resource and risks an outage; failing open protects availability and quietly disables the control at exactly the moment an attacker would most like it disabled. Whichever you choose, choose it explicitly, write it down, and test it — this behaviour is almost never exercised before the day it matters.


User Access

For human access, the change is from “the VPN grants a network route” to “the proxy grants one application, per request”. An identity-aware proxy terminates the connection, authenticates the user against the identity provider, evaluates policy, and only then forwards to the application. The user never gets a route to anything else, so a compromised laptop cannot scan a subnet it was never given.

Device posture becomes a first-class policy input: is this device managed, encrypted, patched, running the controls the organisation requires. This is what makes the difference between authenticating a person and authorizing a session, and it is the input most likely to cause friction — a posture rule that blocks legitimate work will be exempted, and the exemption will be permanent. Step-up authentication is the pressure valve: allow ordinary work on ordinary evidence, and require a stronger factor for the small set of sensitive operations.


Workload to Workload

Most traffic in a modern estate is service to service, and most segmentation projects stop at the north-south edge. The east-west requirement is the same three ingredients: each workload has a verifiable identity, calls are mutually authenticated, and authorization is expressed per caller and per operation rather than per IP address.

Identity for workloads is issued by the platform — short-lived certificates or tokens tied to what the workload is, not to a secret someone pasted into a configuration file. Mutual TLS then gives both authentication and transport protection in one mechanism. A service mesh is one way to deliver this; libraries and sidecars are others. The requirement is identity, policy, and telemetry — not any particular implementation.

Segmenting by identity rather than by address is the point. Addresses are recycled, ephemeral, and forged; a rule written against a workload identity survives a redeploy, and a rule written against a subnet decays into a permanent allow-any that nobody dares remove.


A Staged Migration

1. inventory        who uses which application, from where, on what
2. identity first   one directory · MFA · joiner-mover-leaver that works
3. one application  put it behind an identity-aware proxy
4. log-only         run the policy without enforcing; read the would-be denies
5. enforce          only once the deny set is understood and owned
6. repeat           by sensitivity, not by ease
7. east-west        workload identity + mTLS for the crown jewels
8. retire the VPN   LAST — once nothing depends on the route itself

The ordering is load-bearing. Attempting enforcement before identity is reliable produces a policy layer built on a directory that does not reflect who works here, which fails in both directions at once. Running enforcement before a log-only period produces an outage on day one and a permanent organisational allergy to the project. And retiring the VPN early, while some application still depends on raw network reachability, guarantees an emergency exception that reinstates the flat network you were trying to leave.


What Zero Trust Does Not Fix

  • Application-layer defects. A correctly authenticated, correctly authorized session still reaches a broken endpoint. Missing object-level authorization inside the application is invisible to the proxy in front of it. That work stays with AppSec and threat modeling.
  • Authorized misuse. Someone exfiltrating data they are entitled to read is, to the policy engine, a normal day. Detection here comes from behavioural analysis and data controls, not from access control.
  • Credential and session theft where everything looks right. If both the identity and the device check out, the request is allowed. Continuous evaluation, short sessions, and phishing-resistant factors narrow the window; they do not close it.
  • Availability. You have added a component to the critical path of every request. That is a real cost, and it must be designed for rather than discovered.
  • Complexity. Policy is code without tests unless you give it tests. A large, unowned policy set is its own kind of outage waiting to happen.

Failure Modes

Anti-patternWhat it producesFix
Buying a gateway and declaring victoryA perimeter with a new logo; internal traffic still implicitly trustedMeasure by what fraction of access decisions are per-request and logged
Enforcing before identity is trustworthyPolicy computed from a directory that is wrong in both directionsFix lifecycle and factors first; see IAM
A legacy VLAN kept “temporarily”The old perimeter survives as the path of least resistanceTrack it as an exception with an owner and an end date
Device posture tuned too strictlyBlocked legitimate work, then broad permanent exemptionsStart with the strongest rule you can support, and use step-up rather than denial
Enforcement everywhere, logging nowhereYou cannot answer who accessed what during an incidentDecision logs are part of the enforcement point, not an add-on
Break-glass path is a flat networkThe emergency path is the softest path, and attackers look for itBreak-glass is monitored, alarmed, time-boxed and rehearsed
Policy with no owner or testsRules accumulate, nobody removes any, effective access is unknowablePolicy as code: versioned, reviewed, tested, with an effective-access query
Undefined behaviour when the policy plane is downSilent fail-open discovered during the incidentChoose the failure mode explicitly and test it

When Not to Do This

A small estate with one directory, enforced multi-factor authentication, host firewalls, and a handful of applications already has most of the achievable benefit. Adding a policy plane there buys complexity and an extra outage surface for a marginal gain.

Operational and industrial networks are a different case: legacy protocols that cannot carry identity, devices that cannot run an agent, and equipment with a service life measured in decades. The correct answer there is usually hard segmentation and monitored gateways, not per-request authorization inside the protocol.

And if identity is not yet reliable — if offboarding takes weeks, if local accounts bypass single sign-on, if shared credentials are still normal — then zero trust is theatre. The policy engine will faithfully authorize the wrong people. Fix identity first; the rest of this is a sequence of manageable migrations once that holds.


Checklist

  • One authoritative directory, with working joiner-mover-leaver
  • Phishing-resistant factors on administrative and sensitive paths
  • Applications reachable through an enforcement point, not a network route
  • Device posture is an input to policy, with a step-up path instead of a hard block
  • Every workload has a platform-issued identity; no shared static keys
  • Service-to-service calls mutually authenticated and authorized per operation
  • Policy stored as code, reviewed and tested like code
  • Decision logs retained and queryable by resource and by principal
  • Policy-plane failure behaviour chosen deliberately and exercised
  • Break-glass and legacy exceptions have owners and expiry dates
  • The VPN is retired only after nothing depends on the route

Related reading in this section: identity and access management is the prerequisite layer, application security and DevSecOps covers the defects a valid session still reaches, threat modeling tells you which trust boundary to enforce first, and compliance engineering turns decision logs into evidence. See Networking for the transport layer underneath and Security & Compliance for the section index.

Jakub Dimitri Rezayev
Jakub Dimitri Rezayev
Founder & Chief Architect • Garnet Grid Consulting

Jakub holds an M.S. in Customer Intelligence & Analytics and a B.S. in Finance & Computer Science from Pace University. With deep expertise spanning D365 F&O, Azure, Power BI, and AI/ML systems, he architects enterprise solutions that bridge legacy systems and modern technology — and has led multi-million dollar ERP implementations for Fortune 500 supply chains.

View Full Profile →