Cloud Network Engineering
Cloud networking from first principles: address planning, routing and egress, connecting networks without transitivity, DNS, stateful versus stateless filtering, and the failure modes that are hardest to diagnose.
Cloud networking is the layer where a virtual network is defined in software but still obeys every rule of the physical one underneath it. Addresses still have to be unique, routes still have to be symmetric, packets still have a maximum size, and a stateful firewall still needs to see both directions of a flow. The abstractions hide the cabling, not the protocol. Most of the hardest production incidents in a cloud estate are networking incidents, because the failure is often partial, intermittent, and located somewhere nobody's dashboard is watching.
The Virtual Network and Address Planning
The core primitive — a VPC on AWS and GCP, a virtual network on Azure — is a private address space you define, carved into subnets. Subnets are zone- or region-scoped depending on the provider, which is the first thing to check when porting a design between clouds: on AWS a subnet belongs to one availability zone, whereas on Azure and GCP a subnet is regional and spans zones.
Address planning is the one decision that is genuinely hard to reverse. Renumbering a running network means touching every firewall rule, every route, every peering, every hard-coded address in configuration, and every DNS record — usually while the service is live. Three rules pay for themselves repeatedly:
- Allocate from a central registry. One authoritative record of which private ranges are assigned to which network, including on-premises and partner networks. Teams that allocate ad hoc will eventually collide.
- Leave room. Reserve more space per network than you need, and allocate on aligned boundaries so a range can grow without fragmenting. Subnets that are sized exactly to today's instance count run out during the first autoscaling event or the first Kubernetes cluster, which consumes addresses at a rate that surprises people.
- Avoid the obvious ranges. The most commonly used private blocks are also the ones a future acquisition, partner, or VPN client will already be using. Picking a less-trafficked range costs nothing now and avoids a collision later.
Kubernetes deserves specific mention because it multiplies address consumption: depending on the networking mode, every pod may consume an address from the subnet, so a cluster can exhaust a range that looked generous when sized for virtual machines.
Routing, Ingress, and Egress
A route table maps destination prefixes to next hops, and longest-prefix-match decides. In cloud networks there is no routing protocol running between your subnets by default — the platform installs a local route for the network's own range, and everything else is what you or the platform adds. A subnet is "public" purely because its route table points the default route at an internet gateway. There is no other property that makes it public, which is why a misconfigured route table is the most common way a resource ends up exposed or isolated when it should be the opposite.
Outbound access from private subnets goes through a NAT service or a proxy. Two properties of NAT matter operationally. First, it is metered: managed NAT typically charges both for existence and for every gigabyte processed, which makes it a recurring cost line for traffic that could often bypass it. Second, it is a shared translation table with a finite number of source ports per destination. A workload that opens very many simultaneous connections to a single destination address and port can exhaust the available translations and see intermittent connection failures while the network looks entirely healthy — an incident that presents as a mysterious application error and is invisible to ping and traceroute.
Reaching provider-managed services privately — through VPC endpoints, private endpoints, or Private Service Connect depending on the cloud — solves both problems at once: the traffic never touches NAT or the public internet, which removes the charge and shrinks the exposure.
On ingress, the practical distinction is between layer 4 and layer 7 load balancing. An L4 balancer forwards connections and preserves protocol transparency, which suits non-HTTP protocols and very high throughput. An L7 balancer terminates the connection, which enables path-based routing, header manipulation, TLS termination, and per-request retries — at the cost of being an active participant in the request and therefore another thing that can fail or be misconfigured. Whichever you use, the health check configuration is the part that actually determines availability: a check that hits a static path proves the process is listening and nothing more, while a check that exercises a downstream dependency can take the entire fleet out of service when that dependency has a transient problem. The right target is usually a shallow endpoint that verifies the instance can serve, not one that verifies the whole system is well.
Connecting Networks
The single most important property to internalise: peering is not transitive on any major cloud. If network A peers with B and B peers with C, A cannot reach C. This is not an oversight; it prevents accidental full-mesh reachability and keeps routing tractable. Every hub-and-spoke topology in every cloud exists because of this fact.
That leaves three ways to connect more than a handful of networks:
- Full-mesh peering is simple and low-latency but scales quadratically in relationships, and every new network requires touching every existing one. Fine for a few networks; unmanageable beyond that.
- A transit hub — a managed transit service or a self-run appliance in a hub network — gives a single attachment point per network and one place to apply inspection and policy. It adds a hop, a dependency, and a bandwidth ceiling to plan for, and it becomes a shared failure domain that deserves the same redundancy thinking as any other critical component.
- Shared networks, where several projects or subscriptions attach workloads to one centrally-owned network, sidestep inter-network routing entirely at the cost of a central team owning the address space.
For hybrid connectivity, the choice is between IPsec VPN over the public internet — fast to stand up, cheap, and subject to internet path variability — and a dedicated circuit such as a private interconnect, which offers predictable bandwidth and latency at the cost of lead time, contractual commitment, and physical redundancy planning. A common and defensible pattern is a dedicated circuit as primary with a VPN as automatic backup, which is only genuinely redundant if the two paths do not share a device, a facility, or a provider.
Overlapping address ranges are the chronic problem in hybrid and post-acquisition environments, where two networks that must now talk both use the same private block. The options are all unpleasant: renumber one side, deploy NAT at the boundary so each side sees the other through a translated range, or proxy at the application layer. Renumbering is the only one that leaves a clean network behind, which is the argument for disciplined address planning before it becomes necessary.
DNS and Service Discovery
DNS is where cloud network problems most often become application problems. A handful of mechanisms account for most of them.
Split-horizon resolution — the same name resolving to a private address inside the network and a public address outside — is how private endpoints work, and it only works if the resolver inside the network is authoritative for that zone. A private link whose DNS still resolves publicly gives you a working public path and the false belief that traffic is private.
Hybrid resolution requires forwarding rules in both directions: cloud workloads resolving on-premises names, and on-premises clients resolving cloud names. It is common to build one direction and discover the other during an incident.
TTL is a failover parameter. A DNS-based failover cannot be faster than the record's TTL plus whatever caching clients and intermediate resolvers apply regardless of what you set. Many runtimes and connection pools also cache a resolved address for the life of the process, so lowering TTL changes nothing for them. If you need fast failover, use an anycast address or a load balancer that keeps a stable address, and treat DNS as the slow path.
Stateful and Stateless Filtering
Cloud platforms offer both kinds of filter and they behave very differently:
| Stateful (security group, NSG) | Stateless (network ACL) | |
|---|---|---|
| Return traffic | Permitted automatically | Must be explicitly allowed |
| Attached to | Instance or interface | Subnet |
| Typical use | Primary workload policy | Coarse subnet-wide deny |
| Common mistake | Over-broad source ranges | Forgetting ephemeral return ports |
Stateful filters should carry the day-to-day policy, because expressing intent as "the application tier may reach the database tier" — by referencing another group or identity rather than an address range — survives autoscaling and re-addressing. Stateless ACLs are a blunt instrument best reserved for a small number of subnet-level denies. Where the platform allows a rule to target a workload identity rather than a tag or address, prefer it: that ties network permission to something the identity system already governs.
Cost and Performance Are the Same Decision
Data movement is metered, and the meter usually follows distance: traffic within a zone is cheapest, traffic between zones in a region costs more, traffic between regions more again, and traffic to the internet most of all. Traffic through a managed NAT or an inspection appliance is charged for processing on top.
This creates a genuine tension rather than a simple optimisation. Keeping chatty services zone-local reduces both latency and transfer cost — and concentrates them in a single failure domain. Spreading across zones buys availability and pays for it in cross-zone traffic on every request. The right answer depends on how chatty the services are and how much availability the workload actually requires, and it is worth deciding deliberately rather than inheriting it from a default. Architectural changes that reduce data movement — caching at the edge, colocating a service with the data it reads, batching, compressing, and filtering at the source instead of the destination — improve both axes at once. See FinOps and cost engineering for how these line items get attributed to the teams that create them.
Common Failure Modes
- Address exhaustion. Subnets sized for today. Discovered during an autoscaling event or a cluster upgrade, when there is no time to renumber.
- Overlapping ranges after a merger or a new partner. No clean fix exists once both sides are in production.
- MTU and black-hole paths. Tunnels and overlays reduce the usable packet size. If path MTU discovery is broken because something drops the ICMP messages that carry it, small requests succeed and large ones hang — which presents as an application bug, not a network one.
- Asymmetric routing. A flow that leaves by one path and returns by another will be dropped by any stateful device that only sees half of it. Common after adding a second gateway or a new peering without reviewing routes.
- NAT port exhaustion. Intermittent connection failures under high connection concurrency to a single destination, with every health signal green.
- Health checks that are too deep or too shallow. Too deep and a dependency blip removes the whole fleet; too shallow and broken instances keep receiving traffic.
- Assuming peering is transitive. Spoke-to-spoke traffic silently fails until a hub route is added.
- Private endpoints without private DNS. The network path is built, the name still resolves publicly, and nothing actually changed.
- Security group rules pinned to addresses. They rot as instances are replaced. Reference groups or identities instead.
- Single NAT gateway or single tunnel. A zone-scoped egress path makes a multi-zone workload zone-fated for anything that needs outbound access.
Key Takeaways
- Address planning is the least reversible decision in the network. Allocate centrally, leave headroom, and account for Kubernetes address consumption.
- A subnet is public only because of its route table. Routing, not naming, determines exposure.
- Peering is never transitive. Hub-and-spoke and transit hubs exist for that reason alone.
- Private service endpoints remove NAT cost and internet exposure simultaneously — but only if DNS resolves to the private address.
- Stateful filters carry policy; stateless ACLs are for coarse denies and must allow return traffic explicitly.
- Data movement is metered by distance, so the zone-locality decision trades latency and cost against availability. Make it deliberately.
- The ugliest incidents — MTU black holes, asymmetric routing, NAT port exhaustion — all look like application bugs. Know their signatures before you need them.