AWS Architecture
How AWS is structured for production work: accounts and regions as blast-radius boundaries, the IAM evaluation model, VPC networking, and the failure modes that catch teams building on it.
Amazon Web Services is a set of independently-operated infrastructure services that you rent through a shared control plane and a single identity system. Almost every difficulty teams hit on AWS comes from misunderstanding one of three things: where the boundaries are (accounts and regions), how permission is actually decided (IAM policy evaluation), and which parts of a service are regional versus global. Get those right and the service catalog is mostly detail.
Accounts and Regions Are the Real Boundaries
A region is an independent deployment of AWS in a geography. Regions are deliberately isolated from one another: a control-plane problem in one region is not supposed to propagate to another, and most services replicate nothing across regions unless you explicitly configure it. Inside a region, Availability Zones are separate facilities with independent power and cooling, connected by high-bandwidth, low-latency links. The AZ is the unit you design around for infrastructure failure; the region is the unit you design around for correlated, large-scale failure.
The subtlety is that not everything is regional. Several services present a global namespace and keep their control plane in a single region — IAM, Route 53, CloudFront, and account billing among them. That means a control-plane disruption in that one region can stop you from changing global configuration even while your workloads in other regions keep serving traffic. This is the practical reason to separate data-plane dependencies from control-plane dependencies in any disaster-recovery plan: a runbook whose first step is "create a new IAM role" has a hidden single point of failure.
The Account as a Blast Radius
The AWS account is the strongest isolation boundary available. Quotas are per-account-per-region, most resource namespaces are per-account, and a compromised or runaway workload is contained by the account it runs in. This is why multi-account is the default enterprise pattern rather than a single account with careful tagging.
AWS Organizations arranges accounts into a tree of organizational units. Service Control Policies attach to that tree and define the maximum permissions available inside it. The critical property to internalize is that an SCP never grants anything — it only bounds what an identity policy can grant. An SCP that allows every action does not give anyone access; a principal still needs an IAM policy. Conversely, an SCP that denies an action makes it unreachable even for an account's own administrator. That is exactly what you want for guardrails like "no one may disable CloudTrail" or "no resources outside approved regions," and exactly what causes confusing access-denied errors when nobody remembers the SCP exists.
IAM: How Access Is Actually Decided
IAM is the service most worth learning properly, because every other service defers to it. Evaluation follows a fixed order: an explicit deny anywhere wins; otherwise an explicit allow is required; otherwise the request is denied by default. Policies do not average out and there is no priority ordering beyond deny-beats-allow.
Two policy families matter. Identity policies attach to a user, group, or role and say what that principal may do. Resource policies attach to the resource — an S3 bucket, a KMS key, an SQS queue — and say who may act on it. Within a single account, either one allowing the action is generally sufficient. Across accounts, both sides must allow: the calling principal needs an identity policy permitting the action, and the resource needs a policy trusting the caller. Cross-account access that "should work" and doesn't is usually one half of that pair missing. KMS is stricter still — the key policy is authoritative, and an identity policy alone cannot grant access to a key whose policy does not delegate to IAM.
Prefer Roles to Long-Lived Keys
Long-lived access keys are the most common source of serious AWS incidents, because they leak into source control, laptops, CI logs, and container images, and they do not expire. The alternative is always a role: instance profiles for EC2, task roles for ECS, IRSA or Pod Identity for EKS, execution roles for Lambda, and OIDC federation for CI systems and external workloads. All of these produce short-lived credentials from STS and remove the secret entirely. Treat any remaining static key as an exception that needs an owner, a rotation mechanism, and a justification.
The Primitives You Actually Compose
The catalog is large, but production systems mostly draw from a small set of primitives, chosen along an axis of how much operational responsibility you keep:
- Compute: EC2 gives you the machine and all of its maintenance. ECS and EKS give you a scheduler over machines you may or may not manage (Fargate removes the node). Lambda removes the machine entirely but imposes an event-driven, time-bounded, stateless execution model.
- Storage: S3 is regional object storage with lifecycle policies and multiple durability/retrieval tiers. EBS is a block volume attached to one instance in one AZ. EFS is a shared file system across AZs. The choice is usually decided by the access pattern, not by cost.
- Data: RDS and Aurora are managed relational engines; DynamoDB is a key-value/document store whose performance depends almost entirely on partition-key design. Choosing DynamoDB and then querying it relationally is a reliable way to build something slow and expensive.
- Messaging and events: SQS for queued work, SNS for fan-out, EventBridge for routing and schema-based integration, Kinesis for ordered streaming with replay.
Networking Inside an Account
A VPC is a regional private network with a CIDR range you choose. Subnets are AZ-scoped slices of that range, and a subnet is "public" only because its route table sends the default route to an internet gateway — there is no other flag. Private subnets reach the internet through a NAT gateway, which is itself AZ-scoped and charges for every gigabyte it processes.
That NAT charge is the mechanism behind a very common surprise: traffic from private subnets to AWS services such as S3 or DynamoDB will traverse NAT and be billed unless you add a VPC endpoint. Gateway endpoints (S3, DynamoDB) are route-table entries; interface endpoints, built on PrivateLink, place an ENI for the service in your subnet. Endpoints also keep the traffic off the public internet, which is usually the stronger argument. For deeper treatment of routing, peering, and hybrid connectivity, see cloud network engineering.
Security groups are stateful — allow inbound and the response is automatically permitted — and they can reference other security groups, which is the cleanest way to express "the app tier may talk to the database tier." Network ACLs are stateless and operate at the subnet level, so any rule set must explicitly permit return traffic on ephemeral ports. Most teams should use security groups as the primary control and reach for NACLs only for coarse, subnet-wide denies.
Common Failure Modes
- Single-AZ assumptions in a "highly available" design. An RDS instance without Multi-AZ, an EBS volume, or a single NAT gateway makes the whole stack AZ-fated regardless of how many application instances you run.
- Quota exhaustion during recovery. Service quotas are per-account-per-region and are not raised automatically. A failover that needs to launch capacity in a secondary region frequently discovers the limit at the worst possible time. Pre-raise quotas in the region you plan to fail into, and test it.
- Runbooks that depend on the control plane. If your recovery requires creating roles, modifying DNS, or deploying new infrastructure, it depends on services that may be degraded in the same event. Pre-provision the recovery path so the failover is a data-plane action.
- IAM sprawl. Wildcard actions and wildcard resources accumulate because they make errors go away. Use access analysis to generate least-privilege policies from observed usage rather than writing them by hand.
- Accidental data exposure. Public access on object storage is now blocked by default at several layers, but a resource policy, an ACL, or a presigned URL can still expose data. Enforce the block at the organization level rather than per-bucket.
- Eventual consistency in the control plane. IAM changes propagate; a role created and immediately assumed may fail. Retry rather than treating the first failure as authoritative.
- Cost driven by data movement, not compute. Cross-AZ traffic, NAT processing, and internet egress are billed per gigabyte and are invisible in an architecture diagram. See FinOps and cost engineering for how to make that visible.
When AWS Is the Right Default — and When It Isn't
AWS is a reasonable default when you need breadth of managed services, deep regional coverage, a large hiring pool, and a mature ecosystem of third-party tooling and reference architectures. It is particularly strong when your architecture is event-driven or when you want to assemble fine-grained primitives yourself.
It is a weaker fit when the organization already runs on another vendor's identity and productivity stack and would rather have one identity plane than two — that argument favors Azure. It is a weaker fit when the dominant workload is large-scale analytics or data-warehouse-first, where GCP's data services often need less assembly. And it is a poor fit for workloads with steady, predictable, high resource utilization and no need for elasticity, where owned or colocated hardware can be cheaper per unit of work — the honest comparison there includes staffing, resilience, and opportunity cost, not just the hardware.
Multi-cloud as a hedge deserves particular scepticism. Running the same workload on two providers means designing to the intersection of their features, doubling the operational surface, and paying to move data between them. It is justified by regulation, by a genuine acquisition, or by a specific service that exists in only one place — rarely by negotiating leverage alone.
Key Takeaways
- Accounts and regions are the boundaries that matter; treat the account as the blast radius and use Organizations plus SCPs to bound it.
- SCPs restrict, they never grant. Access requires an identity policy, and cross-account access requires the resource side to agree as well.
- Explicit deny beats explicit allow beats implicit deny. There is no other precedence rule to learn.
- Replace long-lived access keys with roles and federation everywhere you can.
- A subnet is public only because of its route table; NAT and cross-AZ traffic are metered, so data movement is an architectural decision.
- Design recovery to be a data-plane action, and raise quotas in the failover region before you need them.