ESC
Type to search guides, tutorials, and reference documentation.

GCP Solutions Architecture

How Google Cloud is structured: the project-centred resource hierarchy, IAM inheritance, the global VPC model, and the data and serverless services that define the platform.

Google Cloud Platform is a cloud built around two structural choices that distinguish it from its peers: the project is the unit that owns almost everything, and the virtual private cloud is a global resource rather than a regional one. Most of what feels unfamiliar when moving to GCP traces back to one of those two facts. The service catalog is smaller than its competitors' and more opinionated, which means less assembly required when your workload matches its shape and more friction when it doesn't.


The Resource Hierarchy

The hierarchy is organization → folders → projects → resources. The organization node represents your domain and is the root for policy. Folders nest arbitrarily and exist to group projects for governance. The project is where the real work happens: it is the boundary for billing attribution, for API enablement, for quota, for service accounts, and for most resource namespaces.

That last point is worth dwelling on. On GCP you do not simply call an API — you first enable that API on the project. A deployment that fails with a "service not enabled" error is not a permission problem, and no amount of IAM will fix it. Likewise, quotas are per-project-per-region, so splitting workloads across many small projects both limits blast radius and multiplies your available headroom, while consolidating into one large project concentrates risk and hits limits sooner.

Projects are also the cleanest lifecycle boundary. Deleting a project removes what it contains, after a recovery window, which makes "one project per environment per service" a practical pattern for ephemeral and per-team infrastructure. The cost of that pattern is a proliferation of projects to govern — which is what folders and inherited policy are for.

Organization Policy Constraints

Separate from IAM, organization policy constraints define what configurations are permissible anywhere in a subtree. They are the guardrail layer: restricting which regions resources may be created in, forbidding external IP addresses on virtual machines, requiring shielded VMs, restricting which domains can be granted access, and disabling service account key creation. Because they are evaluated at resource-creation time and inherited downward, they hold even against a project owner. As with any deny-shaped control, roll them out in a permissive posture first and measure what would break before enforcing.

IAM and Inheritance

GCP IAM binds a principal to a role at a node in the hierarchy, and the binding is inherited by everything beneath it. The effective permission set for a principal on a resource is the union of every binding from the organization down to that resource. There are also deny policies, which are evaluated before allow bindings, but the common case is purely additive inheritance — and that is the trap. A viewer role granted at the folder level applies to every project that will ever be created under that folder. Grant high in the tree only what is true for the whole subtree.

Roles come in three flavours. Basic roles (owner, editor, viewer) are coarse legacy roles that should almost never be used in production; editor in particular carries a very wide permission set. Predefined roles are service-specific and maintained by Google. Custom roles let you assemble a permission list yourself, at the cost of maintaining it as services add permissions.

Service Accounts Are the Main Risk Surface

A service account is both an identity and a resource, which means there are two distinct grants to reason about: what the service account may do, and who may act as the service account. The second is frequently overlooked, and a principal with permission to impersonate a highly-privileged service account effectively holds that privilege.

Downloadable service account keys are the single most common serious exposure on GCP, for the same reason static keys are on any platform: they are long-lived files that end up in repositories and images. The platform provides alternatives for every common case — the attached service account on a Compute Engine instance or Cloud Run service, Workload Identity for GKE pods, and Workload Identity Federation for external systems such as CI runners or another cloud. Disable key creation by organization policy and treat any surviving key as an exception with an owner and an expiry.

The Global VPC Model

A GCP VPC is global. Subnets are regional, and resources in different regions on the same VPC route to each other over Google's backbone without a peering relationship, a gateway, or a transit appliance. Teams arriving from other clouds routinely over-engineer this, building per-region networks and inter-region plumbing that the platform does not require.

Shared VPC extends this further: a host project owns the network, and service projects attach workloads to its subnets. This lets a central network team own address space, routing, and firewall policy while application teams retain full control of their own projects. VPC Network Peering, by contrast, behaves like peering everywhere else — it is non-transitive and requires non-overlapping ranges.

Firewall rules are VPC-level, stateful, evaluated by priority, and target resources by network tag or by service account. Targeting by service account is the stronger pattern: tags are freely assignable by anyone who can edit an instance, whereas binding a rule to a service account ties the network permission to an identity that IAM already governs.

For reaching Google APIs without public addresses, Private Google Access allows instances with only internal IPs to reach Google service endpoints, and Private Service Connect publishes a service behind an address inside your own VPC. Global external load balancers use a single anycast address announced from many points of presence, so client traffic enters Google's network close to the user and traverses the backbone rather than the public internet for the remaining distance. See cloud network engineering for the vendor-neutral treatment of these patterns.

Compute and the Data Platform

Compute spans the usual range: Compute Engine for virtual machines with managed instance groups for autoscaling; GKE for Kubernetes, in Standard mode where you manage node pools or Autopilot where you do not; Cloud Run for request-driven containers that scale to zero; Cloud Functions for event handlers. Cloud Run is the distinctive one — it accepts an ordinary container image and a port, removes node management entirely, and is a reasonable default for stateless HTTP services that do not need the full Kubernetes surface.

The data services are where GCP's opinions are strongest. BigQuery separates storage from compute and charges either for bytes scanned by a query or for reserved capacity. That mechanism has a direct design consequence: partitioning and clustering a table reduces the bytes a query must read, and SELECT * on a wide table is expensive in a way that has no analogue in a conventional database. Spanner offers horizontally scalable relational storage with strong consistency across regions. Bigtable is a wide-column store for very high write throughput where the row key design determines everything. Pub/Sub provides at-least-once messaging with a global endpoint. Cloud Storage is the object store, with buckets that can be regional, dual-region, or multi-region.

Common Failure Modes

  • Forgetting that an API must be enabled. It looks like a permission failure and is not one. Enable services declaratively in your infrastructure code.
  • Over-granting at the folder or organization level. Inheritance is silent and permanent until someone audits it. Grant at the project unless the grant is genuinely universal.
  • Service account keys in source control and images. Disable key creation by policy and migrate to attached identities and federation.
  • Default service accounts left with broad roles. A workload running as a default compute service account with editor is a privilege-escalation path from any code execution bug.
  • Unbounded BigQuery scans. Because cost follows bytes read, an unpartitioned table plus an ad-hoc dashboard is a recurring bill. Require partition filters on large tables.
  • Firewall rules targeted by network tag. Anyone who can edit an instance can attach a tag and grant themselves the rule. Target by service account instead.
  • Per-project quota surprises. Quotas are per-project-per-region. A recovery plan that concentrates workloads into one project may exceed a limit that was never visible in steady state.
  • Multi-region buckets assumed to be a backup. Geographic redundancy protects against infrastructure loss, not against a delete or an overwrite. Versioning and retention policies are separate controls.

When GCP Fits — and When It Doesn't

GCP is the strongest fit when analytics is the centre of gravity — when the architecture is fundamentally "get data into a warehouse and query it" — because BigQuery removes a large amount of infrastructure that would otherwise need building and operating. It is a good fit for Kubernetes-native organisations, for teams that want container-first serverless without adopting a proprietary function model, and for machine-learning workloads that benefit from tight integration between storage, warehouse, and accelerators.

It is a weaker fit when you need the broadest catalog of niche managed services or the deepest supply of existing reference architectures and hiring pool, which favours AWS. It is a weaker fit when the organisation's identity and productivity stack is Microsoft and consolidating identity matters more than platform preference, which favours Azure. And as everywhere, a workload with flat, fully-utilised, entirely predictable demand gets the least value from any elastic platform.

Key Takeaways

  • The project is the unit of billing, quota, API enablement, and lifecycle — design project layout before anything else.
  • IAM is additive and inherits downward. A grant at a folder applies to projects that do not exist yet.
  • Service accounts have two questions attached: what they can do, and who can act as them.
  • The VPC is global and subnets are regional; do not rebuild inter-region connectivity you already have.
  • BigQuery cost follows bytes scanned, so partitioning and clustering are cost architecture, not tuning.
  • Target firewall rules at service accounts rather than network tags so that network policy inherits IAM's governance.