Google Cloud Platform Services and Architecture Questions
Google Cloud Platform's core services and architecture: Compute Engine, Cloud Run, GKE, Cloud Storage, VPC, managed databases (Cloud SQL, Spanner, Firestore, Bigtable), and BigQuery-adjacent data services. Covers GCP service selection, networking, IAM and security specifics, cost and quota management, and reference patterns for building on the platform. For provider-agnostic compute, storage, or networking concepts, see the cross-cloud entries.
Design a secure GKE cluster for an enterprise with many teams sharing it. Cover network policies, RBAC, namespace isolation, Workload Identity, and how you'd enforce that only scanned, trusted images get admitted. How would you explain this design to a security team without making developers' lives miserable?
Sample Answer
Direct answer
Four independent layers, each closing a different attack surface: network policies restrict pod-to-pod traffic to only what's declared necessary, containing lateral movement; role-based access control (RBAC) and namespace isolation restrict who or what can act on the Kubernetes API itself; Workload Identity replaces long-lived service account keys inside pods with short-lived, per-workload credentials; and image admission control ensures only scanned, trusted images get scheduled at all. To a security team, these map onto four plain questions, can an attacker move sideways, can an attacker escalate privilege through the API, can a leaked credential be used broadly, and can untrusted code run here in the first place, which is a much stronger pitch than a bare feature list.
The four layers
Network policies, for lateral movement
Kubernetes NetworkPolicy objects are deny-by-default once any policy exists in a namespace, define exactly which pods can talk to which other pods, on which ports, instead of the flat, fully-open pod network GKE (Google Kubernetes Engine) gives you by default. This stops a compromised pod in one namespace from freely reaching a database-tier pod in another namespace just because they share the same cluster network.
RBAC and namespace isolation, for API-level control
RBAC governs who, or which service account, can perform which verbs on which API resource types, scoped to a namespace via RoleBindings rather than cluster-wide via ClusterRoleBindings wherever possible. Combined with namespace isolation, each team or environment in its own namespace with its own RBAC bindings, this stops a developer or a compromised CI/CD identity in one team's namespace from reading secrets or modifying deployments in another team's namespace.
Workload Identity, for credential exposure
Workload Identity binds a Kubernetes service account to a Google Cloud service account, so a pod authenticates to Google Cloud APIs with short-lived, automatically rotated credentials tied to its specific Kubernetes service account, with no downloaded key living inside the pod or the container image. This closes the same leaked-key risk that makes downloaded service account keys dangerous in general, applied specifically to workloads running inside the cluster.
Admission control, for image trust
An admission controller, Policy Controller for general policy or Binary Authorization specifically for image provenance, enforces that a pod can only be scheduled if its image passes policy: signed by a trusted build pipeline, scanned with no unresolved critical vulnerabilities, sourced from an approved registry. This answers "how do we know nothing untrusted is even running," independent of how well-behaved that code turns out to be once it is running.
Presenting this without making developers miserable
Frame each control around what it does not require developers to do manually: Workload Identity means developers stop thinking about service account keys at all, strictly less toil, not more. Image scanning happens in the CI pipeline before a developer's normal merge and deploy flow, so a failure surfaces as a familiar pipeline check, not a mysterious runtime rejection. Network policies should ship with a sensible default, allow within your own namespace, deny cross-namespace unless declared, so most teams never need to hand-write one. The one place friction is real, and should be named honestly, is a new cross-namespace dependency: a team that legitimately needs to call another team's service now has to explicitly request that network policy, a deliberate, reviewable speed bump, not an accident, and it's worth saying that out loud rather than pretending there's zero cost anywhere in this design.
Worked example
A developer's pod is compromised through an application-level vulnerability, unrelated to any of these four controls, they don't prevent the initial compromise, only what happens next. Network policy stops the compromised pod from reaching the database tier in a different namespace it was never declared to talk to. Even if it could reach something, RBAC means the pod's own service account has no permission to read Secrets or list workloads outside its own namespace. If the attacker tries to pivot to Google Cloud resources, Workload Identity means there's no downloaded key to steal, only a scoped, short-lived credential tied to that workload's narrow IAM grant. And the image that was compromised had already passed Binary Authorization, meaning the vulnerability was likely a runtime or application flaw rather than a known catalogued vulnerability (a CVE, a publicly documented security issue) the scan should have caught, a useful, honest thing to tell a security team, these controls narrow the blast radius of a compromise, they don't claim to prevent one.
Trade-offs and pitfalls
Rolling out NetworkPolicy cluster-wide with a default-deny stance and no migration plan is a common pitfall, it breaks every existing cross-namespace dependency at once and is exactly the kind of change that gives security controls a reputation for breaking things, roll it out namespace by namespace with existing traffic mapped first. Using ClusterRoleBindings out of convenience, because scoping every binding to a namespace is more setup work, quietly reintroduces the cross-team blast radius these controls exist to eliminate. Binary Authorization policies that are so strict any dependency update blocks a deploy, with no fast, reviewable exception path, teach teams to route around the control entirely, which defeats it. Being honest with the security team about the real, small, controlled friction this design does introduce buys credibility for the large amount of protection that costs developers nothing.
Design a CI/CD pipeline for container images that enforces supply-chain security before anything reaches GKE: vulnerability scanning, image signing, and blocking anything that fails policy at admission time using something like Binary Authorization. How would you prove provenance for an image that does get deployed?
Sample Answer
Direct answer
Build the image once, scan it, sign it with an attestor at each gate it clears, and promote that exact signed digest across environments (never rebuild between them). Binary Authorization (a GKE, Google Kubernetes Engine, admission controller that only lets a workload run if its image carries the attestations a policy requires) denies anything missing the right signatures at deploy time, on every cluster, not just the ones you remember to check by hand. Provenance for a deployed image comes from independently verifying the build metadata Cloud Build attaches to that image, not from trusting whoever tells you it's safe.
Pipeline stages
- Build. Cloud Build (GCP's managed continuous integration and continuous delivery, CI/CD, service) runs in a locked builder image with no arbitrary network egress, builds the container, and pushes it to Artifact Registry (GCP's managed image and package registry) addressed by its content digest (
sha256:...), never by a mutable tag alone. - Vulnerability scan. Artifact Analysis (Artifact Registry's built-in scanning service, formerly Container Analysis) scans the pushed digest. The pipeline fails above a defined severity threshold, for example any CRITICAL finding, or any HIGH finding with a published fix.
- Sign. Two distinct Binary Authorization attestors sign the same digest at different points: a
vuln-scan-passedattestor signs right after step 2 clears, and aqa-passedattestor signs only after smoke tests pass in staging (step 4). Each attestation is a cryptographic signature, backed by a Cloud KMS (Cloud Key Management Service) key, over that specific digest plus a note identifying which check ran. - Promote across environments. The identical digest moves dev to staging to prod. Each promotion re-runs smoke tests against that same digest in the next environment; passing adds an environment-scoped attestation. Rebuilding "just to bump a label" between environments breaks the whole chain, since the digest that gets attested is no longer the digest that gets deployed.
- Admission. Each GKE cluster enforces a Binary Authorization policy naming which attestations it requires: staging requires
vuln-scan-passedandqa-passed; prod additionally requires aprod-smoke-passedattestation from that environment's own smoke run. Anything missing one is denied and logged. An emergency bypass needs a scoped, time-boxed policy exception on record, not a manual click-through, or the control gets disabled the first time it's inconvenient.
flowchart LR
A[Source push] --> B[Cloud Build: unit test]
B --> C[Vulnerability scan: Artifact Analysis]
C -->|clean| D[Sign image: Binary Authorization attestor]
C -->|CVEs above threshold| X[Block build]
D --> E[Push to Artifact Registry]
E --> F[Deploy request to GKE]
F --> G{Admission check: Binary Authorization policy}
G -->|attestations present| H[Pod scheduled]
G -->|missing attestation| X
Proving provenance
Cloud Build generates build provenance metadata, in the SLSA (Supply-chain Levels for Software Artifacts) sense, alongside the image: the exact source commit, the builder identity, and the build steps that ran. For a deployed image, pull that digest's provenance (via the Artifact Registry image metadata or the Container Analysis Occurrence API) and confirm the source repository, commit, and builder service account match what you expect, then check that the same digest carries the Binary Authorization attestations your policy requires. Because every check keys off the immutable digest rather than a tag, an attacker who repoints a tag in the registry cannot silently change what a "trusted" reference actually points to.
Trade-offs and pitfalls
Binary Authorization only guards the GKE admission path. It does nothing if someone gets a shell on a node or a pod's service account is over-privileged, so it has to sit alongside node and workload hardening, not replace it. The real trust boundary is who can produce a valid signature: if the same CI service account that builds the image also holds signing rights for every attestor, a compromised pipeline can attest its own bad image, so the qa-passed attestor's signing identity should be separate from the one that runs the build and scan. The most common way this design gets quietly defeated is attesting to a tag instead of a digest, or rebuilding between staging and prod for a trivial reason, both erase the chain of custody the whole pipeline exists to create.
Architect a streaming ingestion pipeline handling roughly 1TB of events a day, landing in BigQuery (or Bigtable where it fits better), using Pub/Sub and Dataflow. How would you handle partitioning and windowing, keep it fault-tolerant under bursty load, and control cost?
Sample Answer
Direct answer
Land events in Pub/Sub as the durable ingestion buffer, process them in Dataflow for windowing and any transformation, and write to BigQuery for aggregate/analytical access or Bigtable for high-throughput point lookups keyed by entity (whichever the downstream access pattern actually needs, and possibly both from the same pipeline). At roughly 1TB (terabyte) a day, this is a moderate, well-trodden scale: fixed or sliding time windows on event time (not processing time) with a bounded allowed-lateness for stragglers, Dataflow's autoscaling plus Pub/Sub's own buffering absorb bursty load without manual capacity planning, and cost is controlled mainly by keeping the pipeline's window and trigger choices from fanning out more output than the business actually needs.
Structured elaboration
Partitioning and windowing. Use fixed windows (for example, 1-minute or 5-minute tumbling windows) on event time if the destination is BigQuery and the consumer wants regular aggregates, or session windows if the goal is grouping bursts of related activity per entity. Set an allowed lateness (how long to wait for a straggling event after the watermark passes) sized to your actual observed event delay, not a guess, and use a write.to_bigquery with a late-arrival trigger firing an update rather than dropping late data outright if correctness matters more than a hard latency bound. Partition the BigQuery destination table by the window's event date so downstream consumers get the pruning benefit for free.
Fault tolerance and retries under bursty load. Pub/Sub decouples the producers from the processing rate: publishers keep succeeding even if Dataflow temporarily falls behind, because messages sit durably in Pub/Sub (retained up to 31 days) until a subscriber pulls and acknowledges them. Pub/Sub's default delivery guarantee is at-least-once, so a Dataflow worker that dies mid-batch (a real risk during a burst, if autoscaling is adding and removing workers) simply causes those in-flight messages to be redelivered once the ack deadline passes, and Dataflow's own checkpointing resumes windowing state correctly rather than losing progress; this combination is what makes retries safe by default rather than something the pipeline author has to build. Dataflow's autoscaling adds workers as backlog grows and removes them as it drains; the main operational risk under a burst is not data loss but a temporarily growing end-to-end latency, which is a capacity/cost trade-off to size for explicitly (how much backlog, and for how long, is acceptable) rather than an afterthought. If a specific entity (say, a single device or user) needs its events processed in the order they were produced, an ordering key scoped to that entity's ID gives that guarantee cheaply for the traffic that actually needs it, without forcing every event in the topic through a single ordering constraint it doesn't need.
Cost control. The three real levers are: Dataflow worker count and machine size (an overprovisioned max-workers setting pays for idle capacity between bursts), the destination write pattern (BigQuery streaming inserts cost more per row than periodic micro-batch loads via the Storage Write API, so batch the writes if sub-second landing latency isn't actually required), and Pub/Sub subscription count (each additional subscription re-reads and re-bills the full message stream, so fan out downstream inside Dataflow rather than creating a new Pub/Sub subscription per consumer when the same transformed data serves multiple sinks).
flowchart LR
A[Event producers] --> B[Pub/Sub topic]
B --> C[Dataflow: windowing + transform]
C -->|aggregates| D[BigQuery]
C -->|point lookups by key| E[Bigtable]
Worked example
1TB/day, if events average 1KB (kilobyte) each, converts to a concrete throughput target:
1 TB=1012 bytes,86,400 s1012 bytes≈1.157×107 bytes/s≈11.57 MB/s 1,024 bytes/event11.57×106 bytes/s≈11,300 events/s averageReal traffic isn't flat: if the daily pattern peaks at, say, 6x the average during a business-hours burst, that's roughly 68,000 events/s at peak, well within a single Pub/Sub topic's throughput ceiling and a modest Dataflow autoscaling range, but it's the number that should actually drive your worker-count and subscription-quota planning, not the 11,300/s daily average.
Trade-offs & pitfalls
- Choosing BigQuery streaming inserts by default because "it's simple" is a common way to overpay: at this event rate, batching writes every 30 to 60 seconds via a Dataflow-managed load into BigQuery (or the Storage Write API's batch mode) is materially cheaper for a negligible latency cost, unless a specific downstream consumer genuinely needs sub-minute freshness.
- Allowed lateness set too short silently drops legitimately late data (mobile clients reconnecting after being offline are a classic source); set too long, it keeps window state open and increases Dataflow's memory footprint per key.
- A separate Pub/Sub subscription per downstream consumer multiplies your Pub/Sub bill by the number of consumers for the same underlying stream; prefer one subscription into Dataflow and fanning out from there when possible.
- This design doesn't need Bigtable at all if every downstream consumer only wants aggregates; adding it "for scale" when nobody needs sub-10ms point lookups is unnecessary operational surface.
A global financial application needs strong consistency for transactions at around 100,000 transactions per second. Compare Cloud Spanner, a sharded Cloud SQL deployment, and building a custom consensus layer, on transaction model, latency, operational complexity, and cost. How would you justify your recommendation to both architects and finance stakeholders?
Sample Answer
Direct answer
At roughly 100,000 transactions per second (TPS) with strong consistency required globally, recommend Cloud Spanner over both a sharded Cloud SQL deployment and a custom consensus layer: it's the only one of the three that gives you strong distributed consistency as a managed service at that scale, whereas sharded Cloud SQL pushes the hard distributed-transaction problem into application code, and a custom consensus layer means building and operating the exact category of distributed system Spanner already is, just without Google's years of production hardening behind it.
Structured elaboration
Cloud Spanner. Transaction model: native distributed ACID (atomicity, consistency, isolation, durability) transactions via a managed consensus protocol, with external consistency across the whole fleet. Latency: higher per-transaction than a single-machine database due to consensus overhead, but that overhead is fixed and well-characterized, since Google operates and tunes the underlying protocol. Operational complexity: primary-key and interleaving design matter a great deal, but once designed correctly, scaling is "add nodes," not "redesign your architecture." Cost: higher baseline cost than either alternative at low scale, but the cost scales roughly linearly with capacity added, a predictable curve at 100,000 TPS.
Sharded Cloud SQL. Transaction model: strong consistency exists only within one shard (one Cloud SQL instance); any transaction spanning two shards (say, a transfer between two accounts that happen to live on different shards) needs application-level distributed-transaction logic (a two-phase-commit-style protocol, or an eventual-consistency-with-compensation pattern) that the team has to design, build, and prove correct themselves. Latency: fast for single-shard transactions, but cross-shard transactions pay whatever latency and complexity cost the home-built coordination layer adds. Operational complexity: significant and grows with shard count, since re-sharding, rebalancing hot shards, and handling partial failures during cross-shard operations are all now the team's problem, not a managed service's. Cost: potentially lower in raw compute at first, but the true cost includes the engineering time to build and maintain a correct distributed-transaction layer, which is easy to underestimate and expensive to get wrong (a subtly incorrect two-phase-commit implementation is one of the harder classes of bugs to catch before it causes real financial damage).
Custom consensus layer. Transaction model: whatever the team implements, at the mercy of the team's own correctness (Paxos and Raft-family protocols are notoriously easy to get subtly wrong even for experienced distributed-systems engineers, particularly around edge cases like network partitions and leader re-election). Latency: potentially competitive with Spanner if implemented well, but "implemented well" for a financial system's consensus layer is a multi-year, dedicated-team effort at companies that have done this (the major cloud providers, some large exchanges), not a project most engineering organizations should take on for a single application. Operational complexity: the highest of the three by a wide margin, since the team now owns not just the application but the correctness and operations of a foundational distributed-systems primitive. Cost: the engineering cost (specialized distributed-systems expertise, extensive testing including formal verification for a financial use case, and ongoing operational burden) dwarfs the infrastructure cost difference versus the other two options for all but the largest, most specialized organizations.
Justifying the recommendation, to two different audiences:
- To architects: Spanner gives you the distributed-transaction correctness guarantee for free (as a managed capability) that the other two options make you build and prove yourselves; at 100,000 TPS with a financial-consistency requirement, that correctness guarantee is exactly the hard part, and it's the part you don't want to be reinventing under time pressure.
- To finance stakeholders: frame it as risk-adjusted cost, not just infrastructure cost. Spanner's higher line-item cost is the cost of eliminating a class of correctness risk (a bug in a home-built distributed-transaction layer causing a financial reconciliation error) whose downside, for a financial application, is disproportionately larger than the infrastructure cost difference between the three options; the cheapest-looking option (sharded Cloud SQL) carries hidden engineering cost and ongoing correctness risk that doesn't show up on an infrastructure invoice.
Worked example
Illustrating the order of magnitude of the "hidden cost" argument for sharded Cloud SQL: if cross-shard transactions are, conservatively, 10 percent of the 100,000 TPS (10,000 TPS) and each one needs a hand-built two-phase-commit round trip across shards, that's 10,000 distributed-transaction coordination events per second that the application-level code, not a managed service, is responsible for getting right, every single time, indefinitely. A single subtle bug in that path (say, a coordinator crash between the prepare and commit phase leaving a shard in an ambiguous state) is the kind of defect that doesn't show up in normal testing and surfaces as a real financial discrepancy in production, exactly the risk a staff-level recommendation to finance should name explicitly rather than only comparing sticker prices.
Trade-offs & pitfalls
- Recommending the custom consensus layer to "avoid vendor lock-in" or "save on Spanner's premium" without being explicit that it requires dedicated, ongoing distributed-systems expertise most teams don't have on staff is understating the real cost and risk of that option.
- Recommending sharded Cloud SQL without flagging that cross-shard transactions are the team's problem to solve correctly, not a solved problem the platform hands you, sets up an expectation gap that surfaces painfully later, typically after the sharding scheme is already load-bearing in production.
- Even with Spanner, primary-key and interleaving design still has to be done correctly at this scale; recommending Spanner isn't a substitute for that design work, it's a reason that work is tractable instead of a from-scratch distributed-systems project.
Propose a GCP resource organization strategy for an enterprise with multiple business units: when would you use the organization node, folders, projects, billing accounts, and labels? How would you map environments (prod, staging, dev) and shared services like CI/CD and logging, and what are the trade-offs between centralized and decentralized billing?
Sample Answer
Direct answer
Use the organization node as the single root for company-wide policy (org policy constraints and platform-admin IAM), folders to mirror how policy and billing actually need to inherit, and one project per workload per environment, never a shared project across a prod and a staging deployment of the same service. Labels cover everything that is a useful query dimension but not a security or billing boundary, a cost-center or team label, for instance, not a substitute for a folder.
Resource hierarchy
flowchart TD
Org[Organization node] --> FProd[Folder: production]
Org --> FNonProd[Folder: non-production]
Org --> FShared[Folder: shared services]
FProd --> PProdA[Project: bu-a-prod]
FProd --> PProdB[Project: bu-b-prod]
FNonProd --> PStaging[Project: staging]
FNonProd --> PDev[Project: dev]
FShared --> PCicd[Project: cicd]
FShared --> PLogging[Project: logging and monitoring]
An environment-first split, as drawn above, puts prod, non-production, and shared services each in their own top-level folder, with each business unit's projects nested underneath. Shared services, the CI/CD (continuous integration and continuous delivery) pipeline and centralized logging, live in their own folder that every environment's workloads read from or write to through explicit, least-privilege cross-project IAM grants, not by duplicating the pipeline per environment.
The four decisions this actually rests on
Folders are the natural place to draw the IAM boundary: an org policy or IAM role granted at the production folder inherits down to every business unit's prod project automatically, without repeating it project by project. Network isolation should follow the same split: a Shared VPC (a Virtual Private Cloud network shared across multiple projects from one designated host project, so those projects' resources can talk to each other privately) attached per environment, so a compromised staging workload cannot reach a production network path by construction, not by convention. Operational ownership is a separate axis from either of those: a business unit typically owns break/fix for its own prod projects, while a shared-services folder's projects (the CI/CD project, the logging project) are owned by a platform team responsible to every business unit that depends on them, worth naming explicitly since it does not follow automatically from the IAM or network split. Billing separation is the fourth and most consequential decision, covered below.
Worked example
Two business units, Payments and Logistics, each get a prod project and a staging/dev pair under the production and non-production folders respectively (payments-prod, payments-staging, and so on). A single shared-services folder holds one cicd project used by both business units' pipelines and one logging project that every other project's audit and application logs export into, so there is one place, not four, to review activity across the whole organization.
Centralized versus decentralized billing
Centralized billing, one billing account for the whole organization, gives a single invoice and lets committed-use or volume discounts pool across every business unit's spend, at the cost of poor per-team cost attribution unless labels and BigQuery (Google's serverless data warehouse) billing export are used deliberately, and a real concentration of risk in whoever can attach or detach projects from that one billing account. Decentralized billing, one billing account per business unit, gives clean chargeback and contains a runaway bill to the business unit that caused it, at the cost of forfeiting pooled discounts and multiplying the number of billing admins who need auditing. I would default to centralized billing with labels and billing export for chargeback, and only split billing accounts when a business unit is a genuinely separate legal or financial entity, a subsidiary, for example, or a contract specifically requires it.
Trade-offs and pitfalls
Nesting environment above business unit, as drawn here, makes an org-wide policy like "no public IPs in any production project" a single one-line constraint at the top of the production folder; nesting business unit above environment instead means repeating that same constraint once per business-unit folder. Neither nesting is wrong in general, but they are not interchangeable once the organization has more than a couple of business units, so this should be a deliberate choice, not a default. The most common wrong turn is treating folders as a purely cosmetic grouping in the console while doing all real IAM and policy work at the project level, which throws away the entire reason folders exist: bulk, inheritable policy.
Unlock Full Question Bank
Get access to all Google Cloud Platform Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.