Google Cloud Platform Services and Architecture Questions
Google Cloud Platform's core services and architecture: Compute Engine, Cloud Run, GKE, Cloud Storage, VPC, managed databases (Cloud SQL, Spanner, Firestore, Bigtable), and BigQuery-adjacent data services. Covers GCP service selection, networking, IAM and security specifics, cost and quota management, and reference patterns for building on the platform. For provider-agnostic compute, storage, or networking concepts, see the cross-cloud entries.
For hybrid connectivity between on-prem and GCP, compare Cloud VPN, Dedicated Interconnect, and Partner Interconnect. What decision criteria (bandwidth, latency, SLA, security, cost) would lead you to recommend one over the others for a large enterprise customer?
Sample Answer
Direct answer
The three hybrid connectivity options trade bandwidth ceiling, latency and path predictability, whether traffic ever touches the public internet, and how much physical infrastructure access the enterprise needs. For most large enterprises the right answer is a primary path on Dedicated (or Partner) Interconnect for its bandwidth, latency, and cost economics at scale, backed by HA VPN as an independent failover path, rather than picking exactly one and running it alone.
Decision framework
| Option | Bandwidth | Latency and path | SLA (service-level agreement, well-configured) | Security posture | Cost shape |
|---|---|---|---|---|---|
| Cloud VPN (HA VPN) | Roughly 1 to 3 Gbps per tunnel depending on average packet size (capped at 250,000 packets per second per tunnel), scalable by adding more tunnels | Routed over the public internet; round-trip time depends on internet peering paths and is variable, not something GCP configuration controls directly | 99.99% with a properly configured two-tunnel (or more) HA setup | Encrypted (IPsec), but the path itself is the public internet | Low fixed cost, minimal setup lead time |
| Dedicated Interconnect | 10, 100, or 400 Gbps circuits, up to 8 bonded per connection for multiple Tbps aggregate | Direct physical cross-connect; bypasses the public internet entirely, giving a deterministic, generally low and stable round-trip time | 99.99% with 4+ connections across two metros; 99.9% with 2 connections in one metro | Private physical link, not routed over the public internet (add MACsec, a link-layer encryption standard, or layer HA VPN on top, if in-transit encryption is also required) | Higher fixed connection and port cost, but materially lower per-GB egress (data leaving Google's network, billed per GB) cost at real scale |
| Partner Interconnect | Smaller increments than Dedicated, from well below 10 Gbps up to tens of Gbps depending on the partner | Private path via the partner's network; latency depends on the partner's own network quality but avoids the public internet | Comparable SLA tiers to Dedicated when configured redundantly, dependent on the partner | Private, routed through the partner's network rather than the public internet | Fixed cost tied to the partner's pricing, generally a lower entry point than Dedicated |
- What actually decides it for a large enterprise: sustained bandwidth needs above what a handful of VPN tunnels comfortably cover, a hard latency requirement that the variability of public-internet routing can't reliably meet, a requirement (contractual, regulatory, or simply risk-averse) to avoid routing sensitive traffic over the public internet even when encrypted, and existing colocation presence (or budget to acquire it) all point toward Interconnect as the primary path. Whether that's Dedicated or Partner comes down entirely to whether the enterprise has, or is willing to build, a physical presence at a Google-supported facility. Cloud VPN's low setup cost and fast provisioning make it the natural failover layer underneath either interconnect option, and the natural primary option for a smaller or newer connectivity need that doesn't yet justify Interconnect's fixed costs.
Worked example
A large enterprise with sustained multi-gigabit replication traffic between its data center and GCP, and an existing presence in a colocation facility that supports Google Cloud interconnection, provisions a Dedicated Interconnect connection sized with headroom above current sustained usage, for example a 10-Gbps circuit for a workload sustaining 4 to 5 Gbps, leaving room for growth and burst rather than sizing exactly to today's peak, configured with a second connection in a different edge availability domain in the same metro for the 99.9% redundant SLA tier, or across two metros for 99.99% if the business justifies the added cost. HA VPN is configured in parallel as a backup path, advertised to Cloud Router with less-preferred routing metrics so it only carries traffic if the Interconnect path fails, giving the enterprise a tested failover story instead of a single point of failure on the primary path.
Trade-offs and pitfalls
- Choosing Cloud VPN alone as the only connectivity path for a genuinely large, sustained enterprise workload is usually the wrong call once volume grows past a few Gbps sustained: the per-tunnel packet-per-second ceiling means scaling further means adding more tunnels, and public-internet transit, even encrypted, introduces path variability Interconnect avoids entirely.
- Assuming Dedicated Interconnect's private physical link is automatically encrypted in transit is a common misunderstanding; it's private in the sense of not touching the public internet, not encrypted by default, and a compliance requirement for encryption in transit needs an explicit answer such as MACsec or HA VPN layered over the Interconnect connection.
- Sizing a single Interconnect circuit exactly to today's peak usage, with no headroom, leaves no margin for growth, bursts, or the added load during a failover event when all traffic that would normally split across two paths concentrates onto one.
- Running only a primary connectivity path, of any type, without a genuinely independent failover path, different physical infrastructure or a different provider network, not just a second circuit through the same facility, leaves the enterprise exposed to a single point of failure that a redundant-sounding configuration doesn't actually protect against.
Design the HTTP(S) load balancing and Cloud CDN setup for a global web application that serves a mix of static assets from Cloud Storage and a dynamic API from GKE. Cover SSL termination, session affinity, cache key design and invalidation, signed URLs for private assets, and how you would route traffic to the nearest backend while keeping some data in one region for residency reasons.
Sample Answer
Direct answer
One global external Application Load Balancer in front of a URL map that splits traffic by path, a Cloud-CDN-enabled backend bucket for static assets and a backend service pointing at GKE-based (Google Kubernetes Engine) Network Endpoint Groups (NEGs, the load balancer's pointer to a set of backend pods or instances) for the dynamic API, gives a single global anycast entry point (anycast means the same IP address is advertised from many locations at once, so each user's traffic naturally routes to the nearest one), automatic routing to the nearest healthy backend, and a natural place to carve out a region-pinned exception for data-residency needs, without separate infrastructure per region for the common case.
Decision framework and design
- SSL termination: terminate SSL/TLS at the load balancer itself, at Google's global edge, so each user's handshake happens close to them rather than at a single origin region; traffic from the load balancer to GKE backends can run as plain HTTP within Google's network, or re-encrypted, depending on the internal security posture required.
- Cache key design and invalidation for the static path: normalize the cache key to the query parameters that actually matter, prefer content-hashed immutable filenames for assets that change on deploy, and reserve manual path-based invalidation for the rare case something needs to be pulled before its TTL (time-to-live) expires.
- Signed URLs for private assets: for static content that isn't meant to be public, user-uploaded files, gated downloads, keep the backend bucket private and require a signed URL, so the only path to the object is through the load balancer and CDN with a valid signature, never directly against the storage API.
- Session affinity: the dynamic API backend can use session affinity, client-IP-based or a generated cookie, if it genuinely needs it, but the stronger design keeps the API stateless by externalizing session data to a shared store such as Memorystore, since affinity works against even load distribution and complicates autoscaling; reach for it only for something that can't reasonably be made stateless, such as a long-lived WebSocket connection.
- Routing to the nearest backend: the global external Application Load Balancer does this automatically via its own health checks and Google's network, sending each user to their nearest healthy backend with no extra configuration, a meaningful advantage over stitching together several regional load balancers with DNS.
- Data residency for one region: add a distinct URL map rule, by path or by a header identifying the tenant or region, that routes that traffic to its own backend service pointing only at NEGs in the required region, backed by a database and any dependent services also confined to that region. The rest of the app keeps routing globally to the nearest region as normal, keeping the residency requirement scoped to exactly the traffic that needs it.
Worked example
flowchart LR
U[Users worldwide] --> LB[Global external Application Load Balancer]
LB --> CDN[Cloud CDN cache for static path]
CDN --> BB[Backend bucket: Cloud Storage]
LB --> BS[Backend service for api path]
BS --> NEG1[GKE NEG region A]
BS --> NEG2[GKE NEG region B]
LB --> BSEU[Backend service EU tenant only]
BSEU --> NEGEU[GKE NEG EU region data resident]
A user in Frankfurt requests the homepage, and separately an enterprise customer contractually required to keep its data in the EU accesses its tenant-specific dashboard. The homepage's static assets serve from Cloud CDN's European edge cache after the first request from that region warms it, and the homepage's dynamic API calls route through the load balancer's default path to whichever healthy NEG across regions is nearest and healthy, most likely a European GKE cluster, with no per-request configuration needed. The EU tenant's dashboard traffic matches a distinct URL map rule keyed on that tenant, a path prefix or a header set by the frontend, and always routes to the EU-only backend service regardless of where an individual employee of that tenant happens to be connecting from, satisfying the residency requirement even while traveling outside the EU.
Trade-offs and pitfalls
Common misconfigurations to check first when this design doesn't behave as expected:
- Forgetting explicit Cache-Control headers on the dynamic API's own responses is a frequent, dangerous misconfiguration: without them, a caching layer sitting in front of both paths can accidentally cache a personalized or sensitive API response meant for one user and serve it to another; verify the API path is either excluded from CDN caching or explicitly marked non-cacheable.
- Enabling Cloud CDN on the backend bucket but leaving the underlying Cloud Storage bucket publicly readable defeats any signed-URL protection, since the object can still be fetched directly from the storage API, bypassing the load balancer entirely.
- A mismatched or missing health check on the GKE NEG backend is a very common cause of unexplained 502 errors after this kind of setup goes live; verify the health check path and expected response match what the application actually serves.
- Session affinity, once turned on for the API backend, quietly undermines even load distribution and can mask autoscaling problems, a hot instance staying hot because affinity keeps sending it the same clients; avoid it unless there's a specific, tested reason it's required.
- The residency-specific URL map rule is easy to get wrong in a way that fails silently: a routing condition too broad or too narrow either leaks non-EU traffic into the EU-only backend, a capacity and cost problem, or leaks the tenant's traffic into the general backend, a compliance problem; test both matching and non-matching cases explicitly.
Design a secure way for serverless workloads on Cloud Run or Cloud Functions to reach a resource inside a VPC, such as a Cloud SQL instance, while minimizing exposure. What are the options (Serverless VPC Access, Private Service Connect, NAT, peering) and how would you choose between them?
Sample Answer
Direct answer
For Cloud Run or Cloud Functions reaching a Cloud SQL instance or anything else inside a Virtual Private Cloud (VPC), the current default recommendation is Direct VPC egress: it routes the service's outbound traffic straight into your VPC without a separate always-on connector, scales network capacity with the service itself, and is the simplest thing to reason about for a new design. Reach for a Serverless VPC Access connector instead in the one documented case where Direct VPC egress has a known cold-start problem: when the same egress path also goes through Cloud NAT. Google's own Cloud Run documentation calls out cold-start delays of 30 seconds or more on instance startup in that specific combination and recommends a connector with Cloud NAT for better startup performance, rather than a general claim that connectors beat Direct VPC at high request volume. Use Private Service Connect specifically when the destination is a managed service like Cloud SQL or a partner/producer service you want to reach by private IP without exposing it, rather than plain VPC connectivity. Cloud NAT and VPC peering are not choices for this specific problem: they solve the opposite direction (your VPC reaching the internet, or two VPCs reaching each other), not a serverless product reaching into a VPC.
Structured elaboration
Serverless VPC Access (the connector)
- A small, dedicated set of VM instances that sit in your VPC and act as a bridge: Cloud Run or Cloud Functions sends VPC-bound traffic to the connector, which forwards it into the subnet.
- Costs money even when idle (a minimum instance count is always running), and adds a hop, but is a long-established, well-understood path. Google's own guidance does not say the connector generally beats Direct VPC on latency under high load; the specific, documented case is Direct VPC egress paired with Cloud NAT, where Google's own docs report cold-start delays of 30 seconds or more on instance startup and recommend switching to a connector with Cloud NAT for better startup performance instead.
Direct VPC egress
- Cloud Run (and now Cloud Functions deployed as 2nd-generation, which runs on Cloud Run's infrastructure) can attach directly to a VPC subnet without a connector; there is no idle connector cost, and network capacity scales to zero along with the service itself.
- This is the newer, generally recommended default for most workloads precisely because it removes a piece of standing infrastructure you'd otherwise have to size and pay for.
Private Service Connect (PSC)
- Solves a different part of the problem: reaching a managed service, such as a Cloud SQL instance, by a private IP address inside your VPC rather than the instance's default public-or-peered-network path, and without needing VPC peering to Google's managed-services network at all.
- For Cloud SQL specifically, PSC is the modern alternative to the older private-services-access (VPC peering-based) connection method; it avoids the shared, harder-to-reason-about peering range that private-services access consumes from your VPC.
Cloud NAT and peering (why they're not the answer here)
- Cloud NAT lets resources without a public IP reach the internet outbound; it is not a mechanism for a serverless product to reach a private resource, so it plays no role in this specific connectivity problem.
- VPC peering connects two VPC networks to each other; it's what private-services access uses under the hood to reach some managed services, but you would not hand-roll peering yourself for this use case when Direct VPC egress or a connector already exists as the supported path.
Worked example
A backend on Cloud Run needs to reach a Cloud SQL Postgres instance for OLTP (online transaction processing) traffic on every request, at a modest, spiky load. Recommended design: Direct VPC egress on the Cloud Run service pointed at the subnet where Cloud SQL's private IP lives, with the Cloud SQL instance itself reached via Private Service Connect rather than legacy private-services-access peering, so there's no standing connector to pay for and no shared peering range to manage. If this design later needs to reach the public internet from the same subnet via Cloud NAT (not just Cloud SQL over private IP), check for the cold-start delays Google's docs specifically describe for Direct VPC egress combined with Cloud NAT; that documented combination, not a generic assumption about high request volume, is the concrete signal to switch to a Serverless VPC Access connector with a pre-warmed minimum instance count instead, not a default starting point.
Trade-offs & pitfalls
The pitfall is picking the connector by default out of habit (it is the older, more familiar pattern) and paying for idle capacity nobody needed, or conversely defaulting to Direct VPC for an extremely latency-sensitive, high-throughput service without testing cold-start behavior under real load first. Mixing up Private Service Connect (reaching a specific managed service privately) with Serverless VPC Access or Direct VPC (reaching your own VPC generally) is a common source of over-engineered designs that stand up a connector and PSC and peering when only one was needed for the actual destination.
Design secure, low-latency hybrid connectivity between on-prem datacenters and multiple GCP regions. Cover when you'd reach for Dedicated Interconnect, Partner Interconnect, or Cloud VPN, how BGP routing and failover would work, and how you'd limit the blast radius if the on-prem network were compromised.
Sample Answer
Direct answer
Treat hybrid connectivity to multiple regions as a hub-and-spoke problem, not a set of point-to-point links: use Network Connectivity Center (Google's hub product for wiring multiple hybrid connections and VPC networks, that is Virtual Private Clouds, GCP's own private network layer, together with centralized route management) as the hub that on-prem connects into once, with Dedicated or Partner Interconnect (a private physical or partner-provided link between on-prem and Google's network, bypassing the public internet) as the primary spokes and HA VPN as backup, and treat the on-prem side as untrusted by default, limiting what a compromised on-prem network can actually reach through deliberately narrow route advertisement and a firewalled transit boundary, not through the connectivity product's own security.
Decision framework
- Choosing among Dedicated Interconnect, Partner Interconnect, and Cloud VPN follows the same bandwidth, latency, private-path, and colocation-access criteria as a single-region hybrid design, applied per site: a large on-prem location with sustained bandwidth needs and colocation access gets Dedicated Interconnect; a site without colocation access but needing a private path uses Partner Interconnect through a supported provider; Cloud VPN fills in as backup everywhere and as the primary for smaller or lower-priority sites where Interconnect's fixed cost isn't justified.
- Reaching multiple GCP regions from those connections: rather than provisioning a separate physical connection per region, connect on-prem into a hub, Network Connectivity Center, which lets each Interconnect or VPN attachment, a spoke, reach every VPC also attached to the hub, including VPCs in different regions, without a full mesh of direct connections between every site and every region. This centralizes route management in one place instead of duplicating routing policy across many point-to-point links.
- BGP routing and failover: Cloud Router runs BGP (Border Gateway Protocol, the routing protocol used to exchange reachable network ranges) sessions with the on-prem router or routers over each VLAN attachment (Interconnect) or tunnel (VPN), exchanging reachable IP ranges dynamically rather than relying on static routes that need manual updates as the network changes. Configure redundant Cloud Routers and redundant physical paths so no single router or circuit failure drops BGP entirely, and use route priority, MED (a BGP attribute expressing a preference between multiple paths to the same destination), so Interconnect paths are preferred and HA VPN only carries traffic when Interconnect is unavailable. Enabling BFD (Bidirectional Forwarding Detection, a lightweight protocol that detects a link failure faster than waiting for BGP's own hold timer to expire) alongside BGP meaningfully speeds up failover detection compared to BGP alone.
- Limiting blast radius if the on-prem network is compromised: because this is a routed Layer 3 connection, the network path itself forwards whatever traffic the routing tables allow, so the real control is in what gets advertised and what gets firewalled, not in the interconnect or VPN product itself.
- Advertise only the specific destination ranges each spoke actually needs to reach, rather than the default of advertising every subnet in every connected VPC; a compromised on-prem network should not be able to route to GCP subnets it has no legitimate reason to reach.
- Land the Interconnect and VPN attachments in a dedicated transit VPC, the hub, not directly into workload VPCs, and use hierarchical firewall policies, rules applied at the organization or folder level above individual VPC firewall rules, at that transit boundary to allow only specific ports and protocols from the on-prem CIDR range (Classless Inter-Domain Routing, the notation for an IP address range, such as 10.0.0.0/16), denying everything else by default.
- Don't rely on network location as the only trust signal: require service-to-service authentication, mutual TLS or an equivalent identity check, for anything crossing the hybrid boundary, so a packet that does make it through the network path still has to authenticate at the application layer.
- Instrument the transit VPC specifically with VPC Flow Logs and firewall rule logging, watching for anomalous volume or unexpected destinations originating from the on-prem CIDR, the fastest way to detect a compromise actually being exploited through this path rather than finding out from its downstream effects.
- Have a tested circuit-breaker runbook: because the ultimate containment boundary here is the BGP route advertisement and the VLAN attachment itself, know in advance exactly how to withdraw the compromised routes or disable the attachment quickly, rather than working that out for the first time during an actual incident.
Worked example
flowchart LR
ONPREM[On-prem network] --> HUB[NCC hub transit VPC]
ONPREM --> VPNB[HA VPN backup path]
VPNB --> HUB
HUB --> FW[Firewall policy deny by default boundary]
FW --> SPOKE1[Spoke VPC region A workloads]
FW --> SPOKE2[Spoke VPC region B workloads]
FW --> MON[Flow logs and anomaly monitoring]
An enterprise with one large on-prem data center needs low-latency access to workloads split across two GCP regions. Rather than building a separate Interconnect circuit into each region, it provisions redundant Dedicated Interconnect circuits into Network Connectivity Center as the hub, with each region's VPC attached as a spoke. Cloud Router advertises only the specific subnet ranges each region's workloads actually need reachable from on-prem, not the full VPC address space, and a folder-level firewall policy on the hub allows only the specific application ports the on-prem systems legitimately call, denying everything else. When VPC Flow Logs later show an unexpected, high-volume connection attempt from an on-prem host to a database subnet it has no advertised route to reach, the request is dropped at the routing layer before it ever reaches a firewall rule to evaluate, and the flow log entry itself becomes the detection signal that something on the on-prem side is behaving abnormally, triggering the incident runbook rather than a successful lateral-movement attempt.
Trade-offs and pitfalls
- Advertising a full default route set, everything in every VPC, because it's the easiest way to make connectivity just work, is the single biggest blast-radius mistake in this design; it trades away exactly the containment property the question is asking for, in exchange for slightly less route configuration up front.
- Landing Interconnect or VPN attachments directly into workload VPCs instead of a dedicated transit VPC removes the one clean chokepoint where a firewall policy can inspect and restrict all hybrid traffic in one place; retrofitting that separation later is a much larger project than designing it in from the start.
- Treating BGP's dynamic routing as purely a convenience feature, without also configuring MED-based path preference and BFD, leaves failover slower and less predictable than it needs to be, and "it failed over eventually" is a weak answer when a stakeholder asks how long a regional path outage actually degraded traffic.
- A well-designed network boundary without an actual tested incident runbook for withdrawing routes or disabling an attachment is a plan that only exists on paper; the first time anyone tries to execute it should not be during a live compromise.
Explain what Cloud Armor does for a web application and how you'd use it against DDoS and layer-7 attacks. What would you watch for when tuning WAF rules so you don't generate a wave of false positives?
Sample Answer
Direct answer
Cloud Armor sits in front of a web application at the load balancer and does two related but distinct jobs: it absorbs and filters high-volume Layer 7 (application-layer) Distributed Denial of Service, or DDoS, traffic using Adaptive Protection, a machine-learning system trained on the app's own baseline traffic, and it acts as a Web Application Firewall (WAF), matching requests against preconfigured rule sets and custom rules to block things like SQL injection or cross-site scripting attempts before they reach the backend. Tuning it well means never flipping a new rule straight to blocking traffic; every rule goes through a preview period against real traffic first.
Structured elaboration
Against DDoS
- Adaptive Protection continuously baselines what normal traffic to the application looks like, then flags anomalous surges consistent with a Layer 7 DDoS attack and can suggest a specific WAF rule to mitigate it, which a human reviews and applies rather than the system silently blocking traffic on its own.
- Rate-based rules (a "throttle" action) cap requests per client IP or other key within a time window and return a 429 (Too Many Requests) status once exceeded, rather than banning the client outright, which is the right default for protecting backend capacity from a spike without permanently locking out a legitimate but overeager client.
Against Layer 7 attacks (the WAF side)
- Preconfigured WAF rules ship dozens of signatures per category (SQL injection, cross-site scripting, remote code execution, and others), each modeled on the ModSecurity Core Rule Set, giving broad OWASP (Open Worldwide Application Security Project) Top 10 coverage without writing rules by hand.
- Custom rules layer on top for anything specific to the application: named IP allow or deny lists, geographic restrictions, or matching on a header or path unique to the app.
Tuning to avoid false positives
- Never enable a new or preconfigured rule directly in "deny" mode. Cloud Armor supports a preview mode where a rule logs what it would have matched without blocking anything, so the actual hit rate against real production traffic is visible before enforcement.
- Read the preview logs for false positives before switching to deny, and where a specific field (a legitimate JSON payload that happens to resemble an injection pattern, for instance) keeps tripping a signature, use a field-level exclusion for that request attribute rather than disabling the whole rule or the whole preconfigured set.
- Preconfigured rules ship with a sensitivity level, sometimes called a paranoia level, and starting at the lowest sensitivity and raising it deliberately, one review cycle at a time, catches more false positives before they reach production than starting at maximum strictness.
Worked example
An e-commerce checkout API is fronted by an external Application Load Balancer with Cloud Armor attached. The SQL injection preconfigured rule is enabled in preview mode first; over a week, the logs show it repeatedly matching on a legitimate notes field where customers type things like O'Brien's order or paste text containing a single quote and a dash, both of which superficially resemble SQL injection syntax. Rather than disabling SQL injection protection for the whole API, an exclusion is added for that specific field, the rule is left enforced for every other field and endpoint, and it's switched from preview to deny only after a second review shows the exclusion eliminated the false positives without missing any real attack signatures in the same log window.
Trade-offs & pitfalls
Turning on every preconfigured rule at high sensitivity and switching straight to deny is the fastest way to break the application for real customers, not attackers: legitimate input that merely resembles an attack pattern (special characters in free-text fields, for instance) is a common source of false positives, and discovering that only after go-live means an outage caused by your own security tooling. On the DDoS side, relying solely on IP-based rate limiting without Adaptive Protection misses distributed attacks that spread load across thousands of source IPs, each individually under any reasonable per-IP threshold.
Unlock Full Question Bank
Get access to all 10 Google Cloud Platform Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.