Google Cloud Platform Services and Architecture Questions
Google Cloud Platform's core services and architecture: Compute Engine, Cloud Run, GKE, Cloud Storage, VPC, managed databases (Cloud SQL, Spanner, Firestore, Bigtable), and BigQuery-adjacent data services. Covers GCP service selection, networking, IAM and security specifics, cost and quota management, and reference patterns for building on the platform. For provider-agnostic compute, storage, or networking concepts, see the cross-cloud entries.
Pub/Sub offers at-least-once delivery by default. What does that actually mean in practice for a consumer, and how would you design your consumer so that a redelivered message does not cause a duplicate side effect?
Sample Answer
Direct answer
At-least-once means Pub/Sub guarantees your subscriber will see every published message at least one time, but makes no promise against seeing the same message more than once: a slow ack, a subscriber crash between processing and acknowledging, or a lease expiring under load can all trigger a redelivery of a message you already handled. In practice this means every consumer has to be written as if duplicates are a normal, expected input, not an edge case, by making the operation the message triggers idempotent: processing the same message twice must produce the same end state as processing it once.
Structured elaboration
Why redelivery happens at all. A subscriber has a limited window (the ack deadline, configurable up to 600 seconds) to acknowledge a message after receiving it. If your processing takes longer than that window, if the process crashes after doing the work but before sending the ack, or if a network blip loses the ack in flight, Pub/Sub has no way to know the work was actually done, so it does the only safe thing: redeliver. This is a deliberate design choice (favoring "definitely processed, maybe twice" over "maybe never processed"), not a bug to work around.
Designing for idempotency, concretely:
- Deduplicate on a business key, not the Pub/Sub message ID. If the message represents "charge order 12345," check whether order 12345 has already been charged before charging it again, using your own datastore as the source of truth, not Pub/Sub's delivery bookkeeping (which only protects against the platform redelivering the identical message, not a publisher that legitimately sends a second, distinct message for the same logical event).
- Use a unique constraint as the actual enforcement mechanism, not just an application-level check-then-act. A database unique constraint on
order_idfor the charges table (or an equivalent conditional write) turns "check if already processed, then act" into an atomic operation, closing the race condition where two redelivered copies of the same message are processed concurrently by two workers and both pass the check before either writes. - Make the side effect itself naturally idempotent where possible. A
SET status = 'shipped'is idempotent by construction (running it twice leaves the same end state); anINSERTwithout a uniqueness constraint, or anincrement balance by X, is not, and needs the explicit dedup key above. - Layer Pub/Sub's exactly-once delivery feature on top, as defense in depth, not a substitute. It prevents the service from redelivering a message it already recorded as successfully acked, which reduces how often you hit the duplicate path in practice, but it doesn't protect against a publisher-side retry sending a second, distinct message for the same event, so the idempotency design above is still required regardless.
Worked example
An inventory service consumes a "reduce stock by 1" message per sale. A naive consumer runs UPDATE inventory SET stock = stock - 1 WHERE product_id = X, which is not idempotent: two deliveries of the same message reduce stock by 2 for one actual sale. The idempotent version tracks a processed-message table keyed by the order's business ID: INSERT INTO processed_orders (order_id) VALUES (X) ON CONFLICT DO NOTHING, checking the row count affected; only if the insert actually happened (meaning this order hasn't been processed before) does the consumer also run the stock decrement, wrapped in the same transaction so both writes commit or neither does. A redelivered duplicate then hits the ON CONFLICT DO NOTHING branch, skips the decrement entirely, and the stock count stays correct regardless of how many times Pub/Sub redelivers that message.
Trade-offs & pitfalls
- A check-then-act pattern without an atomic uniqueness guarantee (checking "have I seen this order?" in one query, then acting in a separate step) still has a race condition under concurrent redelivery; the dedup check and the side effect need to be atomic together, not sequential.
- Tracking every processed message ID forever grows the dedup table unbounded; scope it with a retention window matched to your maximum plausible redelivery delay (informed by your ack deadline and retry policy), and expire old entries rather than keeping them indefinitely.
- Relying solely on Pub/Sub's exactly-once delivery feature without also handling business-key-level duplicates is a common mistake, since it only closes the platform-redelivery gap, not the "the publisher legitimately sent this twice" gap that a retrying producer, a duplicate webhook trigger, or a replayed batch job can all cause.
A data team runs daily queries over a 5TB table and is complaining about high cost and slow queries. What would you actually recommend, and how would you measure the before-and-after improvement to justify the cost savings to the team?
Sample Answer
Direct answer
Before recommending anything, find out why the queries are slow and expensive: in BigQuery, on-demand cost is driven purely by bytes scanned, and "slow and expensive on a 5TB table" almost always means the query is scanning far more data than the question actually needs (no partition pruning, SELECT * pulling unused columns, no clustering on the filter/join keys). The recommendation is a diagnose-then-fix sequence: pull the job history to see what's actually being scanned, add time-based partitioning and clustering on the columns the team actually filters and joins on, trim SELECT * down to the needed columns, and add a materialized view if the same aggregation runs repeatedly. Only after that, re-evaluate on-demand versus a slot reservation, because sizing a reservation around an unoptimized workload just turns a variable waste into a fixed one. Measure the improvement using BigQuery's own accounting (bytes scanned and slot-milliseconds from INFORMATION_SCHEMA.JOBS), not a stopwatch, because wall-clock time is affected by caching and concurrent load and isn't something you can hand to finance as proof.
Structured elaboration
1. Diagnose before prescribing. Query INFORMATION_SCHEMA.JOBS_BY_PROJECT (or the reservation-scoped equivalent) filtered to this table's queries: look at total_bytes_billed, whether the SQL text uses SELECT *, and whether the WHERE clause could align with a natural partition column (an event date or ingestion timestamp). A bq query --dry_run (or the "this query will process X bytes" estimate shown before you run a query in the console) gives you a free, repeatable baseline number for any specific query without paying for it again.
2. Fixes, ranked by leverage:
- Partitioning on a date or timestamp column the team already filters on (
PARTITION BY DATE(event_ts)). If the daily queries mostly look at recent data, partition pruning turns a full-table scan into a scan of a handful of partitions. - Clustering on the high-cardinality columns used in
WHERE/JOIN(for examplecustomer_id), so BigQuery can skip blocks within a partition too. - Column pruning. BigQuery is columnar and bills by column, not by row, so a
SELECT *against a 40-column table when the report only needs 4 columns scans roughly 10x more data than necessary regardless of partitioning. - Materialized views if the "daily queries" are really the same aggregation run over and over. A materialized view precomputes and incrementally refreshes the result, so the team ends up querying a small precomputed table instead of the 5TB base table.
3. Only then, revisit the pricing model. On-demand bills per TiB (tebibyte, 2^40 bytes) scanned; a reservation (BigQuery editions: Standard, Enterprise, Enterprise Plus) bills per slot-hour regardless of bytes scanned. A reservation only pays off once your optimized, steady-state monthly bytes-scanned bill is higher than the equivalent slot-hour commitment for enough capacity to keep queries fast. Do this math after step 2, not before, or you're sizing a commitment around a wasteful workload.
Worked example
Assume (stated up front so the numbers are reproducible, not measured): the table has 40 roughly-evenly-sized columns, the report currently does SELECT *, spans 400 days of history, and the actual business question only needs 4 columns and the trailing 7 days.
Over a 30-day month, at the published on-demand rate of roughly $6.25 per TiB scanned (confirm the current rate for your region and billing account, since list prices do change):
Before: 5×30=150 TiB/monthAfter: 0.00875×30=0.2625 TiB/month⇒150×6.25≈$937.50/month⇒0.2625×6.25≈$1.64/monthThis is an illustrative calculation from stated assumptions, not a claim about this specific team's table, but it shows the shape of the argument you'd actually make: partitioning and column pruning together can cut scanned bytes (and therefore on-demand cost) by two or more orders of magnitude on a query that's currently scanning everything. To present the real "before and after" to the team, run the exact same query text through --dry_run before and after the schema change and compare the two bytes-scanned numbers directly, which sidesteps result caching (a second identical query might be served from cache for free within 24 hours and would falsely look like the fix worked when it didn't). When you take this to finance rather than just the data team, lead with the bytes-scanned reduction and its dollar equivalent at the current on-demand rate, not a wall-clock speedup: finance can audit a bytes-scanned number against the billing export directly, and that same number is what later tells you whether a slot reservation would ever be worth signing.
Trade-offs & pitfalls
- Partitioning requires either an existing date/timestamp column or falling back to ingestion-time partitioning (
_PARTITIONTIME), which only helps if the team's filters actually align with load time rather than business event time. - Materialized views only help repeated, subsumable aggregation patterns; they do nothing for genuinely ad hoc analyst queries.
- Migrating an existing 5TB table to a partitioned schema usually means a
CREATE TABLE ... PARTITION BY ... AS SELECTrebuild, which itself costs one full scan of the old table (a one-time expense worth calling out explicitly when you propose this), plus updating any views or dashboards pointed at the old table name. - Moving to a flat-rate reservation before the workload is optimized locks in the inefficiency as a fixed monthly bill that's politically harder to walk back than an on-demand bill you can just stop paying by fixing the query.
Compare GCP's load balancer types (global external HTTP(S), regional TCP/SSL proxy, and internal) on latency, session affinity, SSL termination, and cross-region failover. Which would you recommend for a public API, an internal microservice mesh, and a low-latency TCP-based gaming service?
Sample Answer
Direct answer
These three GCP load balancer types trade scope (global vs. regional), protocol layer (application vs. transport), and how much of the connection Google's infrastructure terminates and re-establishes, and the right pick for a public API, an internal microservice mesh, and a low-latency TCP gaming service follows directly from those three axes, not from a single "best" option.
Decision framework
| Load balancer type | Layer | Scope | SSL termination | Session affinity | Cross-region failover |
|---|---|---|---|---|---|
| Global external Application Load Balancer (the current name for what used to be called the HTTP(S) Load Balancer) | 7, application: routes on host, path, headers | Global, one anycast IP, edge-terminated near the user | Terminates SSL/TLS at Google's global edge, close to the client | Supported (client IP or generated cookie), best used sparingly on a stateless backend | Built in: health checks automatically route around an unhealthy region to the next-nearest healthy one |
| Regional external TCP/SSL Proxy Network Load Balancer | 4, transport, but terminates and re-establishes the connection at the proxy, with SSL/TLS offload | Regional | Terminates SSL/TLS at Google's proxy layer, offloading it from the backends | Client-IP-based affinity at the connection level | Not automatic; regional, so multi-region failover needs an external mechanism such as DNS-based routing |
| Internal Load Balancer (Application or passthrough Network, depending on whether Layer 7 routing is needed) | 7 or 4, depending on the variant | Regional, private RFC1918 (the standard defining non-internet-routable private IP ranges, such as 10.0.0.0/8 and 192.168.0.0/16) addresses only | Application variant can terminate SSL/TLS internally; passthrough variant forwards without terminating | Application variant supports header or cookie affinity; passthrough variant relies on consistent connection hashing rather than a true affinity mechanism | Regional only; needs private DNS or a service-mesh layer for multi-region internal traffic |
- Public API: the global external Application Load Balancer, because a public API benefits directly from edge SSL termination close to each client, Layer 7 routing for versioning or multiple services behind one domain, and its built-in, health-check-driven failover across regions, the closest thing to cross-region failover available without building a separate DNS failover strategy.
- Internal microservice mesh: an Internal Application Load Balancer for HTTP-based service-to-service traffic within a VPC, since Layer 7 routing between service versions and the ability to terminate TLS internally both matter more here than raw connection-level performance; a passthrough internal Network Load Balancer is the better fit specifically for internal Layer 4 traffic that isn't HTTP-shaped.
- Low-latency TCP-based gaming service: among the three types this question names, the regional external TCP/SSL Proxy Network Load Balancer is the closest fit, since it's built for TCP traffic and offloads SSL termination from the game servers. Worth flagging honestly: for the most latency-sensitive real-time traffic, a passthrough Network Load Balancer, which forwards packets directly to the backend without terminating and re-establishing the connection at a Google proxy, and preserves the client's source IP, removes an extra hop the proxy variant adds, and is generally the better choice specifically when shaving every millisecond matters more than offloading TLS termination.
Worked example
A public API serving clients in North America, Europe, and Asia gets one global external Application Load Balancer with a single anycast IP; a client in Tokyo has its TLS handshake terminated at a nearby Google edge location rather than round-tripping to wherever the backend runs, and if the nearest healthy backend region becomes unhealthy, the load balancer's own health checks redirect that client's traffic to the next-nearest healthy region automatically, with no DNS change or manual failover step. Contrast that with the regional TCP/SSL Proxy Network Load Balancer serving the gaming service in one region: if that region has a problem, there's no automatic failover to another region built into the load balancer itself, since it's inherently regional; a multi-region gaming deployment needs its own failover strategy, such as DNS-based routing between regional deployments, accepting the DNS TTL as a floor on failover speed.
Trade-offs and pitfalls
- Choosing the global external Application Load Balancer for internal, VPC-only traffic just because it has the most features is a real anti-pattern: it's designed for internet-facing traffic, and internal microservice communication should use the internal variants, keeping traffic on private addressing.
- Assuming any of these load balancers gives automatic cross-region failover is only true for the global external Application Load Balancer; both other types here are regional by design, and treating them as if they failover across regions on their own is a design gap that surfaces exactly when a region has a real problem.
- For the gaming scenario, defaulting to the proxy variant without at least considering the passthrough alternative trades away real latency for TLS offload convenience.
- Session affinity on any of these is a tool for a specific, real requirement, a stateful protocol, a long-lived connection, not a default; turning it on for a stateless HTTP API undermines even load distribution for no benefit.
Design an Anthos-based hybrid architecture that lets a customer run and migrate stateful workloads between on-prem and GCP with minimal disruption. How would multi-cluster service discovery, centralized policy enforcement, and config sync fit together, and what's the plan for storage persistence during upgrades?
Sample Answer
Direct answer
The hard part of this design isn't the Kubernetes layer, which the fleet model (GKE's mechanism for managing multiple clusters, on-prem and in GCP alike, as one logically grouped unit with shared configuration and policy) handles reasonably directly, it's that stateful workloads carry data that has to physically move or replicate between on-prem and GCP. Minimal disruption means the migration plan has to be built around that data's replication and cutover timeline, with multi-cluster service discovery and centralized config sync handling the "make it look like one fleet" part around the edges of that core data problem.
Designing the hybrid architecture
Multi-cluster service discovery
Fleet-level multi-cluster services let a service in the GCP cluster be discoverable and callable from the on-prem cluster, and vice versa, using the same service name and mesh-level routing regardless of which cluster it's actually running in. This is what lets a stateless caller of the stateful workload migrate independently of migrating the stateful workload itself, the caller doesn't need to know or care which cluster currently hosts the data tier it calls.
Centralized policy enforcement and config sync
Config Sync keeps both the on-prem and GCP clusters' RBAC (role-based access control), admission policy, and namespace conventions identical by continuously reconciling both clusters against the same Git-sourced configuration. A stateful workload migrating from one cluster to the other lands in an environment enforcing exactly the same policy it left, no surprise admission-control rejection or RBAC gap on the new side.
Storage persistence, two distinct problems
This is the crux of the question, and it splits into two related but distinct problems. First, persistence through a routine cluster upgrade, not a migration: stateful workloads rely on PersistentVolumes (the Kubernetes objects that give a pod durable disk storage that survives the pod being rescheduled or restarted), so the storage class (the setting that tells Kubernetes which underlying disk technology and behavior to provision a PersistentVolume from) and underlying backend need to support live migration or reattachment of volumes across nodes during a surge-style upgrade, the same surge-upgrade concept that keeps capacity from dipping during any routine GKE (Google Kubernetes Engine) cluster upgrade, extended to the on-prem cluster too, so a node being upgraded doesn't strand a volume. Second, persistence across a migration from on-prem to GCP, where the two sides genuinely use different storage backends: this needs its own continuous data-replication mechanism, database-level replication for a database workload, the same kind of continuous, monitored replication a cross-region database failover setup relies on, or a storage-level replication tool for other stateful workloads, run continuously ahead of cutover, with a validated cutover step rather than a single-shot copy that risks losing anything written during the copy window.
Sequencing for minimal disruption
Stand up the destination GCP cluster with Config Sync already applying the same policy baseline, establish continuous data replication for the stateful workload from on-prem to GCP, migrate stateless callers first using multi-cluster service discovery so they can reach the still-on-prem data tier transparently, then cut the data tier over last, once replication lag is at zero and validated, mirroring the same cutover discipline any careful database migration needs, but applied to whatever storage technology backs this specific workload.
Worked example
A stateful order-database workload running on-prem needs to move to GCP with minimal disruption. The team stands up the GCP-side GKE cluster under the same fleet, with Config Sync already enforcing identical namespace and admission policy. They configure continuous database-level replication from the on-prem primary to a GCP-hosted replica, following the same discipline any careful database migration needs, continuous replication, monitored lag, a defined point of no return. Meanwhile the stateless API layer reading and writing this database migrates to the GCP cluster first, using multi-cluster service discovery to keep calling the still-on-prem database transparently, no code change required to reach across clusters. Only once replication lag is confirmed at zero does the team promote the GCP-side database, cut the API layer's connection over, and decommission the on-prem database, keeping the real risk concentrated at the single validated moment rather than spread across the whole migration window.
Trade-offs and pitfalls
Treating this as a pure Kubernetes-manifest migration, and only discovering the storage persistence problem once the stateful workload actually needs to move, is the biggest pitfall, the multi-cluster service discovery and config sync pieces are the easy 80 percent, data replication and cutover discipline is the hard 20 percent that actually determines whether this is disruptive. Relying on config sync alone to make the two environments "the same," while skipping validation that the on-prem and GCP storage backends actually behave equivalently for PersistentVolumes, is another, a storage class that behaves differently on each side can silently break assumptions the workload depends on. A third pitfall, familiar from any database migration, is having no defined point of no return for the data cutover, once writes start landing on the GCP side, rolling back to on-prem is a second migration, not an undo.
Design a secure, scalable, cost-effective analytics platform for a fintech customer using Pub/Sub, Dataflow, Cloud Storage, BigQuery, and IAM. What would you do about encryption, access control, audit logging, and data lineage to satisfy regulatory constraints like PCI or GDPR?
Sample Answer
Direct answer
Land raw transaction events in Pub/Sub (a managed publish/subscribe service for ingesting streams of events), use Dataflow (a managed service for processing those event streams in real time or in batch) as the one place any PAN (primary account number) or other personal data gets tokenized or masked, before anything downstream ever sees it, keep a minimally-retained raw archive in Cloud Storage for replay and audit only, and serve analysts and dashboards exclusively out of a curated BigQuery layer with row- and column-level security. PCI DSS (Payment Card Industry Data Security Standard) and GDPR (General Data Protection Regulation) both get easier to satisfy through the same architectural choice: minimize where sensitive data physically lands, rather than bolting access control onto every copy of it after the fact.
Pipeline
flowchart LR
SRC[Transaction events] --> PS[Pub/Sub topic]
PS --> DF[Dataflow: validate, mask PAN, enrich]
DF --> GCS[Cloud Storage: raw archive]
DF --> BQ[BigQuery: curated dataset]
BQ --> AN[Analysts and dashboards]
DF -.audit event.-> LOG[Cloud Audit Logs]
BQ -.access event.-> LOG
IAM[IAM least-privilege roles] -.governs.-> DF
IAM -.governs.-> BQ
IAM -.governs.-> GCS
- Encryption. TLS in transit across Pub/Sub, Dataflow worker traffic, and BigQuery; CMEK (customer-managed encryption keys via Cloud KMS, Cloud Key Management Service) on the Cloud Storage archive and on any BigQuery dataset holding data before it has been tokenized.
- Access control. IAM (Identity and Access Management) applied at the dataset and table level in BigQuery, not just at the project level, plus column-level security through BigQuery policy tags so a PAN column stays masked for anyone without an explicit role granting visibility into it. Each pipeline stage runs under its own service account, so a compromised Dataflow job's blast radius (the scope of what an attacker or bug could reach and damage from that single foothold) is limited to what that one stage was ever granted, not the whole pipeline.
- Audit logging. Cloud Audit Logs' Data Access logs enabled on BigQuery and Cloud Storage; PCI DSS requires tracking who accessed cardholder data, and GDPR requires being able to answer who accessed a given data subject's records.
- Data lineage. Tracked at two levels: pipeline-level, the Dataflow job graph and a catalog (Cloud Data Catalog or Dataplex) tagging which BigQuery tables descend from which raw Pub/Sub topics, and record-level, an ingest timestamp and pipeline-version column carried on every row. GDPR's right to erasure and right of access both require finding every downstream copy of one data subject's records, which is only tractable if lineage was tracked from day one rather than reconstructed after a request arrives.
Worked example
Tokenizing in Dataflow, illustratively: a transform reads the raw PAN field, computes a deterministic token (for example via a keyed hash or a tokenization vault lookup) before the record is ever written anywhere durable, and only the token, never the raw PAN, is what lands in Cloud Storage or BigQuery. Anyone querying the curated BigQuery table sees the token; recovering the original PAN requires a separate, tightly-restricted lookup against the vault, which is itself a smaller, more auditable surface than the whole analytics warehouse.
Trade-offs and pitfalls
Tokenizing inside Dataflow, in the stream, versus tokenizing after landing in BigQuery trades pipeline complexity for a materially stronger guarantee: tokenize-after leaves a window, however brief, where raw PAN sits in BigQuery, which is exactly the exposure PCI DSS scoping is trying to eliminate, so tokenize-in-Dataflow is worth the extra code. GDPR's right to erasure is awkward against an append-only analytics warehouse and a replay-based raw archive; if erasure requests are expected at any real volume, design the schema around a pseudonymous identifier from the start, so erasure means deleting one mapping row rather than rewriting historical BigQuery partitions after the fact. The most common wrong turn is treating IAM alone as the compliance control and deferring data lineage until an auditor asks for it, at which point reconstructing lineage across months of Dataflow job history is often only partially possible.
Unlock Full Question Bank
Get access to all Google Cloud Platform Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.