Google Cloud Platform Services and Architecture Questions
Google Cloud Platform's core services and architecture: Compute Engine, Cloud Run, GKE, Cloud Storage, VPC, managed databases (Cloud SQL, Spanner, Firestore, Bigtable), and BigQuery-adjacent data services. Covers GCP service selection, networking, IAM and security specifics, cost and quota management, and reference patterns for building on the platform. For provider-agnostic compute, storage, or networking concepts, see the cross-cloud entries.
A customer wants to protect sensitive data in Cloud Storage and BigQuery from exfiltration. What is VPC Service Controls actually doing for them, when would you recommend setting up a service perimeter, and what limitations should they go in expecting?
Sample Answer
Direct answer
VPC Service Controls (VPC-SC) draws a service perimeter around a set of projects and the Google-managed APIs they use (Cloud Storage, BigQuery, and others), and by default blocks any request that would move protected data across that boundary, even if the caller's IAM (Identity and Access Management) permissions would otherwise allow it. Recommend setting up a perimeter whenever the real threat is a legitimate, authenticated identity (an employee with valid credentials, or credentials that leak) exfiltrating data to a personal project or an external account, since that is exactly the class of attack IAM alone cannot stop: IAM answers "is this identity allowed to touch this data," not "is this data allowed to leave this boundary."
Structured elaboration
What it's actually doing
- IAM controls who can call an API; VPC-SC controls where the data such a call touches is allowed to flow, independent of who's asking.
- Concretely, it blocks things like copying a BigQuery table or a Cloud Storage object to a project outside the perimeter, or accessing protected data from an unauthorized network location, even from someone who has full IAM permissions on both sides.
- Every perimeter can run in dry-run mode first: it logs what would have been blocked without actually blocking it, which is how you find out what breaks before you find out the hard way.
When to recommend it
- Data with real exfiltration consequences (customer PII, financial records, anything under a compliance mandate) sitting in Cloud Storage or BigQuery, where the org already trusts IAM to gate access but wants a second, independent control against a valid credential being misused or a config mistake exposing a bucket.
- Not a reflexive recommendation for every project: it adds real setup and ongoing maintenance cost (below), so it earns its place where the data's sensitivity justifies it.
Limitations to set expectations on up front
- VPC-SC does not support every Google Cloud service; an unsupported service inside a protected project can behave unpredictably or simply not work once the perimeter is enforced, so the supported-products list has to be checked against the project's actual footprint before enabling enforcement, not after.
- It is not designed to comprehensively control metadata movement (things like resource names or IAM policy metadata, as opposed to the data itself); IAM and custom roles remain the tool for governing metadata access.
- It only governs traffic to Google-managed APIs; it does not replace VPC firewall rules or network-level controls for traffic between your own VMs.
Worked example
A company keeps customer financial statements in a BigQuery dataset. Today, any employee with BigQuery Data Viewer on that dataset can bq extract or run a query with a destination table in any project their account can reach, including a personal sandbox project, entirely within IAM's rules. Wrapping the dataset's project in a VPC-SC perimeter (started in dry-run mode, monitored for a few weeks against real traffic, then switched to enforced) makes that same export attempt fail with a perimeter violation, unless the destination project is inside the same perimeter or an explicit perimeter bridge has been configured to allow it.
Trade-offs & pitfalls
The most common practical failure is enabling enforcement before checking which services in the protected project are unsupported, which breaks something like a CI/CD pipeline or a third-party integration that reaches a Google API from outside the perimeter, and does so in a way that reads as a mysterious outage rather than an obvious permissions error. Skipping dry-run mode to save time is the single riskiest shortcut here; it trades a controlled, observable rollout for an unplanned one.
A customer with a MySQL deployment sharded across regions wants to migrate to Cloud Spanner to cut operational overhead. Provide a migration plan: schema translation considerations, migration technique (bulk load, change data capture, or dual-write), what changes in the application for distributed transactions and latency, and a pilot and rollback approach.
Sample Answer
Direct answer
Treat the migration as schema translation plus a choice of cutover strategy, run in that order: convert the sharded MySQL schema to Spanner's model first (this is where the real design work is, since Spanner's primary keys determine physical data placement), then move the data with a technique matched to how much downtime the business will tolerate (bulk load for a clean cutover window, change data capture for near-zero downtime), update the application for the transaction and latency differences Spanner introduces, and validate on a pilot before a full cutover with a concrete rollback plan.
Structured elaboration
Schema translation. The open-source Spanner migration tool (formerly known as HarbourBridge, now part of the Cloud Spanner ecosystem) can auto-generate a starting Spanner schema from the existing database's schema; it supports MySQL, PostgreSQL, SQL Server, and Oracle as sources (the broader Spanner migration tooling also covers Cassandra), so the same plan below applies essentially unchanged if the source were a sharded Cloud SQL for PostgreSQL deployment instead of MySQL, with Datastream equally able to stream ongoing PostgreSQL changes for the CDC path described later. The auto-generated result needs deliberate revision, not blind acceptance, regardless of source engine: MySQL's typical auto-incrementing integer primary keys are exactly the pattern that hotspots a single Spanner shard (every new row's key sorts adjacent to the last, so all recent writes land on the same physical split), so primary keys usually need redesigning (a hash prefix, or a naturally well-distributed business key) specifically for Spanner, even where the MySQL schema worked fine as-is. This is also the point to decide which parent-child relationships should become Spanner interleaved tables (physically co-locating related rows, such as an order and its line items, for fast, cheap joins) versus staying as separate, non-interleaved tables.
Migration technique, chosen by downtime tolerance:
- Bulk load, the offline approach (export from MySQL, transform, load into Spanner) is simplest and fastest to build, but requires a maintenance window sized to however long the load and validation take, since the source keeps changing after the export snapshot unless writes are paused for the duration.
- Change data capture, the online approach, using Datastream to stream ongoing inserts/updates/deletes from MySQL, lets you do the bulk historical load once and then keep Spanner continuously caught up with live changes, which is what makes a near-zero-downtime cutover possible: you cut traffic over once the CDC (change data capture) stream shows negligible lag, rather than needing a long pause. Whether offline or online is right is really a question of how much downtime the business will tolerate against how much additional pipeline complexity (a CDC stream to monitor and keep healthy) the team is willing to operate; for a system that can accept a scheduled maintenance window, the offline path is genuinely simpler and has fewer moving parts to get wrong.
- Dual-write (the application writes to both MySQL and Spanner during a transition period) avoids relying on any single replication pipeline, but pushes the burden of keeping both stores consistent onto application code, which is generally more error-prone and harder to reason about than a dedicated CDC pipeline; reserve it for cases where CDC genuinely can't be made to work for the source, not as a default.
What changes in the application. Distributed transactions: a MySQL transaction that happened to touch two rows will, if those rows now live on different Spanner shards, pay a cross-shard coordination cost; the interleaving decisions made during schema translation directly determine how often that happens in practice. Latency: Spanner's per-transaction latency is generally higher than a single-machine MySQL instance's for equivalent transactions, since even a single-shard Spanner write goes through Spanner's replication and consensus path rather than one machine's local commit; application code (and any latency-sensitive SLAs, service-level agreements) needs to be validated against this, not assumed unaffected. Any code that relied on MySQL-specific behavior (auto-increment semantics, specific SQL dialect features, stored procedures) needs explicit rework, since Spanner's SQL dialect and feature set, while broadly ANSI SQL compatible, isn't a drop-in replacement for MySQL-specific extensions.
Pilot and rollback. Migrate one shard (or one low-risk customer segment) first, run it against Spanner in production for a defined observation period while keeping the CDC pipeline running in reverse-compatible mode (or keeping the old MySQL shard warm and unmodified) so a rollback is a traffic-routing change, not a data-recovery emergency, and only proceed to the remaining shards once the pilot's latency, correctness, and cost profile are validated against real production traffic, not just synthetic load tests.
flowchart TD
A[Sharded MySQL] --> B[Schema translation: redesign keys, decide interleaving]
B --> C[Bulk historical load]
A --> D[Datastream CDC]
D --> E[Spanner: continuously caught up]
C --> E
E --> F[Pilot: one shard/segment]
F -->|validated| G[Full cutover]
F -->|issues found| H[Rollback: route traffic back to MySQL shard]
Worked example
Say the sharded MySQL deployment uses an auto-incrementing order_id as the primary key on each shard. Migrated naively to Spanner with the same key, every new order across the entire migrated dataset would insert into the same narrow, ever-increasing key range, concentrating all write traffic on one Spanner split regardless of how many nodes the instance has, the exact hotspot Spanner is otherwise good at avoiding. The corrected schema instead uses a key like hash(customer_id) + order_id (or swaps to a naturally distributed business key such as customer_id as a leading key component, interleaving orders under customers), spreading new writes across the keyspace and letting Spanner's automatic splitting distribute load across nodes as intended, while also making "a customer's orders, interleaved with the customer row" a cheap, co-located read instead of a cross-shard join.
Trade-offs & pitfalls
- Accepting the auto-generated schema from the migration tool without redesigning primary keys for Spanner's physical model is the single most common way these migrations perform far worse than expected after cutover, since the symptom (one hot shard) often doesn't show up clearly until production write volume, not a test load, hits it.
- Dual-write designs that don't carefully handle partial failure (a write succeeding on MySQL but failing on Spanner, or vice versa) silently drift the two stores apart; if dual-write is used, it needs an explicit reconciliation process, not an assumption that both writes always succeed together.
- Skipping the pilot phase to move faster removes the one mechanism (a small blast radius, meaning any mistake stays contained to a narrow, easily-undone slice of the system, paired with a known-good rollback path) that catches a bad interleaving or key design decision before it's load-bearing across the entire migrated dataset.
Design a Zero Trust architecture for an enterprise that spans multiple clouds, using Identity-Aware Proxy, BeyondCorp principles, VPC Service Controls, and Private Service Connect. How would authentication, service-to-service communication, and short-lived credentials work together, and what would a pilot with a single application team look like?
Sample Answer
Direct answer
This is the perimeter-and-identity layer of zero trust, distinct from a service mesh's job of securing service-to-service traffic inside a single compute platform like GKE or Cloud Run. BeyondCorp is Google's underlying philosophy, access decisions based on device and user context rather than network location. Identity-Aware Proxy (IAP) is the enforcement point implementing that philosophy for human access to applications. VPC Service Controls draws a data-exfiltration perimeter around managed GCP services regardless of identity or IAM (Identity and Access Management) permissions. Private Service Connect gives services in different networks, including a different cloud, private connectivity to each other without that traffic ever traversing the public internet.
How the four pieces fit together
BeyondCorp and IAP: human access, replacing the VPN model
BeyondCorp's core claim is that being on the corporate network should grant zero implicit trust, every access decision is made per request based on who the user is, what device they're on, and other context, not on network location. IAP is the practical enforcement point specific to application access, it sits in front of an application and checks identity, device, and context against policy before the request is allowed through, replacing "connect to the VPN, then you're trusted" with a per-request check that works the same whether the user is in an office or on a coffee-shop network.
VPC Service Controls: the data-exfiltration perimeter
Where IAM answers "does this identity have permission," VPC Service Controls answers a different question entirely: "even with valid credentials, can data cross this boundary at all." It creates a perimeter around a set of GCP services such that data cannot be copied out to a resource outside the perimeter, even by someone with legitimate IAM permissions on both sides, which specifically defends against a stolen or overly broad credential being used to exfiltrate data, something IAM policy alone cannot stop since IAM can only say yes or no to an action, not "yes, but only within this boundary."
Private Service Connect: private cross-network connectivity
Lets a service in one network reach a service in another network, a different VPC, a different cloud, via a private internal address rather than the public internet or a full network peering that would expose more of each network to the other than necessary. For a multi-cloud enterprise, this is what lets a service on GCP privately reach a service hosted on another cloud, or vice versa via that cloud's own equivalent feature, without opening the traffic to the public internet.
Authentication, service-to-service, and short-lived credentials together
Human authentication runs through IAP plus the organization's identity provider via single sign-on, enforced per request, with no standing VPN trust. Service-to-service authentication uses Workload Identity for GCP-hosted services and Workload Identity Federation for services on another cloud calling GCP APIs, or that cloud's own equivalent mechanism for the reverse direction, all short-lived and scoped, with no long-lived credential ever copied between environments. Short-lived credentials are the common thread across all of it, whether it's a human's IAP-mediated session, a service's Workload Identity token, or a cross-cloud federated token, none of them remains valid and unrevoked for weeks or months, which is what makes "zero trust" more than a label, a compromised credential in this design has a short, bounded window of usefulness.
Piloting with a single application team
Pick one team with real cross-boundary access needs, calling another team's service, touching a sensitive data store, or serving both internal and external users, a team with a trivial access pattern won't exercise these mechanisms meaningfully. Put IAP in front of just that team's application first to confirm legitimate users aren't locked out, configure a VPC Service Controls perimeter scoped to just that team's sensitive services rather than organization-wide, so a misconfiguration has a contained blast radius (a small, bounded scope of what breaks, not an organization-wide outage), and set up Private Service Connect only for the specific cross-network dependency that team actually has. Measure the pilot on two things: did access break for anyone who should have had it, and can a deliberate attempt to move data across the new perimeter with valid but out-of-scope credentials be shown to fail.
Worked example
A finance application team needs employees to access sensitive financial reports, human access, IAP's job, while the application's backend also needs to privately call an analytics service hosted on a partner's account on another cloud, cross-cloud service-to-service, Private Service Connect's job, and the whole system must not let anyone, even with valid access, copy the underlying financial dataset out to an external storage bucket, VPC Service Controls' job. The pilot stands up IAP in front of the reporting application, a VPC Service Controls perimeter around the project holding the financial dataset, and Private Service Connect for the specific path to the partner's analytics service, all scoped to this one team. A test where a team member with legitimate IAM read access on the dataset attempts to export it to a personal storage bucket outside the perimeter fails as intended, concrete, checkable proof the pilot needs before expanding beyond one team, not just an absence of complaints so far.
Trade-offs and pitfalls
Rolling out VPC Service Controls organization-wide before understanding which legitimate data flows currently cross what will become perimeter boundaries is the most common way such a rollout causes an unplanned outage, some legitimate integration nobody inventoried gets blocked. Treating IAP as sufficient zero-trust coverage on its own is another mistake, IAP only covers human access to the applications it fronts, service-to-service traffic and data-exfiltration risk still need the other two mechanisms. A third pitfall is picking a pilot team with a simple, low-boundary-crossing access pattern because it's easier, that produces a pilot that "succeeds" without ever really testing the parts of the design that matter.
Explain how BigQuery reservations and flex slots work, including when on-demand pricing beats flat-rate. If several teams are sharing a reservation, how would you size it, monitor it, and avoid one team's heavy queries starving the others?
Sample Answer
Direct answer
Reservations let you buy dedicated BigQuery compute (slots) instead of paying per byte scanned. You choose an edition (Standard, Enterprise, or Enterprise Plus), commit slots on a monthly/annual term for a discount, or use Flex commitments that bill per second with no long-term lock-in for short bursts. On-demand still wins when usage is low, spiky, and hard to predict; a reservation wins once your steady-state slot-hour spend is lower than the equivalent per-byte cost of the same workload, and once you value predictable query latency (no queueing behind other customers) over metered simplicity. When several teams share one reservation, size it from peak concurrent demand, give each team its own assignment and baseline (a protected floor), and turn on idle-slot sharing so a quiet team's unused capacity can help a busy one without ever letting one team eat into another's guaranteed floor.
Structured elaboration
Editions and autoscaling. Standard edition is autoscale-only: cheapest per slot, no guaranteed floor, no advanced workload management. Enterprise and Enterprise Plus add a baseline (an always-on floor of slots) plus autoscaling up to a max, and unlock advanced workload management: idle capacity sharing, target job concurrency, and (Enterprise Plus) managed disaster recovery and compliance controls.
Flex slots. A Flex commitment buys slot capacity with a very short minimum window (on the order of a minute) that you can cancel any time after that minimum, billed per second while active, versus a monthly/annual commitment that needs advance notice to unwind. Flex is for short bursts (a quarter-end reporting push, evaluating whether a workload benefits from reserved capacity at all) rather than everyday steady-state use, because it carries the highest per-slot unit price of any option.
On-demand versus flat-rate crossover. On-demand bills a fixed rate per TiB (tebibyte) scanned regardless of how busy your projects are; a reservation bills a fixed rate per slot-hour regardless of how many bytes any given query scans. The crossover is workload-specific: estimate your monthly TiB scanned at the on-demand rate, compare that dollar figure to the fully-loaded cost of enough slots to keep your queries at acceptable latency, and let that arithmetic decide, not a rule of thumb. In general, many small, unpredictable ad hoc queries favor on-demand; a handful of teams running large, recurring batch/ETL (extract-transform-load) and BI (business intelligence) workloads at predictable volume favor a reservation, because performance stops depending on how busy other BigQuery customers are that hour.
Sizing, monitoring, and isolation for multiple teams:
| Lever | What it actually does |
|---|---|
| Reservation size (total slots) | The compute ceiling for everything assigned to it |
| Assignment (project/folder/org) | Routes a team's jobs into a specific reservation, so cost and usage are attributable |
| Baseline vs. autoscale max (Enterprise+) | Guarantees each team a floor no one else can take, and caps how far it can burst |
| Idle slot sharing | Lets a reservation temporarily borrow another reservation's unused slots |
| Target job concurrency | Caps how many jobs one reservation runs at once, trading throughput for consistent per-job latency |
Size the baseline from the P50 (50th percentile) steady demand per team using slot utilization history (INFORMATION_SCHEMA.JOBS slot_ms, or the reservation monitoring dashboards in Cloud Monitoring), and size the autoscale max from the P95 (95th percentile) peak, not the theoretical worst case, since that just buys idle capacity nobody uses. Give each team its own reservation with its own protected baseline, then enable idle-slot sharing between the reservations: this is the actual mechanism that prevents starvation, because a heavy team can only consume other teams' idle surplus, never their guaranteed floor, and the floor snaps back the instant its owner needs it. Alert on sustained saturation of a team's baseline plus autoscale ceiling (a signal it's genuinely undersized) and on jobs sitting in PENDING (a queueing signal that capacity, not query design, is the bottleneck).
Where query optimization fits in. A reservation buys a compute ceiling; it does not fix a wasteful query. A team running unpartitioned full-table scans inside a large reservation is just paying to do the wasteful thing faster. The standing order is to optimize first, the same levers as any BigQuery cost problem (partitioning, clustering, materialized views, column pruning), and size the reservation for the optimized load, otherwise you're locking a fixable inefficiency into a fixed monthly bill.
Worked example
Three teams share one Enterprise-edition reservation: Team A runs BI dashboards with steady daytime concurrency, Team B runs a nightly ETL batch that bursts hard for about 30 minutes, and Team C runs low-volume, unpredictable ad hoc analyst queries. A sane sizing pattern (illustrative, not a specific customer's real numbers): a modest shared baseline sized to cover Team A's steady daytime need plus Team C's floor, since Team B's burst happens overnight when A and C are near-idle and can be borrowed from via idle-slot sharing, with the autoscale max set high enough to absorb Team B's nightly spike. Concretely: monthly reservation cost = (baseline slots x current edition slot-hour rate x 24 hours x 30 days) as the fixed floor, plus metered autoscale usage only during Team B's nightly burst window. That's the correct shape of the calculation (fixed floor plus metered burst); plug in your edition's current slot-hour rate from the console to turn it into a dollar figure, since list prices vary by edition and can change.
Trade-offs & pitfalls
- Sizing to the peak-of-all-peaks instead of the P95 buys idle capacity that just sits there as sunk cost.
- Turning on idle-slot sharing without giving each team its own protected baseline recreates on-demand's noisy-neighbor problem, just inside a paid reservation.
- Standard edition can't do assignment hierarchies, idle sharing, or target concurrency, so moving multiple teams onto a cheap Standard reservation "to save money" can make fairness worse, not better.
- A months-long flat-rate commitment signed before the underlying queries are optimized bakes today's inefficiency into a bill that's much harder to unwind than simply fixing the queries under on-demand.
How would you design secret management for an application that spans GKE, Cloud Run, and Cloud Functions, using Secret Manager and Cloud KMS? Think about access control, rotation, how secrets get into a CI/CD pipeline safely, and how you'd keep them out of logs and container images.
Sample Answer
Direct answer
Keep every secret in Secret Manager as the single source of truth, reach it at runtime through each compute platform's native identity (Workload Identity for GKE, short for Google Kubernetes Engine, or the built-in service identity for Cloud Run and Cloud Functions) rather than long-lived credentials, grant access to individual secrets (not "Secret Manager" broadly) via IAM (Identity and Access Management) scoped per service account, and never let a secret's plaintext value pass through a build step, a CI/CD log, or a container image layer.
Structured elaboration
Access control
- Each secret is its own IAM-protected resource; grant
roles/secretmanager.secretAccessoron the specific secret to the specific service account that needs it, not project-wide access to all secrets. A leaked or over-scoped service account then only exposes the one secret it was meant to use. - On GKE, the workload's Kubernetes (the open-source system that orchestrates groups of containers, which GKE runs as a managed service) service account is bound to a Google service account via Workload Identity Federation for GKE, so a pod authenticates as that identity without any credential file baked into the image or mounted from a Kubernetes Secret.
- Cloud Run and Cloud Functions carry their own runtime service account by default; the same per-secret IAM binding pattern applies directly to that identity.
How secrets actually reach the workload
- Cloud Run and Cloud Functions can reference a Secret Manager secret version directly as an environment variable or a mounted volume in the service or function's own configuration, meaning the secret's plaintext never appears in the deployment manifest, the console, or Cloud Build logs; only a reference (the secret's resource name and version) does.
- On GKE, the equivalent is mounting the secret as a volume via the Secret Manager integration (rather than a native Kubernetes Secret, which is only base64-encoded, not encrypted with a customer-managed key, and is visible to anyone with read access to the Secret object in the cluster), so the actual value is fetched at pod-start time from Secret Manager under the pod's Workload Identity.
Rotation
- Store the current value as a new secret version rather than overwriting the old one, and have the application read a version alias (such as
latest, or better, an explicit "current" alias you control) so a rotation event is a new version plus an alias flip, not a coordinated redeploy across every consumer. - For anything with an external rotation lifecycle (a database password, a third-party API key), automate the rotation itself (a scheduled job that calls the external system's rotation API, writes the new value as a new Secret Manager version, and updates the alias) rather than relying on someone remembering to do it manually.
Getting secrets into CI/CD safely
- The CI/CD system should authenticate to Google Cloud via Workload Identity Federation (letting GitHub Actions, GitLab CI, or similar exchange a short-lived OIDC, or OpenID Connect, token for GCP credentials) instead of holding a long-lived service account key, which removes an entire class of "a CI secret leaked" incident.
- The pipeline fetches secrets it needs (for example, to run integration tests against a real staging credential) directly from Secret Manager at run time, using narrowly scoped, short-lived credentials, and never writes the fetched value into a build artifact, a cached layer, or an exported log.
Keeping secrets out of logs and images
- Never pass a secret as a Dockerfile
ARGorENVbaked at build time; both persist in the image's layer history and are recoverable withdocker historyor by anyone who can pull the image, even if a later layer "removes" the file. - Fetch secrets at container start (an entrypoint script calling Secret Manager, or the platform-native env/volume injection described above) rather than at build time.
- Apply log redaction or structured-logging discipline so a secret value can't end up printed by an unhandled exception or a debug log statement; treat any code path that could log a full request or environment dump as a place a secret could leak until proven otherwise.
Worked example
An application's Postgres password lives as a Secret Manager secret with an alias current pointing at its latest version. The Cloud Run service references that secret and version alias directly in its service configuration as an environment variable, so the value never appears in the Cloud Build log or the Cloud Run console's revision diff, only the reference does. A GKE-hosted worker for the same application reads the same secret via a Secret Manager-backed volume mount under its own Workload Identity binding, scoped to secretAccessor on that one secret only. When the database password rotates, a scheduled job writes a new secret version, updates the current alias, and both the Cloud Run service (on its next cold start or restart) and the GKE workload (on its next scheduled pod restart) pick up the new value without a manual redeploy or a shared coordination step.
Trade-offs & pitfalls
The most common real-world mistake is granting a service account broad secretmanager.secretAccessor at the project level "to keep things simple," which turns a single compromised workload into access to every secret in the project instead of just its own. A close second is treating Kubernetes-native Secrets as sufficient on GKE: they're a convenient API, but without additional encryption and access controls layered on, they don't provide the same audit trail or per-secret IAM granularity that routing through Secret Manager does. Baking a secret into a container image, even briefly during a build stage, is effectively permanent: the image needs to be rebuilt and every copy of it purged, not just the running containers redeployed.
Unlock Full Question Bank
Get access to all Google Cloud Platform Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.