Google Cloud Platform Services and Architecture Questions
Google Cloud Platform's core services and architecture: Compute Engine, Cloud Run, GKE, Cloud Storage, VPC, managed databases (Cloud SQL, Spanner, Firestore, Bigtable), and BigQuery-adjacent data services. Covers GCP service selection, networking, IAM and security specifics, cost and quota management, and reference patterns for building on the platform. For provider-agnostic compute, storage, or networking concepts, see the cross-cloud entries.
You need to propose a migration plan for an on-prem OLTP MySQL database to GCP. What target options would you consider (Cloud SQL, Spanner, or self-managed MySQL on a VM), what migration methods are available, and how would you tailor the recommendation for a typical enterprise-sized database?
Sample Answer
Direct answer
For most on-prem transactional (OLTP, online transaction processing) MySQL migrations, the default target is Cloud SQL for MySQL, because it keeps the same engine, drivers, and application code path and mainly removes the operational burden of running the database yourself. Reach for Spanner only when the actual pain point is horizontal write scale or multi-region strong consistency that a single-writer MySQL topology genuinely cannot provide, and reach for self-managed MySQL on a Compute Engine virtual machine only when there is a specific, named capability (a storage engine plugin, an unsupported extension, a licensing constraint) that Cloud SQL's managed surface cannot accommodate.
What actually decides between the three targets
| Target | Choose when | What you give up |
|---|---|---|
| Cloud SQL for MySQL | Same engine, want managed backups, HA, and patching, vertical scale is enough | Some MySQL flags and plugins aren't exposed; a single-region-primary write ceiling |
| Cloud Spanner | Need horizontal write scale, multi-region strong consistency, a very high service-level agreement (SLA, the uptime guarantee a vendor commits to) tier | Different query dialect and consistency model; effectively an application rewrite, not a drop-in MySQL replacement |
| Self-managed MySQL on Compute Engine | Need a specific engine version, plugin, or extension Cloud SQL doesn't expose | You own patching, backups, HA, and failover yourself, all of it |
Migration methods per target
For a Cloud SQL target, Database Migration Service (DMS) gives continuous, binary-log-based replication for minimal-downtime migration, or a logical dump and restore for smaller, simpler databases where a short maintenance window is acceptable. For a Spanner target, this is a re-platform, not a lift-and-shift, tooling can convert schema and replicate changes, but application code touching joins, transactions, or auto-increment keys needs rework because Spanner's data model and consistency guarantees differ from MySQL's. For a self-managed VM target, DMS's Cloud SQL-specific tooling does not apply the same way, since the destination isn't a Cloud SQL API, this becomes native MySQL replication or a backup-and-restore, the same approach as any on-prem-to-on-prem move.
Tailoring for a typical enterprise-sized database
For an enterprise-sized OLTP database (tens of gigabytes to low terabytes, steady transactional load, application code already written against MySQL, a team that would rather not run database infrastructure), Cloud SQL for MySQL via DMS is the right recommendation in the large majority of cases. The bar for recommending Spanner instead is a demonstrated scaling ceiling, write throughput or dataset growth a single Cloud SQL primary genuinely cannot absorb even with a larger machine type and read replicas, not "it's Google's flagship database." The bar for self-managed is a specific named incompatibility with Cloud SQL, not general discomfort with managed services.
Worked example
A retailer's 800GB on-prem order-processing database runs roughly 2,000 write transactions per minute at peak, which works out to about 33 transactions per second. A right-sized Cloud SQL for MySQL instance with a read replica for reporting comfortably absorbs that load, there is no scaling signal here that justifies Spanner's re-platforming cost. If instead this were a global ledger needing consistent reads across three continents at ten times that write rate, the calculus flips toward Spanner despite the application rewrite, because no single-writer MySQL topology, managed or not, solves that problem.
Trade-offs and pitfalls
A common pitfall is recommending Spanner because it is technically impressive in the abstract while ignoring that migrating to it is a genuinely different data model, that rewrite cost has to be justified by an actual scaling need, not by reputation. A second pitfall is recommending self-managed MySQL on a VM "for control" without naming the specific capability Cloud SQL lacks, this just re-creates the operational burden the customer was trying to escape in the first place. A third pitfall is treating "enterprise" as a synonym for "needs Spanner," database size and the word "enterprise" do not by themselves correlate with a genuine write-scaling ceiling.
Propose a GCP resource organization strategy for an enterprise with multiple business units: when would you use the organization node, folders, projects, billing accounts, and labels? How would you map environments (prod, staging, dev) and shared services like CI/CD and logging, and what are the trade-offs between centralized and decentralized billing?
Sample Answer
Direct answer
Use the organization node as the single root for company-wide policy (org policy constraints and platform-admin IAM), folders to mirror how policy and billing actually need to inherit, and one project per workload per environment, never a shared project across a prod and a staging deployment of the same service. Labels cover everything that is a useful query dimension but not a security or billing boundary, a cost-center or team label, for instance, not a substitute for a folder.
Resource hierarchy
flowchart TD
Org[Organization node] --> FProd[Folder: production]
Org --> FNonProd[Folder: non-production]
Org --> FShared[Folder: shared services]
FProd --> PProdA[Project: bu-a-prod]
FProd --> PProdB[Project: bu-b-prod]
FNonProd --> PStaging[Project: staging]
FNonProd --> PDev[Project: dev]
FShared --> PCicd[Project: cicd]
FShared --> PLogging[Project: logging and monitoring]
An environment-first split, as drawn above, puts prod, non-production, and shared services each in their own top-level folder, with each business unit's projects nested underneath. Shared services, the CI/CD (continuous integration and continuous delivery) pipeline and centralized logging, live in their own folder that every environment's workloads read from or write to through explicit, least-privilege cross-project IAM grants, not by duplicating the pipeline per environment.
The four decisions this actually rests on
Folders are the natural place to draw the IAM boundary: an org policy or IAM role granted at the production folder inherits down to every business unit's prod project automatically, without repeating it project by project. Network isolation should follow the same split: a Shared VPC (a Virtual Private Cloud network shared across multiple projects from one designated host project, so those projects' resources can talk to each other privately) attached per environment, so a compromised staging workload cannot reach a production network path by construction, not by convention. Operational ownership is a separate axis from either of those: a business unit typically owns break/fix for its own prod projects, while a shared-services folder's projects (the CI/CD project, the logging project) are owned by a platform team responsible to every business unit that depends on them, worth naming explicitly since it does not follow automatically from the IAM or network split. Billing separation is the fourth and most consequential decision, covered below.
Worked example
Two business units, Payments and Logistics, each get a prod project and a staging/dev pair under the production and non-production folders respectively (payments-prod, payments-staging, and so on). A single shared-services folder holds one cicd project used by both business units' pipelines and one logging project that every other project's audit and application logs export into, so there is one place, not four, to review activity across the whole organization.
Centralized versus decentralized billing
Centralized billing, one billing account for the whole organization, gives a single invoice and lets committed-use or volume discounts pool across every business unit's spend, at the cost of poor per-team cost attribution unless labels and BigQuery (Google's serverless data warehouse) billing export are used deliberately, and a real concentration of risk in whoever can attach or detach projects from that one billing account. Decentralized billing, one billing account per business unit, gives clean chargeback and contains a runaway bill to the business unit that caused it, at the cost of forfeiting pooled discounts and multiplying the number of billing admins who need auditing. I would default to centralized billing with labels and billing export for chargeback, and only split billing accounts when a business unit is a genuinely separate legal or financial entity, a subsidiary, for example, or a contract specifically requires it.
Trade-offs and pitfalls
Nesting environment above business unit, as drawn here, makes an org-wide policy like "no public IPs in any production project" a single one-line constraint at the top of the production folder; nesting business unit above environment instead means repeating that same constraint once per business-unit folder. Neither nesting is wrong in general, but they are not interchangeable once the organization has more than a couple of business units, so this should be a deliberate choice, not a default. The most common wrong turn is treating folders as a purely cosmetic grouping in the console while doing all real IAM and policy work at the project level, which throws away the entire reason folders exist: bulk, inheritable policy.
A customer is considering Anthos to run GKE both on-prem and on GCP. What are the real benefits and trade-offs here, and what operational considerations, networking, policy, security, licensing, should factor into your recommendation?
Sample Answer
Direct answer
Anthos, the capability Google now markets mostly under the GKE Enterprise umbrella (Anthos is still the name most engineers recognize it by) buys a customer one consistent Kubernetes operating model, API surface, and policy and security tooling across on-prem and GCP. That is genuinely valuable when the customer's real goal is operational consistency across environments they are already committed to running, and a poor fit if the goal is simply "run some workloads on-prem cheaply," because it adds a real subscription cost and operational learning curve on top of running Kubernetes on-prem directly, a cost that only pays for itself when the consistency it provides is something the customer would otherwise have to build and maintain by hand.
Weighing the recommendation
Real benefits
One control-plane experience and API surface whether a cluster runs on GCP, on-prem through VMware or bare metal, or, via attached clusters, on another cloud's Kubernetes, so platform tooling, RBAC (role-based access control, the system that governs which identities can perform which actions) conventions, and upgrade tooling don't fork into "the GCP way" and "the on-prem way." Centralized configuration and policy tooling (Config Sync and Policy Controller under current branding) applies the same kind of guardrails a well-run shared Kubernetes platform needs, namespace quotas, admission policy, RBAC conventions, consistently across every cluster from one source of truth, instead of each environment's platform team reinventing its own version. A service mesh (a dedicated infrastructure layer that handles service-to-service traffic, security, and observability so individual services don't each have to implement it themselves), Cloud Service Mesh in GCP's current branding, works the same way across environments too, extending the mTLS (mutual TLS, a handshake where both sides of a connection present and verify a certificate, not just the server as in ordinary TLS) and traffic-management story from the multi-region GKE design to on-prem clusters, not just GCP-hosted ones.
Real trade-offs
Licensing and cost is a genuine, ongoing line item on top of whatever compute the customer already runs, historically priced per unit of managed infrastructure, that a customer running plain on-prem Kubernetes today doesn't currently pay, and it needs to be weighed against the engineering time it would otherwise take to build equivalent cross-environment consistency by hand. There's a real operational learning curve too, teams fluent in plain on-prem Kubernetes need to learn this platform's specific config-management and fleet-management model, that's a genuine ramp-up cost, not a drop-in replacement for what they already run. Networking and connectivity matter as well, on-prem clusters need reliable, sufficiently low-latency connectivity back to GCP for fleet management and telemetry functions, an unreliable link between the customer's data center and GCP becomes a platform reliability risk that didn't exist when everything ran purely on-prem.
Policy and security as an operational consideration, not just a benefit
The tooling above gives you the enforcement mechanism for consistent policy, but someone still has to author and own those organization-wide policies, and the security team needs to learn this specific tooling rather than whatever policy engine or process they already use on-prem. That adoption cost is real and distinct from the licensing line item, and it deserves its own line in the recommendation rather than being folded silently into "benefits."
Recommendation framing
Recommend this when the customer's actual pain is fragmented policy, security, and tooling across environments they're already committed to running, on-prem for latency, data residency, or existing hardware investment reasons, alongside a larger GCP footprint. Recommend against it when the on-prem footprint is small, temporary, or already well-served by a lightweight Kubernetes setup, in that case the subscription cost buys consistency the customer doesn't have enough surface area to benefit from yet.
Worked example
A manufacturing customer runs latency-sensitive control systems on factory-floor Kubernetes clusters, which must stay on-prem for the life of those factories, alongside a much larger GCP footprint for everything else. Without this platform, their security team maintains two separate policy models and two separate ways of enforcing "no privileged containers" and "images must be scanned," and platform engineers context-switch between two different toolsets. With it, the same admission-control and config-management approach a security-conscious, shared Kubernetes platform needs is defined once and applied to both the factory-floor clusters and the GCP clusters, and the subscription cost is justified against the real alternative, paying platform engineers to build and maintain that consistency by hand across two genuinely different environments indefinitely.
Trade-offs and pitfalls
Recommending this as a default "best practice" for any hybrid setup without sizing the actual on-prem footprint is a common pitfall, a customer with one small on-prem cluster and a short timeline to fully leave on-prem may never recoup the subscription cost before that cluster is decommissioned. Underestimating the connectivity requirement is another, a customer whose data-center-to-GCP link is unreliable or bandwidth-constrained will feel that as platform instability, not just a networking inconvenience, name that risk explicitly before recommending anything, not after. A third pitfall is conflating "we adopted this platform" with "our security posture is automatically better," it gives you the tooling to enforce consistent policy, it does not enforce good policy by itself, someone still has to define the actual rules.
A customer runs a 4 vCPU, 16GB VM at around 20% average CPU utilization, handling 1000 requests per minute with 2 second average response times. Propose a cost-optimized hosting option across Compute Engine, GKE with horizontal pod autoscaling, and Cloud Run, and say what you'd recommend for a mid-sized company and why.
Sample Answer
Direct answer
A 4 vCPU VM sitting at 20% average utilization while still taking 2 seconds to answer a modest 1000 requests a minute is a workload that's mostly waiting, not computing, which points away from paying for a dedicated, always-on VM and toward a platform that only charges for the request-handling it actually does: Cloud Run, sized to share that light CPU load across many concurrent requests per instance.
Decision framework
- Read the numbers before picking a platform. 1000 requests/minute is about 16.7 requests/second (RPS); combined with a 2-second average response time, Little's Law (a queueing-theory relationship: average concurrent work in a system equals arrival rate multiplied by average time each unit spends in the system) says the system typically has around 33 requests in flight at once, not 1000 and not 17. That's a moderate concurrency need, not a high one.
- 20% CPU utilization on a 4 vCPU (virtual CPU core) box means roughly 0.8 vCPU worth of continuous work, spread across those ~33 concurrent requests. That combination, long response time, low CPU use, moderate concurrency, is the signature of an I/O-bound workload: most of each request's 2 seconds is spent waiting on something else (a downstream call, a database query, a lock), not computing.
- Compute Engine here means paying for 4 vCPUs and 16 GB continuously, 24 hours a day, to use about a fifth of the CPU: the least cost-efficient of the three options for this utilization profile, and it only becomes competitive again if traffic is steady enough that a smaller, right-sized VM plus committed use discounts beats the alternatives, worth checking but not the way to bet given the stated 20% figure.
- GKE (Google Kubernetes Engine) with horizontal pod autoscaling (HPA, which adds or removes pod replicas based on observed load) earns its value once there are multiple services or a need for fine-grained scheduling across shared node infrastructure; for one workload sitting mostly idle, it adds Kubernetes' operational overhead, nodes to patch and size, HPA metrics to tune, without a clear efficiency win over a simpler request-based platform.
- Cloud Run fits this shape best: since the workload is I/O-bound rather than CPU-bound, a single instance can safely serve many concurrent requests (Cloud Run lets concurrent requests per instance be configured, appropriate here because those requests are mostly waiting, not competing hard for CPU), so a handful of instances can absorb the ~33 concurrent in-flight requests. Billing follows actual request-handling load rather than a fixed 4 vCPU reservation, directly targeting the waste implied by 20% utilization.
Worked example
L=λ×W=(601000 req/s)×2 s≈33.3 concurrent requests busy vCPU=4×0.20=0.8 vCPU equivalentWith roughly 33 concurrent, mostly-waiting requests needing only about 0.8 vCPU of actual compute between them, a Cloud Run deployment configured for, say, 40 to 80 concurrent requests per instance would need only one to a small handful of instances to cover that load, each billed close to the fraction of a vCPU it actually consumes, instead of one VM reserving all 4 vCPUs regardless of use.
Trade-offs and pitfalls
- This recommendation depends on the workload actually being I/O-bound, which the low CPU utilization strongly suggests but doesn't prove by itself; a workload with occasional CPU-heavy spikes hidden inside a low average needs real profiling before committing to a high concurrency-per-instance setting, since packing too many CPU-bound requests onto one instance would degrade latency instead of improving cost.
- For a mid-sized company specifically, Cloud Run's simplicity, no cluster to run, is worth real money in avoided operational headcount, a genuine factor even when the raw compute-hours math is close between options.
- If traffic later becomes far more predictable and sustained at higher utilization, revisit Compute Engine with committed use discounts; this recommendation is a function of the stated 20% utilization figure and would change if that figure changed materially.
- Session affinity, or any assumption that a given client always lands on the same backend instance, breaks cleanly on Cloud Run's per-request scaling model; sticky in-memory sessions need to move to an external store like Memorystore before this migration works correctly, not after.
Design an Anthos-based hybrid architecture that lets a customer run and migrate stateful workloads between on-prem and GCP with minimal disruption. How would multi-cluster service discovery, centralized policy enforcement, and config sync fit together, and what's the plan for storage persistence during upgrades?
Sample Answer
Direct answer
The hard part of this design isn't the Kubernetes layer, which the fleet model (GKE's mechanism for managing multiple clusters, on-prem and in GCP alike, as one logically grouped unit with shared configuration and policy) handles reasonably directly, it's that stateful workloads carry data that has to physically move or replicate between on-prem and GCP. Minimal disruption means the migration plan has to be built around that data's replication and cutover timeline, with multi-cluster service discovery and centralized config sync handling the "make it look like one fleet" part around the edges of that core data problem.
Designing the hybrid architecture
Multi-cluster service discovery
Fleet-level multi-cluster services let a service in the GCP cluster be discoverable and callable from the on-prem cluster, and vice versa, using the same service name and mesh-level routing regardless of which cluster it's actually running in. This is what lets a stateless caller of the stateful workload migrate independently of migrating the stateful workload itself, the caller doesn't need to know or care which cluster currently hosts the data tier it calls.
Centralized policy enforcement and config sync
Config Sync keeps both the on-prem and GCP clusters' RBAC (role-based access control), admission policy, and namespace conventions identical by continuously reconciling both clusters against the same Git-sourced configuration. A stateful workload migrating from one cluster to the other lands in an environment enforcing exactly the same policy it left, no surprise admission-control rejection or RBAC gap on the new side.
Storage persistence, two distinct problems
This is the crux of the question, and it splits into two related but distinct problems. First, persistence through a routine cluster upgrade, not a migration: stateful workloads rely on PersistentVolumes (the Kubernetes objects that give a pod durable disk storage that survives the pod being rescheduled or restarted), so the storage class (the setting that tells Kubernetes which underlying disk technology and behavior to provision a PersistentVolume from) and underlying backend need to support live migration or reattachment of volumes across nodes during a surge-style upgrade, the same surge-upgrade concept that keeps capacity from dipping during any routine GKE (Google Kubernetes Engine) cluster upgrade, extended to the on-prem cluster too, so a node being upgraded doesn't strand a volume. Second, persistence across a migration from on-prem to GCP, where the two sides genuinely use different storage backends: this needs its own continuous data-replication mechanism, database-level replication for a database workload, the same kind of continuous, monitored replication a cross-region database failover setup relies on, or a storage-level replication tool for other stateful workloads, run continuously ahead of cutover, with a validated cutover step rather than a single-shot copy that risks losing anything written during the copy window.
Sequencing for minimal disruption
Stand up the destination GCP cluster with Config Sync already applying the same policy baseline, establish continuous data replication for the stateful workload from on-prem to GCP, migrate stateless callers first using multi-cluster service discovery so they can reach the still-on-prem data tier transparently, then cut the data tier over last, once replication lag is at zero and validated, mirroring the same cutover discipline any careful database migration needs, but applied to whatever storage technology backs this specific workload.
Worked example
A stateful order-database workload running on-prem needs to move to GCP with minimal disruption. The team stands up the GCP-side GKE cluster under the same fleet, with Config Sync already enforcing identical namespace and admission policy. They configure continuous database-level replication from the on-prem primary to a GCP-hosted replica, following the same discipline any careful database migration needs, continuous replication, monitored lag, a defined point of no return. Meanwhile the stateless API layer reading and writing this database migrates to the GCP cluster first, using multi-cluster service discovery to keep calling the still-on-prem database transparently, no code change required to reach across clusters. Only once replication lag is confirmed at zero does the team promote the GCP-side database, cut the API layer's connection over, and decommission the on-prem database, keeping the real risk concentrated at the single validated moment rather than spread across the whole migration window.
Trade-offs and pitfalls
Treating this as a pure Kubernetes-manifest migration, and only discovering the storage persistence problem once the stateful workload actually needs to move, is the biggest pitfall, the multi-cluster service discovery and config sync pieces are the easy 80 percent, data replication and cutover discipline is the hard 20 percent that actually determines whether this is disruptive. Relying on config sync alone to make the two environments "the same," while skipping validation that the on-prem and GCP storage backends actually behave equivalently for PersistentVolumes, is another, a storage class that behaves differently on each side can silently break assumptions the workload depends on. A third pitfall, familiar from any database migration, is having no defined point of no return for the data cutover, once writes start landing on the GCP side, rolling back to on-prem is a second migration, not an undo.
Unlock Full Question Bank
Get access to all Google Cloud Platform Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.