Multi-Cloud and Hybrid Cloud Architecture Questions
Designing systems that span multiple cloud providers or bridge cloud and on-premises. Covers cloud-agnostic abstraction, workload placement across providers or environments, cross-cloud networking and identity federation, data gravity, infrastructure-as-code and centralized observability that span providers, and the operational cost of avoiding vendor lock-in versus the risk of accepting it. Also covers keeping a system correct once it spans providers: leader election, distributed transactions, rate limiting, and service discovery across cloud or cluster boundaries. Resilience patterns here are scoped to crossing a provider or on-prem/cloud boundary (for example failover from one provider to another, or from on-prem to cloud). Resilience across regions of a single provider, with no second provider or on-prem leg involved, is a different topic (multi-region architecture) and is out of scope here.
Design a multi-cloud Kubernetes deployment strategy for a SaaS product that must run on both AWS and GCP for redundancy and customer choice. Address CI/CD, secrets management, cluster networking, stateful data replication, and how you'd ensure consistent policy and observability across clouds while minimizing divergence in operational workflows.
Sample Answer
Direct answer
I would run a single GitOps-driven deployment pipeline that pushes the same Kubernetes manifests to both EKS (Elastic Kubernetes Service, AWS's managed Kubernetes offering) and GKE (Google Kubernetes Engine, GCP's managed Kubernetes offering), back both clusters with one central secrets manager and one central observability stack rather than each cloud's native equivalent, and accept asynchronous, eventually-consistent stateful data replication between the clusters rather than trying to force synchronous consistency across a cross-cloud link. The design goal is that an engineer deploying a change should not need to think "which cloud is this for," except at the small number of points where a provider difference is unavoidable.
Structured elaboration
CI/CD: build container images once, push to a registry both clusters can pull from (or mirror to a registry in each cloud to avoid cross-cloud pull latency and egress on every deployment), and use a GitOps tool (ArgoCD or Flux) with one Git repository as the single source of truth, syncing the same manifests to both clusters. Provider-specific differences (a storage class name, a load balancer annotation) should live in a small, explicitly-labeled overlay per cluster (Kustomize overlays work well here) rather than scattered conditionals throughout the base manifests, so the divergence is visible and auditable rather than implicit.
Secrets management: use one central secrets manager (HashiCorp Vault, or a cloud-agnostic wrapper over each cloud's native secret store) rather than AWS Secrets Manager for EKS and GCP Secret Manager for GKE as two separate systems, specifically because "consistent policy across clouds" is very hard to audit when the policy lives in two different systems with two different permission models. Both clusters authenticate to the same Vault using their own native workload identity (IRSA for EKS, Workload Identity for GKE) federated to Vault, so secret access is centrally auditable while credential issuance still uses each cloud's short-lived identity mechanism.
Cluster networking: EKS's VPC CNI and GKE's native VPC-native networking both assign pod IPs from the cluster's own network, and the two clusters need a path to reach each other for stateful replication traffic and any direct service-to-service calls: a VPN or dedicated interconnect between the AWS VPC and GCP VPC, with non-overlapping CIDR ranges (checked before either cluster's network is provisioned, not after), and a small, explicit allowlist of what's permitted to cross that link rather than a flat, fully-open peering.
Stateful data replication: for a SaaS product's primary database, run the true source of truth in one cloud (chosen by where the majority of write traffic naturally originates, or by contractual/regulatory pull) with logical replication to a read replica in the other cloud, rather than attempting multi-primary writes across a cross-cloud link, which introduces conflict-resolution complexity that is rarely worth it unless the product specifically needs multi-region active-active writes. This means a full failover of the primary to the other cloud is a real, tested runbook operation (promoting the replica), not something that happens automatically and silently.
Consistent policy and observability: apply the same Kubernetes-native policy engine (OPA Gatekeeper or Kyverno) as an admission controller in both clusters, synced from the same Git-managed policy set, so "no privileged containers" or "all images must be signed" is enforced identically rather than configured twice and drifting. Ship metrics, logs, and traces from both clusters to one central observability backend (rather than CloudWatch for EKS and Cloud Monitoring for GKE as two separate places an engineer has to check), so an incident spanning both clusters can be diagnosed from one dashboard instead of two.
flowchart LR
subgraph AWS["AWS us-east-1"]
EKS["EKS cluster"]
RDS["Aurora Postgres primary"]
end
subgraph GCP["GCP us-central1"]
GKE["GKE cluster"]
CloudSQL["Cloud SQL Postgres replica"]
end
GitOps["Git repo, single source of config"] -->|ArgoCD sync| EKS
GitOps -->|ArgoCD sync| GKE
RDS -->|logical replication| CloudSQL
GlobalLB["Global HTTP load balancer"] --> EKS
GlobalLB --> GKE
Vault["Central secrets manager"] --> EKS
Vault --> GKE
Obs["Central observability stack"] --> EKS
Obs --> GKE
Worked example
A SaaS product needs to run on both AWS and GCP so customers who require a specific cloud for their own compliance reasons can choose either. The Git repository holds base Kubernetes manifests plus two Kustomize overlays (overlays/aws, overlays/gcp) that differ only in storage class and ingress annotation. ArgoCD running in each cluster watches the same repo and applies its own overlay. Aurora Postgres in AWS is the primary; logical replication streams to a Cloud SQL Postgres replica in GCP, giving GCP-hosted customers read access to the same data with a typically sub-second replication lag, while writes for all customers (regardless of which cloud they're pinned to for compute) go to the AWS primary through an internal API rather than directly to the database, keeping the "one writer" invariant simple even though compute is genuinely multi-cloud. Vault, federated to both clusters' workload identities, is the only place secrets are managed; Grafana backed by a central metrics store receives data from both clusters so an on-call engineer sees one unified dashboard regardless of which cloud is having the problem.
Trade-offs and pitfalls
The central trade-off is operational uniformity versus native-cloud convenience: using each cloud's own secrets manager and observability stack would be less work to set up initially (no Vault to operate, no central metrics pipeline to build) but produces exactly the "operational workflow divergence" the question asks to minimize, an engineer or on-call responder now needs fluency in two full toolchains instead of one. The most common pitfall is letting the Kustomize overlays grow beyond a small, well-understood set of genuine provider differences into a place where meaningful application logic quietly diverges between clouds, at which point you no longer have one product deployed twice, you have two products that happen to share a Git repository. A second pitfall is skipping a real, tested failover runbook for the stateful database because "it's just a read replica for now": the day a genuine need arises to promote GCP to primary (an AWS regional outage, for instance), an untested promotion procedure is a bad time to discover an edge case.
Design a global service discovery mechanism for ephemeral Kubernetes workloads running in clusters across multiple clouds and regions. Consider how services find each other, handle latency, deal with stale entries, and maintain security (authentication and authorization) across cluster boundaries.
Sample Answer
Direct answer
I would build global service discovery as a service mesh control plane that federates per-cluster local registries into one global view, rather than trying to run a single flat registry across three clouds: each cluster keeps a fast local registry for its own pods, a lightweight sync layer propagates short-lived leases (not permanent records) between clusters, and every cross-cluster call is authenticated with mutual TLS (mTLS) so cluster boundaries are a trust boundary, not just a network hop.
Structured elaboration
How services find each other: each pod registers itself with a local agent on join and deregisters (or lets its lease expire) on termination; this is the layer that must be fast because it is on every pod's lifecycle path. A separate, slower propagation layer syncs a summarized view (which services exist, their health, and their cluster) to the other clusters' control planes, so a caller in Cluster A's mesh sidecar first checks its local view before querying the global layer, keeping the common case (same-cluster call) cheap and the cross-cluster case (uncommon, but must work) correct.
Latency handling: rank candidate endpoints by measured round-trip latency, not by a static preference list, and default to same-region/same-cluster targets whenever a healthy local instance exists. Ephemeral workloads mean the set of valid endpoints changes constantly, so latency-aware routing has to be a continuous measurement (piggybacked on existing health checks or mesh telemetry), not a one-time configuration.
Handling stale entries: use short TTL (time-to-live) leases rather than permanent registration records; an entry that isn't refreshed within its TTL window is dropped automatically rather than requiring an explicit deregistration message that an abruptly killed pod will never send. Pair this with active health checking at the mesh sidecar level as a second line of defense: even a technically-not-yet-expired entry gets marked unhealthy and skipped if its last few health probes failed, which matters because a TTL long enough to tolerate normal network jitter is also long enough to route a few requests to a pod that already died.
Security across cluster boundaries: every cross-cluster call goes through mTLS with certificates issued by a shared root of trust (a central certificate authority the mesh control plane manages, with per-cluster intermediate certificates so a compromised cluster doesn't require rotating the global root). Authorization is enforced per-service, not per-cluster: a service in Cluster A being allowed to call a service in Cluster B does not imply every service in A can call every service in B, which is the natural but wrong assumption once you have federated the registries.
Worked example
A checkout service running in the AWS cluster needs to call a pricing service that exists in both the GCP and Azure clusters for redundancy. Its local mesh sidecar first checks whether a healthy pricing-service instance exists in the AWS cluster; if not, it queries the federated registry, gets back a latency-ranked list (GCP instance at 45ms measured p50, Azure instance at 130ms measured p50 from this cluster's vantage point), and routes to the GCP instance. Each pricing-service registration carries a 15-second TTL lease refreshed every 5 seconds; if three consecutive refreshes are missed (an abrupt pod kill during a node eviction, for instance), the entry disappears from the federated view within 15 to 20 seconds without any explicit deregistration message. The mTLS certificate presented by the GCP pricing-service instance is validated against the shared mesh root of trust, and a per-service authorization policy explicitly allows checkout.aws-cluster to call pricing.* but does not implicitly allow every other AWS-cluster service to do the same.
flowchart TB
subgraph AWS_Cluster["Cluster A (AWS, us-east-1)"]
SvcA["Service X pods"]
AgentA["Registry agent"]
end
subgraph GCP_Cluster["Cluster B (GCP, us-central1)"]
SvcB["Service X pods"]
AgentB["Registry agent"]
end
subgraph Azure_Cluster["Cluster C (Azure, eastus)"]
SvcC["Service X pods"]
AgentC["Registry agent"]
end
SvcA --> AgentA
SvcB --> AgentB
SvcC --> AgentC
AgentA -->|mTLS heartbeat, TTL lease| GlobalRegistry["Global service registry (multi-cluster mesh control plane)"]
AgentB -->|mTLS heartbeat, TTL lease| GlobalRegistry
AgentC -->|mTLS heartbeat, TTL lease| GlobalRegistry
Client["Caller in Cluster A"] -->|1 resolve Service X| LocalProxy["Local mesh sidecar"]
LocalProxy -->|2 query, prefer local| GlobalRegistry
GlobalRegistry -->|3 endpoint list ranked by latency| LocalProxy
LocalProxy -->|4 mTLS call| SvcB
Trade-offs and pitfalls
The core tension is propagation latency versus registry load: propagating every registration change to every cluster instantly keeps the global view fresh but means a busy cluster's churn floods the sync layer; batching or debouncing propagation reduces load but widens the window where a caller in a remote cluster sees a stale view. A TTL that's too short causes needless re-registration traffic and flapping under normal network jitter; a TTL that's too long means dead endpoints stay discoverable longer than the failure they're meant to route around. The security pitfall that's easy to miss is treating "the clusters trust each other" as equivalent to "every service trusts every other service": once a global registry exists, the temptation is to skip per-service authorization because the network-level trust is already established, which turns a single compromised low-value service into a lateral-movement path across all three clouds.
Design a GitOps-based change management model for hybrid network configurations: define IaC for network resources, pull-request based approvals, automated policy checks (OPA/terraform-compliance), drift detection and remediation, and integration with cloud provider config APIs and on-prem device management. Outline how emergency changes are handled and audited.
Sample Answer
Direct answer
Treat every hybrid network change, cloud-side and on-prem-side, as a pull request against a Git repository
of declared intent, gated by automated policy checks before a human ever reviews it, applied only after
merge, and continuously reconciled against live state afterward so drift is caught within minutes rather than
at the next audit.
Structured elaboration
IaC for network resources. Represent every network object, VPCs/VNets, route tables, firewall rules,
on-prem switch and firewall configuration where the vendor supports a declarative interface, as code in one
repository, so a router ACL (access control list, a set of allow/deny rules) change and a cloud security
group change go through the identical review process instead of two different ones with two different bars
for scrutiny.
Pull-request based approvals. Every change is a PR (pull request) with a diff a reviewer can actually
read: "widen this security group from port 443 to all ports" is visible in the diff itself, not buried in a
console click. Require at least one independent approval before merge, and for genuinely high-risk changes
(anything touching the on-prem-to-cloud boundary itself) require two.
Automated policy checks. Run OPA (Open Policy Agent, a general-purpose policy engine) or
terraform-compliance against every PR before a human reviews it, checking rules like "no security group may
allow inbound 0.0.0.0/0 on a management port" automatically, so reviewers spend their attention on intent and
architecture, not on manually re-deriving policy violations a machine can catch in seconds.
Drift detection and remediation. Run a scheduled job (hourly or more frequent for the boundary itself)
that compares live configuration against the last-applied Git state. A detected drift either auto-reverts
for low-risk, easily reversible changes, or pages a human for anything touching routing or the security
boundary, since auto-reverting a security-relevant emergency change someone made by hand could itself cause
an outage.
Integration with provider config APIs and on-prem device management. The same pipeline that calls cloud
provider APIs (via Terraform providers, for example) also needs a path to on-prem device management,
whether that is a vendor's own API, Ansible against network devices, or a network automation platform;
without that second path, "GitOps for hybrid" quietly becomes "GitOps for the cloud half only," and the
on-prem half stays manually managed and invisible to drift detection.
Emergency changes. Define a break-glass path explicitly: a named, time-boxed exception that allows a
direct change outside the normal PR flow during an active incident, with a mandatory retroactive PR within a
fixed window (for example, one business day) that brings the emergency change back into Git as the source of
truth, and an audit log entry generated automatically the moment break-glass access is used, not manually
written after the fact.
flowchart TD
Dev[Engineer opens PR] --> Policy[OPA / terraform-compliance checks]
Policy -- fail --> Dev
Policy -- pass --> Review[Peer review + approval]
Review --> Merge[Merge to main]
Merge --> Apply[Pipeline applies IaC]
Apply --> CloudAPI[Cloud provider config APIs]
Apply --> OnPremMgmt[On-prem device management system]
Drift[Drift detector, scheduled] --> Compare{Live state matches Git?}
CloudAPI --> Drift
OnPremMgmt --> Drift
Compare -- no --> Remediate[Auto-revert or alert]
Compare -- yes --> Idle[No action]
Worked example
An engineer opens a PR widening an on-prem firewall rule to let a new cloud burst pool reach an internal
API. The OPA policy check fails immediately because the rule as written allows the port from any source, not
just the burst pool's specific CIDR; the engineer narrows the rule, the check passes, a peer reviews and
approves, and the pipeline applies it through the on-prem device management API. Two weeks later, an
on-call engineer makes an emergency, direct change to that same rule at 2 a.m. to resolve an active incident,
using the documented break-glass path. The drift detector flags the divergence from Git within the hour; the
audit log already shows who used break-glass and why, and the engineer files the retroactive PR the next
morning, restoring Git as the single source of truth for that rule.
Trade-offs and pitfalls
The most common failure is building this pipeline for the cloud side only, because cloud provider APIs are
easy to automate against and on-prem device management often is not, which quietly defeats the point:
the boundary itself, the thing this question is actually about, ends up as the one part still managed by
hand. A second pitfall is auto-reverting every detected drift unconditionally: a well-intentioned emergency
fix made outside the pipeline during a live incident can get silently undone by the drift remediator,
turning one incident into two. Scope auto-revert to genuinely low-risk changes and alert-only for anything
near the security boundary.
Cross-cloud deployment strategy: Describe a blue/green or canary deployment pattern that operates across three Kubernetes clusters in different providers. Explain how you would manage traffic shifting, schema migrations, and rollback while avoiding data corruption and maintaining customer experience.
Sample Answer
Direct answer
I would run a canary rollout coordinated by a global traffic-management layer sitting above the three clusters, gate schema migrations to be backward-compatible with both the old and new application version for the entire rollout window (never a hard cutover), and treat the three clusters as independent rollout targets that are promoted one at a time rather than in lockstep, so a bad canary in one cluster never becomes a bad canary in all three simultaneously.
Structured elaboration
Traffic shifting: a global load balancer or service mesh with cross-cluster awareness shifts a small percentage (commonly starting at 1 to 5%) of traffic to the new version in one cluster first, monitors error rate and latency against the baseline for a defined bake time, and only then increases the percentage or promotes to the next cluster. Canary (gradual percentage-based) is generally preferable to blue/green (instant full cutover) for a cross-cluster rollout specifically because it limits blast radius to a small slice of traffic in one cluster at a time, rather than blue/green's all-or-nothing cutover per cluster, which is higher-stakes precisely when you have three independent clusters that could each go wrong differently.
Schema migrations: the single hardest part of this pattern, and the part most rollout designs get wrong. A schema migration must be expand and contract: first deploy a migration that adds the new schema element while the application still works correctly against the old schema (expand phase), deploy the new application version that can read/write both old and new schema shapes, let it bake and roll out fully, then only after the old application version is completely retired everywhere, run a second migration that removes the old schema element (contract phase). Never deploy a migration that breaks the currently-running old version, because during a canary rollout, the old version and new version are both live simultaneously by design, often for hours, and across three separately-deployed clusters that may not all be at exactly the same migration state at the same moment.
Rollback: because schema changes are expand/contract, rolling back the application to the old version is always safe as long as the contract migration hasn't run yet, the old version simply ignores the new schema elements it doesn't know about. This is the actual payoff of the expand/contract discipline: rollback becomes a traffic-shift operation, not a data-recovery operation.
Avoiding data corruption: the risk is a new-version instance and an old-version instance both writing to the same row/document concurrently in incompatible ways. Guard against this with additive-only schema changes during the bake window (new nullable columns, new optional fields, never renaming or repurposing existing ones) and, where the two versions genuinely cannot coexist safely (a field whose meaning is changing, not just its presence), consider dual-writing or a versioned-record approach until the old version is fully retired.
Maintaining customer experience: bake time and automated rollback criteria (error rate or latency crossing a threshold triggers automatic traffic-weight reduction back to zero for the canary, without waiting for a human to notice) matter more here than in a single-cluster canary, because three clusters means three chances for something to go subtly wrong, and a human watching three dashboards simultaneously is a weaker safety net than an automated guardrail per cluster.
Worked example
Rolling out v2 of an API across EKS, GKE, and AKS clusters: the pipeline first deploys v2 to 5% of pods in the EKS cluster only, with the global load balancer weighting 5% of EKS-bound traffic to those pods; a monitoring gate watches error rate and p99 latency (the 99th-percentile response time: the value 99% of requests are faster than, a common way to track worst-case-leaning tail latency instead of just the average) against the v1 baseline for a 15-minute bake window. If the gate passes, weight increases to 25%, then 100% within EKS, at which point the pipeline moves to GKE and repeats the same staged rollout independently, then finally AKS. Throughout this entire process (which might span hours across all three clusters), the database schema is in its expand phase: a new column added in a prior migration is nullable and ignored by v1, populated by v2, so v1 pods still running in GKE and AKS while v2 is fully live in EKS continue working correctly against the same shared database. Only after v2 reaches 100% in all three clusters does a separate, later migration drop the old column that v1 used and v2 has fully superseded.
Trade-offs and pitfalls
The most common and most damaging mistake is a migration that assumes the rollout is instantaneous and fully synchronized across clusters, dropping or renaming a column the moment the "new" version is deployed anywhere, which breaks every old-version pod still serving traffic in the clusters that haven't been promoted yet, exactly the multi-cluster version of the classic single-cluster canary schema mistake, just with a longer exposure window because cross-cluster rollouts take longer end to end. A second pitfall is promoting all three clusters in lockstep to save time; this feels efficient but means a canary issue specific to one cluster (a resource limit misconfigured only in that cluster's manifests, for instance) surfaces simultaneously everywhere instead of being caught and stopped in the first cluster before it ever reaches the other two. The trade-off of doing this safely is time: a fully staged, one-cluster-at-a-time canary with real bake windows is slower than a simultaneous rollout, and that slowness is the actual safety mechanism, not an inefficiency to optimize away.
Design a logging and aggregation strategy for a hybrid environment with on-prem servers, AWS EC2, and GKE clusters. Include log collection agents, transport, central storage, index/search considerations, retention policies, and how you would secure log access for compliance.
Sample Answer
Direct answer
I would standardize on one log shipping agent deployed uniformly across on-prem servers, AWS EC2 (Elastic Compute Cloud, AWS's virtual machine service), and GKE clusters (Fluent Bit is the pragmatic default: lightweight, a native GKE logging path, and a broad set of AWS output plugins), ship everything to one central store rather than three separate per-environment logging stacks, and treat retention and access control as first-class design decisions from the start rather than an afterthought bolted onto whatever the default log store does.
Structured elaboration
Collection agents: run Fluent Bit as a DaemonSet on GKE (one instance per node, standard practice for container log collection) and as a systemd-managed agent, systemd being the standard Linux service manager, on on-prem servers and EC2 instances. Using the same agent everywhere means one configuration language, one set of parsing rules, and one team that actually knows the tool, instead of splitting expertise across CloudWatch Logs agent, a GKE-native path, and a bespoke on-prem syslog setup.
Transport: agents buffer locally (to survive a network blip to the central store) and forward over TLS to a message layer that decouples producers from the store, commonly Kafka or a managed equivalent (Amazon MSK, Confluent Cloud). This buffer matters specifically in the hybrid case: an on-prem link to the cloud is far more likely to have transient connectivity issues than an intra-cloud path, and losing logs during exactly the window an incident is happening defeats the purpose of the system.
Central storage and indexing: an object-storage-backed system (Loki, or Elasticsearch/OpenSearch with a hot/warm/cold tiering strategy) gives you full-text search on recent logs while keeping older data cheap. Design the index/search tier around your actual query pattern: if most searches are "find all logs for this request ID/trace ID in the last 24 hours," an index optimized for structured field lookups outperforms one optimized for full-text search, and vice versa if engineers primarily grep for free-text error strings.
Retention policy: separate operational retention (14 to 30 days of fast, searchable storage for day-to-day debugging) from compliance retention (the longer window, often a year or more, in cheap cold storage that does not need to be instantly searchable). Enforce both with lifecycle rules on the storage tier rather than manual cleanup.
Securing log access for compliance: logs frequently contain sensitive data (customer PII, tokens accidentally logged, internal IPs), so access must be role-based and audited: separate read permissions from the ability to export or delete, redact or hash known-sensitive fields at the collection agent before they ever reach central storage where possible, and keep an immutable audit trail of who queried what, which itself needs its own retention policy to satisfy an actual compliance auditor asking "who looked at customer data in these logs."
Worked example
Concretely: Fluent Bit tails on-prem application logs and EC2 instance logs, running as a systemd service with a local buffer directory sized for roughly two hours of log volume at peak rate (bounding the blast radius of an unexpected network partition to the central store). On GKE, the same Fluent Bit binary runs as a DaemonSet reading container stdout/stderr via the container runtime's log driver. All three paths forward over TLS to a managed Kafka topic partitioned by environment, which a log-processing pipeline consumes into OpenSearch: a hot tier (7 days, fast SSD-backed nodes) for active debugging, and a warm/cold tier (up to 30 days) for slower ad hoc queries, with anything older archived to object storage (S3 or GCS) under a lifecycle rule for a one-year compliance retention window that satisfies a typical SOC 2 (a common third-party security and compliance audit standard) audit trail requirement. Field-level access control in OpenSearch restricts a customer_pii index pattern to a small security-team role, and every query against that pattern is itself logged to a separate append-only audit index.
Trade-offs and pitfalls
A single unified agent and pipeline is more operationally coherent but concentrates risk: if the shared Kafka layer degrades, all three environments lose their logging path simultaneously, so the local buffer on each agent is not optional, it is the thing that keeps a shared-transport outage from becoming a total blind spot during exactly the incident you'd want logs for. The most common pitfall in hybrid logging designs is treating on-prem and cloud as symmetric when the network path is not: on-prem to cloud egress is often the lowest-bandwidth, highest-latency hop in the whole system, so log sampling or local pre-aggregation (counting instead of shipping every debug-level line) is frequently necessary on-prem in a way it is not for intra-cloud traffic. A second common gap is designing retention purely for cost and forgetting the compliance dimension until an auditor asks for logs from eight months ago that were already deleted by a well-intentioned cost-saving lifecycle rule.
Unlock Full Question Bank
Get access to all 13 Multi-Cloud and Hybrid Cloud Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.