Multi-Cloud and Hybrid Cloud Architecture Questions
Designing systems that span multiple cloud providers or bridge cloud and on-premises. Covers cloud-agnostic abstraction, workload placement across providers or environments, cross-cloud networking and identity federation, data gravity, infrastructure-as-code and centralized observability that span providers, and the operational cost of avoiding vendor lock-in versus the risk of accepting it. Also covers keeping a system correct once it spans providers: leader election, distributed transactions, rate limiting, and service discovery across cloud or cluster boundaries. Resilience patterns here are scoped to crossing a provider or on-prem/cloud boundary (for example failover from one provider to another, or from on-prem to cloud). Resilience across regions of a single provider, with no second provider or on-prem leg involved, is a different topic (multi-region architecture) and is out of scope here.
Explain what SD-WAN is and how it changes the design of hybrid connectivity between branch offices, on-prem data centers, and cloud provider networks. Cover central control vs local forwarding, policy-based path selection, encryption, and common deployment models (managed appliance, virtual edge, overlay vs underlay awareness).
Sample Answer
Direct answer
SD-WAN (Software-Defined Wide Area Network) separates the control plane, deciding which path a packet should take, from the data plane, actually forwarding it, and centralizes that control plane so policy is written once and pushed to every site, instead of every branch router being configured by hand. That single change is what lets a hybrid network use several cheap underlay links (broadband, LTE) intelligently instead of relying on one expensive, purpose-built circuit (MPLS, Multiprotocol Label Switching) for everything.
Structured elaboration
Central control, local forwarding
A central controller, or controller cluster, holds the policy: which application goes on which path, what happens when a path degrades, what gets encrypted. Each site's SD-WAN edge appliance still forwards packets locally, at line rate, using the policy it was last given, so a brief loss of contact with the controller does not stop traffic from moving, it just means the site cannot receive a policy update until contact is restored.
Policy-based path selection
Instead of one static route toward the data center, an SD-WAN edge continuously measures loss, latency, and jitter across every available underlay and steers each application's traffic to whichever path currently meets that application's requirement, re-evaluating in near real time rather than only failing over when a link goes fully down.
Encryption
Traffic is encrypted, typically IPsec (IP Security, a protocol suite that encrypts and authenticates traffic between two endpoints), between SD-WAN edges by default, which is what makes it safe to route real branch traffic over ordinary business broadband instead of a private circuit: the link itself does not need to be trusted if every packet crossing it is encrypted and authenticated.
Deployment models
| Model | What it looks like |
|---|---|
| Managed appliance | A vendor-supplied hardware box at each site, typically bundled with a managed service contract |
| Virtual edge | The same software running as a VM (Virtual Machine) or container, at a branch, inside a cloud VPC/VNet, or both |
| Overlay-underlay awareness | The SD-WAN overlay actively measures the real underlay links beneath it (broadband, LTE, occasionally MPLS) rather than treating the underlay as an opaque, always-available pipe |
Worked example
A branch has two underlays: a business broadband circuit and an LTE backup. Under a purely static design, all traffic uses broadband until it fails completely, then everything moves to LTE. Under SD-WAN, a bulk file transfer stays on broadband, cheap, high bandwidth, latency does not matter much, while a voice call is steered onto whichever link is currently measuring under roughly 30 ms of jitter and 1 percent loss, which might mean voice moves to LTE for a few minutes while broadband is congested, then moves back once broadband's measured jitter drops again, all without a human intervening or the file transfer even noticing.
Trade-offs and pitfalls
SD-WAN does not remove the underlying physics of the internet: a broadband link with genuinely poor upstream connectivity is still a poor link, SD-WAN just detects that faster and reacts to it, it does not make the link better. It also introduces a new dependency, the controller and its reachability, so a design has to consider what a site does during an extended controller outage; in practice, keep forwarding on last-known policy, which is why that property is a requirement, not a nice-to-have, when selecting a platform. Finally, overlay-underlay awareness only helps if the SD-WAN platform is actually measuring the real underlay characteristics continuously; a platform that treats every link as equivalent until it is completely down is not meaningfully different from the static routing SD-WAN is supposed to replace.
Describe redundancy and failover patterns for dedicated cloud interconnects (for example active/active Direct Connect or ExpressRoute, link aggregation groups, dual-provider setups). Explain how you would configure routing (BGP attributes, local-pref, MED, BFD) to ensure fast failover and avoid traffic blackholing during a link event.
Sample Answer
Direct answer
Run two independent physical circuits, ideally from two different providers or at least two different physical paths, active/active, with BGP (Border Gateway Protocol) preferring one path via local-preference under normal conditions, paired with BFD (Bidirectional Forwarding Detection) so a link failure is detected in well under a second instead of waiting for BGP's own, much slower, default timers to notice.
Structured elaboration
Link Aggregation Groups
Where the provider supports it, bond multiple physical connections, for example, two 10 Gbps Direct Connect or ExpressRoute ports, into one LAG (Link Aggregation Group) so a single port or fiber failure does not drop the whole circuit, just its share of bandwidth, and BGP does not even need to react because the LAG absorbed the failure below the routing layer.
Active/active with BGP attributes
- Local-preference: a locally significant attribute deciding which of several outbound paths a network prefers; set it higher on the primary Direct Connect or ExpressRoute path so, under normal conditions, all outbound traffic prefers that path over the backup.
- MED (Multi-Exit Discriminator): tells a neighboring AS (Autonomous System, a routing domain with its own BGP identity) which of several entry points it should prefer for inbound traffic, but it only compares meaningfully between routes learned from the same neighboring AS, a detail that trips people up when they expect it to work like local-preference in reverse.
- BFD: a lightweight, sub-second failure-detection protocol running underneath BGP that tells it immediately when a path is actually dead, instead of relying on BGP's own keepalive and hold timers, commonly 60 seconds and 180 seconds by default on many implementations, which would otherwise leave traffic being sent into a dead path for tens of seconds after it failed.
Dual-provider setups
Using two different interconnect providers, not just two circuits from the same provider, protects against a provider-wide event, a fiber cut affecting all of one provider's capacity in a region, a billing or provisioning error, that two circuits from the same provider would not protect against.
Worked example
| Primary (Direct Connect, provider A) | Backup (a comparable path, provider B) | |
|---|---|---|
| Local-preference | 200 | 100 |
| Effect | Preferred for all outbound traffic under normal conditions | Used only when the primary's local-preference-200 routes are withdrawn |
| BFD interval | Sub-second detection configured | Same |
| With local-preference set to 200 on the primary and 100 on the backup, every router in the local AS prefers the primary path for anything reachable both ways. If BFD detects the primary link is down, it immediately tells BGP to withdraw the local-preference-200 routes, and traffic shifts to the local-preference-100 backup path within the BFD detection interval, typically well under a second, rather than waiting out BGP's default hold timer. |
Trade-offs and pitfalls
The classic blackholing pitfall is a mismatch between the two circuits' configurations: if the backup path is not actually configured to accept and advertise the same prefixes as the primary, a primary failure does not fail over cleanly, it fails silently, because BGP has nothing valid to switch to even though the physical backup link is up. The MED nuance is a real, common mistake: because MED is only compared between routes from the same neighboring AS, using it to influence path preference from two entirely different upstream providers does nothing, local-preference or a different mechanism is needed there instead. Test the failover, not just the configuration: the only way to know a backup path actually works under BGP is to fail the primary over in a maintenance window and watch it happen, not to trust that the configuration looks symmetric on paper.
Explain the concept of data gravity and how it impacts decisions to replicate, move, or leave data in place when designing a multi-cloud or hybrid architecture. Provide two practical examples where data gravity should dominate architectural choices.
Sample Answer
Direct answer
Data gravity is the observation that as a dataset grows larger, it becomes progressively harder and more expensive to move, so applications and services end up migrating toward the data instead of the other way around, the same way a massive object's gravity pulls smaller objects toward it rather than the reverse. In multi-cloud and hybrid design, data gravity means the location of your largest, most actively-used datasets should usually be decided first, and compute placement, replication strategy, and even which cloud you pick for a new service should be decided around that anchor, not independently of it.
Structured elaboration
Why data becomes harder to move as it grows
Three costs scale with data size and none of them scale favorably: egress cost (cloud providers typically charge to move data out, and that charge is per gigabyte, so moving a petabyte costs roughly a thousand times what moving a terabyte costs, not a fixed fee), transfer time (even a fast one-hundred gigabit connection has a real throughput ceiling, so a large enough dataset takes so long to move that the data changes again before the move finishes), and dependent-system coupling (a dataset that has been in one place for years usually has dozens of downstream jobs, indexes, and integrations built assuming it stays there, each of which is its own migration task, not just a data copy).
The practical design consequence
Instead of asking "which cloud should this new service run in," the data-gravity-aware question is "where does the data this service needs already live, and what does it cost in latency and egress fees to reach it from anywhere else." A new analytics service is usually cheaper and faster to build next to the multi-terabyte warehouse it queries than to build wherever is otherwise convenient and pay a per-query cross-cloud or cross-environment data transfer cost indefinitely, since that ongoing cost compounds every day the service runs, unlike a one-time migration cost.
Worked example (two practical cases where data gravity should dominate)
Case 1, a data warehouse anchoring a business intelligence platform. A company has an 800-terabyte data warehouse on one cloud provider, built up over six years of daily loads. A new team wants to adopt a different provider's business intelligence tool that happens to integrate more smoothly with that provider's own data warehouse. Moving the 800 terabytes to the new provider to get that smoother integration is the wrong call. Using an illustrative egress rate of $0.02 per gigabyte (published cross-cloud egress rates vary by provider and region and should be checked against the current price list before committing to a number, but this order of magnitude is representative):
800,000 GB×$0.02/GB=$16,000 for the one-time transfer
That is before accounting for the transfer window (at even a sustained 10 gigabits per second, 800 terabytes takes roughly 800,000 GB×8/10≈640,000 seconds≈7.4 days of continuous transfer) or the re-validation effort on the far end, and every future day of ingestion still needs to update it wherever it actually lives, so a one-time $16,000 estimate is a floor, not the full cost. The data-gravity-correct answer is to bring the business intelligence tool to the warehouse, either by using a connector that queries across clouds or by choosing a tool native to where the warehouse already sits, rather than moving eighty terabyte-years of accumulated data to chase a marginally better tool integration.
Case 2, a machine learning training pipeline next to its source data. A retailer's clickstream data lands continuously in cloud object storage on one provider at multiple terabytes per day. A data science team wants to train models using a managed machine learning platform that happens to be more mature on a different provider. Training there would mean either re-copying multiple terabytes daily (a recurring egress cost that never stops, unlike a one-time migration) or accepting stale, batch-copied data that lags the live clickstream by however long the copy takes. The data-gravity-correct answer is to run training compute in the same provider and region as the clickstream data, even if that provider's machine learning tooling is a notch behind the competitor's, because the recurring cost and staleness of moving a continuously-growing dataset dominates a one-time tooling preference.
Trade-offs & pitfalls
- Data gravity is a strong default, not an absolute rule: a small, slow-growing dataset (gigabytes, updated weekly) does not generate meaningful gravity, and optimizing its placement around "where the data lives" when the data is small is over-engineering.
- The most common mistake is estimating data gravity's cost using only the one-time migration price and ignoring the recurring cost of a permanent cross-environment dependency, which is usually the larger number over any multi-year horizon.
- Data gravity can become an excuse for never modernizing tooling. The right response to a genuinely superior tool on another provider is usually to evaluate whether the tool has an equivalent on the data's current provider, or whether the tool can be pointed at the data remotely, before defaulting to "we can never move."
- Compliance-driven data residency requirements (data that legally cannot move, a separate but related force) reinforce data gravity but are a distinct concern: gravity is about cost and practicality, residency is about legal obligation, and a design can face either, both, or neither.
Design a GitOps-based change management model for hybrid network configurations: define IaC for network resources, pull-request based approvals, automated policy checks (OPA/terraform-compliance), drift detection and remediation, and integration with cloud provider config APIs and on-prem device management. Outline how emergency changes are handled and audited.
Sample Answer
Direct answer
Treat every hybrid network change, cloud-side and on-prem-side, as a pull request against a Git repository
of declared intent, gated by automated policy checks before a human ever reviews it, applied only after
merge, and continuously reconciled against live state afterward so drift is caught within minutes rather than
at the next audit.
Structured elaboration
IaC for network resources. Represent every network object, VPCs/VNets, route tables, firewall rules,
on-prem switch and firewall configuration where the vendor supports a declarative interface, as code in one
repository, so a router ACL (access control list, a set of allow/deny rules) change and a cloud security
group change go through the identical review process instead of two different ones with two different bars
for scrutiny.
Pull-request based approvals. Every change is a PR (pull request) with a diff a reviewer can actually
read: "widen this security group from port 443 to all ports" is visible in the diff itself, not buried in a
console click. Require at least one independent approval before merge, and for genuinely high-risk changes
(anything touching the on-prem-to-cloud boundary itself) require two.
Automated policy checks. Run OPA (Open Policy Agent, a general-purpose policy engine) or
terraform-compliance against every PR before a human reviews it, checking rules like "no security group may
allow inbound 0.0.0.0/0 on a management port" automatically, so reviewers spend their attention on intent and
architecture, not on manually re-deriving policy violations a machine can catch in seconds.
Drift detection and remediation. Run a scheduled job (hourly or more frequent for the boundary itself)
that compares live configuration against the last-applied Git state. A detected drift either auto-reverts
for low-risk, easily reversible changes, or pages a human for anything touching routing or the security
boundary, since auto-reverting a security-relevant emergency change someone made by hand could itself cause
an outage.
Integration with provider config APIs and on-prem device management. The same pipeline that calls cloud
provider APIs (via Terraform providers, for example) also needs a path to on-prem device management,
whether that is a vendor's own API, Ansible against network devices, or a network automation platform;
without that second path, "GitOps for hybrid" quietly becomes "GitOps for the cloud half only," and the
on-prem half stays manually managed and invisible to drift detection.
Emergency changes. Define a break-glass path explicitly: a named, time-boxed exception that allows a
direct change outside the normal PR flow during an active incident, with a mandatory retroactive PR within a
fixed window (for example, one business day) that brings the emergency change back into Git as the source of
truth, and an audit log entry generated automatically the moment break-glass access is used, not manually
written after the fact.
flowchart TD
Dev[Engineer opens PR] --> Policy[OPA / terraform-compliance checks]
Policy -- fail --> Dev
Policy -- pass --> Review[Peer review + approval]
Review --> Merge[Merge to main]
Merge --> Apply[Pipeline applies IaC]
Apply --> CloudAPI[Cloud provider config APIs]
Apply --> OnPremMgmt[On-prem device management system]
Drift[Drift detector, scheduled] --> Compare{Live state matches Git?}
CloudAPI --> Drift
OnPremMgmt --> Drift
Compare -- no --> Remediate[Auto-revert or alert]
Compare -- yes --> Idle[No action]
Worked example
An engineer opens a PR widening an on-prem firewall rule to let a new cloud burst pool reach an internal
API. The OPA policy check fails immediately because the rule as written allows the port from any source, not
just the burst pool's specific CIDR; the engineer narrows the rule, the check passes, a peer reviews and
approves, and the pipeline applies it through the on-prem device management API. Two weeks later, an
on-call engineer makes an emergency, direct change to that same rule at 2 a.m. to resolve an active incident,
using the documented break-glass path. The drift detector flags the divergence from Git within the hour; the
audit log already shows who used break-glass and why, and the engineer files the retroactive PR the next
morning, restoring Git as the single source of truth for that rule.
Trade-offs and pitfalls
The most common failure is building this pipeline for the cloud side only, because cloud provider APIs are
easy to automate against and on-prem device management often is not, which quietly defeats the point:
the boundary itself, the thing this question is actually about, ends up as the one part still managed by
hand. A second pitfall is auto-reverting every detected drift unconditionally: a well-intentioned emergency
fix made outside the pipeline during a live incident can get silently undone by the drift remediator,
turning one incident into two. Scope auto-revert to genuinely low-risk changes and alert-only for anything
near the security boundary.
Design a secure pattern to expose a cloud-hosted service to on-prem customers using PrivateLink (AWS) or Private Endpoint (Azure) so traffic does not traverse the public internet. Include DNS resolution, endpoint types (interface vs gateway), access controls, firewall considerations, and auditing recommendations.
Sample Answer
Direct answer
Expose the service through a provider-managed private connectivity primitive (AWS PrivateLink or Azure Private Link/Private Endpoint) so the consumer's traffic reaches the service over the cloud provider's internal network fabric instead of the public internet: the consumer creates an interface endpoint (an elastic network interface, or ENI, with a private IP address inside their own VPC (AWS's Virtual Private Cloud) or VNet (Azure's Virtual Network)) that forwards traffic to the provider's endpoint service, which in turn fronts a load balancer in the service owner's account. Because the endpoint is just a private IP address inside the consumer's own network, on-prem clients reach it exactly the way they reach any other private resource: over the existing hybrid connection (Direct Connect, AWS's dedicated private circuit into its cloud; ExpressRoute, Azure's equivalent; or VPN), with DNS resolving the service's name to that private IP rather than a public one.
Structured elaboration
flowchart LR
subgraph OnPrem["On-prem"]
U[Corporate user] --> R[On-prem DNS resolver]
end
subgraph ProviderVPC["Consumer VPC"]
R -->|conditional forwarder| PHZ[Private hosted zone]
PHZ --> EP[Interface endpoint / ENI]
end
subgraph ServiceVPC["Provider VPC"]
EP -->|PrivateLink| NLB[Network load balancer]
NLB --> SVC[Service instances]
end
FW[Endpoint security group] -.guards.-> EP
LOG[VPC flow logs + endpoint policy audit] -.records.-> EP
DNS resolution. The service is published under its normal public-looking domain name (for example api.internal.example.com), but a private hosted zone (AWS) or private DNS zone (Azure) overrides that name to resolve to the interface endpoint's private IP for anything querying from inside the VPC/VNet or from on-prem through a configured resolver. On-prem clients need a conditional forwarder pointed at the cloud's inbound DNS resolver endpoint (Route 53 Resolver inbound endpoint on AWS, Azure DNS Private Resolver on Azure) so that queries for that specific domain get forwarded into the cloud's private DNS rather than resolved publicly; every other domain continues resolving normally. Getting this forwarding rule wrong is the most common way this pattern silently fails: the endpoint exists and works, but on-prem clients still resolve the public name and hairpin out to the internet.
Endpoint types, interface versus gateway. An interface endpoint is an ENI with a private IP, works for most AWS services and any customer-hosted service behind a network load balancer, and is reachable from on-prem over the hybrid link because it is a normal routable private IP. A gateway endpoint (AWS-specific, for Amazon Simple Storage Service (S3) and DynamoDB only) is not an IP address at all; it is a route-table entry that keeps traffic to those specific services on the AWS network without ever leaving the VPC's route table, which means it is only reachable from within that VPC (and peered VPCs with the right route propagation), not from on-prem, so it does not apply to the "expose to on-prem customers" requirement in this question at all. For a customer-hosted service, the answer is always an interface endpoint backed by an endpoint service.
Access controls. Two independent layers gate who can reach the service: the endpoint service owner controls which consumer accounts or VPCs are allowed to create an endpoint at all (an allow-list, with or without a manual acceptance step), and the consumer's security group on the interface endpoint's ENI controls which of their own resources can send traffic to it. On the service side, a network load balancer in front of the actual service instances still needs its own security group and, for anything sensitive, an application-layer authorization check; PrivateLink solves network reachability, not authentication.
Firewall considerations. Because traffic now flows over a private IP that looks like any other internal address, existing on-prem firewall rules written for "internal versus internet" traffic need an explicit rule for this new private range, and any deep-packet-inspection appliance in the path needs to know this destination is legitimate rather than flagging a new private-looking destination as anomalous. On the cloud side, if the service instances sit behind additional security appliances, the network load balancer's health checks and the actual data path both need to traverse those appliances consistently, or failover between them will disagree about which path is healthy.
Auditing recommendations. Enable VPC flow logs (or VNet flow logs on Azure) on the subnet containing the interface endpoint, which records every connection attempt to and from the endpoint's private IP regardless of application-level logging. Combine that with the endpoint service's own connection acceptance log (which consumer accounts/VPCs are attached, and when) and the load balancer's access log on the service side, so an auditor can trace a specific request from "which on-prem host initiated it" through "which endpoint accepted it" to "which backend instance served it" without gaps.
Worked example
An on-prem trading-reconciliation team needs to call an internal pricing API hosted in the cloud, without that traffic ever touching the public internet. The service owner creates a network load balancer in front of the API's instances and publishes it as an endpoint service, allow-listing only the consumer's VPC account. The consumer creates an interface endpoint in their VPC, which allocates a private IP, say 10.20.3.45, in a dedicated endpoint subnet. A private hosted zone maps pricing-api.internal.example.com to that IP for anything resolving inside the VPC. On the on-prem side, the existing DNS server gets a conditional forwarder: any query for internal.example.com is forwarded to the cloud's inbound DNS resolver endpoint over the existing Direct Connect link, instead of being resolved publicly. An on-prem host runs nslookup pricing-api.internal.example.com and gets back 10.20.3.45, its request traverses the Direct Connect link and lands on the interface endpoint's ENI (allowed by that ENI's security group, which only permits the on-prem CIDR range (a block of IP addresses written in Classless Inter-Domain Routing notation, such as 10.20.0.0/16) on the API's port), forwards to the network load balancer, and reaches a backend instance. VPC flow logs record the connection at the endpoint, and the load balancer's access log records which backend served it, giving a complete, auditable path with no public IP address anywhere in it.
Trade-offs & pitfalls
- Forgetting the on-prem conditional-forwarder step is the single most common failure: the private endpoint and DNS zone are both configured correctly, but on-prem clients still resolve the public name because nothing told the on-prem resolver to forward that specific domain into the cloud.
- Gateway endpoints do not solve this problem for on-prem access at all; reaching for one because "it's free and simpler" for S3/DynamoDB access from inside the VPC is a different use case than the on-prem exposure this question asks about.
- An interface endpoint is billed per hour plus per-GB of data processed in every consumer VPC that attaches one, so a design with many consuming teams each creating their own endpoint multiplies that cost; a shared, centrally-managed endpoint (fronting a transit VPC or hub) is usually cheaper at scale than one endpoint per team.
- PrivateLink secures the network path, not the request itself: pairing it with a strong endpoint policy and application-level authentication/authorization is still necessary, since a private IP is not, by itself, an access-control decision.
Unlock Full Question Bank
Get access to all Multi-Cloud and Hybrid Cloud Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.