Multi-Cloud and Hybrid Cloud Architecture Questions
Designing systems that span multiple cloud providers or bridge cloud and on-premises. Covers cloud-agnostic abstraction, workload placement across providers or environments, cross-cloud networking and identity federation, data gravity, infrastructure-as-code and centralized observability that span providers, and the operational cost of avoiding vendor lock-in versus the risk of accepting it. Also covers keeping a system correct once it spans providers: leader election, distributed transactions, rate limiting, and service discovery across cloud or cluster boundaries. Resilience patterns here are scoped to crossing a provider or on-prem/cloud boundary (for example failover from one provider to another, or from on-prem to cloud). Resilience across regions of a single provider, with no second provider or on-prem leg involved, is a different topic (multi-region architecture) and is out of scope here.
Write a Python script outline (pseudocode acceptable) that listens to BGP route updates via ExaBGP, validates incoming prefixes against an allowed-prefix list, and applies safe updates to cloud route tables via AWS/GCP APIs. Include idempotency handling, concurrency controls, rate limiting, and safety checks to prevent accidental route hijacks.
Sample Answer
Direct answer
Structure the listener as three separated concerns: a thin ExaBGP (an open-source BGP implementation that exposes route updates as simple text lines over stdin/stdout, commonly used for exactly this kind of programmatic route control) input loop, a pure validation function that checks every incoming prefix against an explicit allow-list and length bounds before anything else happens, and an idempotent, rate-limited apply step that is the only code allowed to touch the actual cloud route table, so a malformed or malicious announcement gets rejected before it ever reaches an AWS or GCP API call.
Approach
Keep validation as a pure function with no side effects, so it can be unit-tested and reasoned about independently of the network and API calls around it: reject anything outside an explicit prefix-length range, this blocks both an accidental default-route announcement and an overly specific announcement that could be a deaggregation attack, reject anything not contained within an explicit list of allowed aggregate prefixes, and only then hand a validated update to an idempotent apply function that tracks last-applied state per prefix, so a duplicate announcement is a safe no-op, behind a lock, so concurrent updates cannot race each other, and a simple rate limiter, so a flapping or compromised peer cannot hammer the cloud API into its own rate limits.
import ipaddress
import threading
import time
ALLOWED_PREFIXES = [
ipaddress.ip_network("10.10.0.0/16"),
ipaddress.ip_network("10.20.0.0/16"),
]
MIN_PREFIX_LEN = 8 # reject anything broader than this (blocks default-route hijacks)
MAX_PREFIX_LEN = 24 # reject anything more specific than this (blocks deaggregation floods)
RATE_LIMIT_PER_MIN = 30
_seen = {} # idempotency: prefix -> last-applied state
_lock = threading.Lock()
_rate_window = [] # sliding-window rate limiter
def is_prefix_allowed(prefix_str):
try:
net = ipaddress.ip_network(prefix_str, strict=True)
except ValueError:
return False, "unparseable prefix"
if net.prefixlen < MIN_PREFIX_LEN or net.prefixlen > MAX_PREFIX_LEN:
return False, f"prefix length {net.prefixlen} outside allowed range [{MIN_PREFIX_LEN},{MAX_PREFIX_LEN}]"
for allowed in ALLOWED_PREFIXES:
try:
if net.subnet_of(allowed):
return True, "ok"
except TypeError:
continue # mixed IPv4/IPv6 comparison: never matches, never crashes
return False, "prefix not contained in an allowed aggregate"
def rate_limit_ok():
now = time.time()
with _lock:
while _rate_window and now - _rate_window[0] > 60:
_rate_window.pop(0)
if len(_rate_window) >= RATE_LIMIT_PER_MIN:
return False
_rate_window.append(now)
return True
def apply_route_update(prefix_str, next_hop, withdrawn=False):
"""The only function allowed to touch the real route table. Here it is
stubbed to just record state, since no cloud credentials exist in this
environment; in production the APPLIED branch calls, for example, the
AWS ec2 client's replace_route or the GCP compute routes API."""
allowed, reason = is_prefix_allowed(prefix_str)
if not allowed:
return False, f"REJECTED {prefix_str}: {reason}"
if not rate_limit_ok():
return False, f"THROTTLED {prefix_str}: rate limit exceeded"
with _lock:
state = "withdrawn" if withdrawn else next_hop
if _seen.get(prefix_str) == state:
return True, f"NOOP {prefix_str}: already in state {state}"
_seen[prefix_str] = state
# cloud API call goes here: AWS replace_route / GCP routes.patch
return True, f"APPLIED {prefix_str} -> {state}"
def handle_exabgp_line(line):
"""ExaBGP's process API emits simple space-separated lines like
'announce 10.10.5.0/24 192.0.2.1' or 'withdraw ...' on stdin."""
parts = line.split()
if len(parts) < 3:
return None
action, prefix, next_hop = parts[0], parts[1], parts[2]
return apply_route_update(prefix, next_hop, withdrawn=(action == "withdraw"))
if __name__ == "__main__":
demo_updates = [
"announce 10.10.5.0/24 192.0.2.1",
"announce 10.10.5.0/24 192.0.2.1",
"announce 0.0.0.0/0 192.0.2.1",
"announce 172.16.0.0/16 192.0.2.1",
"announce 2001::/16 192.0.2.1",
"withdraw 10.10.5.0/24 192.0.2.1",
]
for line in demo_updates:
print(handle_exabgp_line(line))
Running this exactly as shown prints:
(True, 'APPLIED 10.10.5.0/24 -> 192.0.2.1')
(True, 'NOOP 10.10.5.0/24: already in state 192.0.2.1')
(False, 'REJECTED 0.0.0.0/0: prefix length 0 outside allowed range [8,24]')
(False, 'REJECTED 172.16.0.0/16: prefix not contained in an allowed aggregate')
(False, 'REJECTED 2001::/16: prefix not contained in an allowed aggregate')
(True, 'APPLIED 10.10.5.0/24 -> withdrawn')
Key points
- Validation, rate limiting, and the actual route-table mutation are three separate functions specifically so validation logic can be tested and reasoned about without a network or cloud SDK anywhere near it.
- The
subnet_ofcontainment check is wrapped in atry/except TypeErrorbecause comparing an IPv6 network against an IPv4 aggregate raisesTypeErrorin Python'sipaddressmodule rather than returningFalse; without that guard, a single stray IPv6 announcement, shown above as2001::/16, would crash the whole listener, an availability bug caused directly by the safety check meant to prevent one. - Idempotency is keyed on whether a prefix already maps to this exact state, not whether the prefix has been seen before, so a legitimate re-announcement of the same route, which ExaBGP peers do periodically, not just on change, is a safe no-op rather than a repeated, wasted API call.
- The rate limiter is a blunt, global instrument on purpose: a route table getting hammered by dozens of updates per minute is itself a signal worth throttling and alerting on, independent of whether any individual update would otherwise have been valid.
Complexity
Validation is O(k) per update, where k is the small, fixed number of allowed aggregate prefixes, since it checks containment against each one in turn. The idempotency and rate-limit paths are O(1) amortized per update, a dict lookup and a bounded sliding-window trim. Memory is O(p), where p is the number of distinct prefixes ever seen, since _seen never evicts old entries; a long-running production version would need a bound or a TTL (Time-To-Live) on that dictionary specifically to avoid unbounded growth from a peer announcing many distinct short-lived prefixes over time, which the outline as shown does not yet handle.
Edge cases
- Mixed address families: an IPv6 announcement against an IPv4-only allow-list is caught by the
TypeErrorguard and rejected cleanly rather than crashing, verified above with the2001::/16line. - A default-route announcement (
0.0.0.0/0), the classic route-hijack shape, is rejected by the minimum-prefix-length check before it ever reaches the aggregate-containment check. - A duplicate announcement of an already-applied prefix is a NOOP, not a repeated cloud API call, protecting both API rate limits and the idempotency of the underlying route table.
- The demo's
_seendictionary and_rate_windowlist are process-local and not thread-safe across multiple listener processes; a production deployment running more than one instance for high availability needs that state in a shared store, such as Redis, or the cloud's own route table treated as the source of truth via a read-before-write, instead of in-process memory, or two instances can each believe they are the first to apply a given prefix.
Design a Terraform module pattern for a reusable, versioned multi-cloud transit network that supports AWS Transit Gateway, Azure Virtual WAN, and GCP Network Connectivity Center. Describe the module input and output interface, optional features (VPN fallback, NAT, inspection), how you would handle provider differences, versioning strategy, and testing in CI/CD.
Sample Answer
Direct answer
Build one consistent input/output interface (variable and output names that mean the same thing regardless of which cloud implements them) and then a separate submodule per provider behind it, rather than one module trying to branch internally on cloud provider, because Transit Gateway (AWS's hub-and-spoke network resource for connecting multiple VPCs), Virtual WAN (Azure's equivalent hub resource for connecting VNets), and Network Connectivity Center (NCC) are genuinely different resource models underneath and forcing them into a single resource block with conditionals produces a module that is hard to read and hard to test. Version the root module and each submodule independently using semantic versioning, pin exact provider versions per submodule (since each targets a different provider), and test with terraform validate (plus a plan against a sandbox account) in continuous integration and continuous delivery (CI/CD) before any tag is published.
Structured elaboration
flowchart TB
RootMod[Root module: cloud_provider input] --> AWSMod[transit-spoke-aws submodule]
RootMod --> AzureMod[transit-spoke-azure submodule]
RootMod --> GCPMod[transit-spoke-gcp submodule]
AWSMod --> TGW[Transit Gateway hub]
AzureMod --> VWAN[Virtual WAN hub]
GCPMod --> NCC[Network Connectivity Center hub]
Remote[(Remote state: one workspace per environment)] -.backs.-> RootMod
Module interface: consistent inputs and outputs across providers. Every provider-specific submodule below exposes the same variable names (spoke_name, hub_id, network_id, enable_appliance_mode, tags) and the same output names (attachment_id, state), even though what each one actually does underneath differs. This is what lets the root module (or a calling workspace) treat "attach a spoke to the transit hub" as one operation regardless of provider, while still letting each submodule use the correct native resource and arguments for its cloud. attachment_id is a genuine, provider-native identifier in all three submodules. state is not equally genuine everywhere, and that limitation is called out explicitly below rather than left to be discovered: checked against the pinned provider schemas (terraform providers schema -json), only the GCP submodule's google_network_connectivity_spoke resource actually exposes a computed lifecycle-state attribute; neither aws_ec2_transit_gateway_vpc_attachment nor azurerm_virtual_hub_connection exposes one at all in the provider versions pinned here, so those two submodules' state outputs are documented placeholders, not real status.
AWS and Azure state output limitation, checked against the actual provider schemas. Running terraform providers schema -json against the pinned hashicorp/aws (> 5.0) and > 3.90) providers confirms neither resource below has a computed status attribute: hashicorp/azurerm (aws_ec2_transit_gateway_vpc_attachment's only computed attributes are arn, id, security_group_referencing_support, tags_all, transit_gateway_default_route_table_association, transit_gateway_default_route_table_propagation, and vpc_owner_id, and the matching data source exposes the same set, nothing that reflects the attachment's actual lifecycle state (available, pending, deleting). azurerm_virtual_hub_connection is narrower still, its only computed attribute is id. This is why the AWS submodule's state output below is wired to vpc_owner_id (the AWS account ID that owns the target VPC, unrelated to attachment status) and the Azure submodule's state output is wired to .name (an echo of the caller's own spoke_name input, not a status either). Both are kept only for interface parity so the output always exists under that name; a caller needing the real lifecycle state has to query it outside this module, for example via aws ec2 describe-transit-gateway-vpc-attachments or the Azure CLI/ARM API for the hub connection, since no plain resource or data-source attribute currently exposes it in either provider.
AWS submodule (Transit Gateway), validated with terraform init and terraform validate against the real hashicorp/aws provider schema:
terraform {
required_providers {
aws = { source = "hashicorp/aws", version = "~> 5.0" }
}
}
variable "spoke_name" {
type = string
description = "Logical name of this spoke, shared across all provider implementations."
}
variable "hub_id" {
type = string
description = "Provider-native ID of the transit hub (Transit Gateway ID here)."
}
variable "attach_resource_ids" {
type = list(string)
description = "Subnet IDs to attach (one per AZ)."
}
variable "network_id" {
type = string
description = "VPC ID owning the subnets."
}
variable "enable_appliance_mode" {
type = bool
default = false
description = "Route all traffic for this attachment through a single AZ (needed for stateful inspection appliances)."
}
variable "tags" {
type = map(string)
default = {}
}
resource "aws_ec2_transit_gateway_vpc_attachment" "this" {
transit_gateway_id = var.hub_id
vpc_id = var.network_id
subnet_ids = var.attach_resource_ids
appliance_mode_support = var.enable_appliance_mode ? "enable" : "disable"
tags = merge(var.tags, { Name = var.spoke_name })
}
output "attachment_id" {
value = aws_ec2_transit_gateway_vpc_attachment.this.id
}
output "state" {
value = aws_ec2_transit_gateway_vpc_attachment.this.vpc_owner_id
}
Azure submodule (Virtual WAN hub connection), independently validated the same way against hashicorp/azurerm:
terraform {
required_providers {
azurerm = { source = "hashicorp/azurerm", version = "~> 3.90" }
}
}
variable "spoke_name" {
type = string
description = "Logical name of this spoke, shared across all provider implementations."
}
variable "hub_id" {
type = string
description = "Provider-native ID of the transit hub (Virtual Hub resource ID here)."
}
variable "network_id" {
type = string
description = "Resource ID of the VNet being connected to the hub."
}
variable "enable_appliance_mode" {
type = bool
default = false
description = "Kept for interface parity with the AWS/GCP spokes; Virtual WAN routes via the hub's route table instead, so this only toggles internet_security_enabled here."
}
variable "tags" {
type = map(string)
default = {}
}
resource "azurerm_virtual_hub_connection" "this" {
name = var.spoke_name
virtual_hub_id = var.hub_id
remote_virtual_network_id = var.network_id
internet_security_enabled = var.enable_appliance_mode
}
output "attachment_id" {
value = azurerm_virtual_hub_connection.this.id
}
output "state" {
value = azurerm_virtual_hub_connection.this.name
}
GCP submodule (Network Connectivity Center spoke), independently validated against hashicorp/google:
terraform {
required_providers {
google = { source = "hashicorp/google", version = "~> 5.30" }
}
}
variable "spoke_name" {
type = string
description = "Logical name of this spoke, shared across all provider implementations."
}
variable "hub_id" {
type = string
description = "Provider-native ID of the transit hub (Network Connectivity Center hub name here)."
}
variable "network_id" {
type = string
description = "Self-link of the VPC network being registered as a spoke."
}
variable "location" {
type = string
default = "global"
}
variable "enable_appliance_mode" {
type = bool
default = false
description = "Kept for interface parity; NCC has no direct equivalent for a VPC spoke, so this is intentionally unused here (see answer notes)."
}
variable "tags" {
type = map(string)
default = {}
}
resource "google_network_connectivity_spoke" "this" {
name = var.spoke_name
location = var.location
hub = var.hub_id
labels = var.tags
linked_vpc_network {
uri = var.network_id
}
}
output "attachment_id" {
value = google_network_connectivity_spoke.this.id
}
output "state" {
value = google_network_connectivity_spoke.this.state
}
All three were run through terraform fmt, terraform init -backend=false, and terraform validate against their real provider schemas (hashicorp/aws ~> 5.0, hashicorp/azurerm ~> 3.90, hashicorp/google ~> 5.30) and each reports "Success! The configuration is valid," which is as far as validation can go without live cloud credentials: plan/apply output cannot be claimed here.
A fuller, production module breakdown separates concerns further than the three spoke submodules above: a vpc (or vnet) submodule owning the tenant's own network and its subnets, a subnet submodule for the specific subnets the transit attachment needs (often requiring dedicated, non-overlapping ranges per provider's transit-gateway conventions), a vpn-connection submodule for the VPN-fallback path when the primary interconnect (the main hub-to-spoke network connection) is down, and a shared-services submodule for cross-cutting resources like a central DNS resolver or firewall inspection VPC that every spoke needs to reach, all composed by an environment-level root module rather than flattened into the three spoke modules above.
Optional features (VPN fallback, NAT, inspection). These are exposed as variables at the root module level (enable_vpn_fallback, enable_nat, enable_inspection) that conditionally include the vpn-connection submodule, a NAT gateway resource, and a route through an inspection VPC/hub respectively, using count or for_each guarded by the boolean so an environment that does not need a given feature does not pay for or provision it.
Handling provider differences. The shared interface variables (spoke_name, hub_id, network_id, enable_appliance_mode, tags) are the contract; each submodule is free to interpret them however its underlying provider requires, which is why enable_appliance_mode means something meaningfully different (AWS: real per-attachment traffic-routing behavior; Azure: internet-security toggle; GCP: unused, since NCC has no equivalent knob for this spoke type) while still being the same variable name at the calling site. Where a provider genuinely has no equivalent concept, the submodule documents that explicitly (as the GCP submodule's variable description does) rather than silently ignoring the input, and the state output above gets the same treatment for exactly this reason: rather than silently returning an unrelated value, the AWS and Azure submodules document that their state output is a placeholder because the underlying resource has no real status attribute to expose.
Versioning strategy. The root module and each provider submodule get independent semantic-version tags in their own source repository path (or a shared monorepo with per-directory tags), so a breaking change to the Azure submodule's interface does not force every AWS-only consumer to also bump their pinned version; calling code pins an exact or constrained version (version = "~> 2.1") per submodule, and CI runs the full validation suite against any proposed version bump before it is tagged.
Testing in CI/CD, unit and integration as two distinct layers. On every pull request: terraform fmt -check, terraform validate for each submodule (as run above), and a unit-test layer using Terraform's native terraform test framework with mock_provider, asserting that a given set of input variables produces the expected planned resource attributes without touching any real cloud account or needing credentials at all. For the AWS submodule above, this unit test (run against the real module with terraform test on Terraform 1.15, no credentials required) passes:
mock_provider "aws" {}
run "creates_attachment_with_expected_name_tag" {
variables {
spoke_name = "test-spoke"
hub_id = "tgw-0123456789abcdef0"
network_id = "vpc-0123456789abcdef0"
attach_resource_ids = ["subnet-aaa", "subnet-bbb"]
enable_appliance_mode = true
}
assert {
condition = aws_ec2_transit_gateway_vpc_attachment.this.appliance_mode_support == "enable"
error_message = "appliance_mode_support should be enable when enable_appliance_mode = true"
}
assert {
condition = aws_ec2_transit_gateway_vpc_attachment.this.tags["Name"] == "test-spoke"
error_message = "Name tag should match spoke_name"
}
}
Running terraform test against the AWS submodule with this file present produces Success! 1 passed, 0 failed., confirming both the appliance-mode toggle and the tag-merging logic behave as intended, entirely offline. Only after unit tests like this pass does a slower integration-test layer run: a terraform plan against a dedicated sandbox account/subscription/project per provider (using ephemeral or long-lived least-privilege sandbox credentials scoped only to test resources) to catch anything unit tests and validate cannot see, such as a resource argument that is syntactically and logically valid but rejected by the actual API, followed by a less frequent full integration cycle (apply into the sandbox, verify the resource exists via a data source or provider API call, then destroy) that catches drift between the module and the provider's live behavior that plan alone cannot.
Worked example
A platform team maintains this module across three repositories (or three directories in one monorepo), transit-spoke-aws at v2.3.0, transit-spoke-azure at v1.4.0, and transit-spoke-gcp at v1.1.0, each independently tagged. A consuming environment's root module pins all three and, based on which clouds that specific environment actually uses, instantiates only the relevant submodules:
module "aws_spoke" {
source = "git::https://example.com/transit-modules.git//transit-spoke-aws?ref=v2.3.0"
spoke_name = "prod-us-east-1"
hub_id = var.aws_tgw_id
network_id = var.aws_vpc_id
attach_resource_ids = var.aws_subnet_ids
tags = local.common_tags
}
A change that adds a new optional argument to the AWS submodule's interface (say, native NAT integration) ships as v2.4.0 (a minor, backward-compatible bump); a change that renames attach_resource_ids would require a v3.0.0 major bump, signaling to every consumer, via their own pinned constraint, that they need to review the change before adopting it, rather than silently picking it up.
Trade-offs & pitfalls
- Forcing one module to branch internally on
cloud_provider(a single resource block withcount = var.cloud_provider == "aws" ? 1 : 0repeated three times) looks like less code up front but produces a module where changing one provider's behavior risks breakingterraform planoutput for the other two, and where the three providers' genuinely different resource models get flattened into variables that fit none of them well; separate submodules behind a shared interface avoid this at the cost of more files to maintain. terraform validatecatches syntax and schema errors but not everything a real API will reject (a value that is the right type but violates a provider-side constraint, for example); treating a cleanvalidateas equivalent to a successfulapplyis a common and costly overconfidence, which is exactly why the CI pipeline needs a sandboxplan, and ideally a periodic realapply/destroycycle, notvalidatealone.- Independent versioning per submodule means a consumer using all three clouds has to track three version numbers instead of one, which is more cognitive overhead than a single monolithic version, but avoids forcing an unrelated provider's consumers to absorb a breaking change they never asked for.
- The
enable_appliance_modevariable meaning three different things across providers (real behavior on AWS, a different real behavior on Azure, nothing on GCP) is a deliberate interface-consistency trade-off; a consumer who assumes it behaves identically everywhere will be surprised on GCP specifically, which is why the GCP submodule's variable description says so explicitly rather than leaving it to be discovered. - The
stateoutput is not a genuine cross-provider abstraction the wayattachment_idis: only GCP's reflects a real computed lifecycle attribute, while AWS's and Azure's are documented placeholders (an unrelated account ID for AWS, an echo of the input name for Azure) because neitheraws_ec2_transit_gateway_vpc_attachmentnorazurerm_virtual_hub_connectionexposes a status attribute in the current provider versions. A consumer who wires orchestration logic (a health check, a readiness gate) off ofstatewithout checking this will get a value that means something on GCP and means nothing at all on the other two clouds, which is exactly the kind of interface promise that looks uniform in the variable list and is not uniform in practice.
Assess the security, privacy, and compliance implications of using a third-party managed SD-WAN provider whose traffic transits shared carrier infrastructure. Address data sovereignty, access to logs by the provider, encryption in transit and at rest, contractual audit rights, and how to demonstrate compliance to auditors.
Sample Answer
Direct answer
Treat a third-party managed SD-WAN (Software-Defined Wide Area Network) provider the same way any subprocessor handling your traffic would be treated: assume the provider can see metadata about every flow crossing its infrastructure even when payload is encrypted, get contractual audit rights and a clear data-residency commitment in writing before signing, and verify encryption in transit and at rest independently rather than accepting a marketing claim.
Structured elaboration
Data sovereignty
Ask, specifically, in which countries the provider's orchestration and control-plane data, not just the customer's actual traffic, is stored and processed, since SD-WAN control planes typically log flow metadata, configuration, and sometimes packet captures for troubleshooting, and that telemetry can cross a jurisdictional boundary the traffic itself never does. This matters most for regulated data, personal data under GDPR (the EU's General Data Protection Regulation), or sector-specific rules such as HIPAA (the Health Insurance Portability and Accountability Act) in healthcare, where applicable law cares about where data is processed, not just where it is encrypted.
Access to logs by the provider
A managed SD-WAN provider's own operations staff almost always have some level of access to flow logs, performance telemetry, and configuration, because that is how the managed part of the service gets delivered. Get explicit answers to: who at the provider can see this data, what triggers an access event, routine monitoring versus a support ticket that was opened, how long logs are retained, and whether the customer is notified of provider-side access, before assuming encrypted answers every access question.
Encryption in transit and at rest
IPsec (IP Security, a protocol suite that encrypts and authenticates traffic between two endpoints) or a comparable tunnel encryption in transit is close to universal among SD-WAN products; the gap to actually verify is encryption at rest for anything the provider's control plane stores about the customer's traffic, configuration backups, log archives, and packet captures kept for troubleshooting. Ask for the specific standard, AES-256 is the common baseline, and where the encryption keys are held, since a provider holding both the encrypted data and the keys is a materially weaker guarantee than one supporting customer-managed keys.
Contractual audit rights
The right to request evidence of compliance is not automatic; it has to be a specific, negotiated clause in the master service agreement, naming what can be audited, SOC 2 Type II reports (a third-party auditor's report on a provider's security controls, examined over a period of time) or ISO 27001 certification (an international certification for an organization's information-security management system), penetration test summaries, how often, and with what notice. A provider that will not commit to a specific audit-rights clause in writing is showing, in advance, how a future compliance question with them will go.
Demonstrating compliance to auditors
Auditors will ask for proof the provider's controls are adequate, not just the provider's word for it, so collect the provider's SOC 2 or ISO 27001 report annually, not once at signing, map their control objectives explicitly onto the specific regulatory requirement being audited against, a SOC 2 report does not automatically prove GDPR or HIPAA compliance, it evidences specific control objectives that may or may not cover what those regulations require, and keep the underlying contract's audit-rights and data-residency clauses on file as the artifact that proves the right to ask existed in the first place.
Worked example
A healthcare organization uses a managed SD-WAN provider whose primary data centers are in a different country from where patient data is generated. The organization's HIPAA risk assessment has to account for the provider seeing flow metadata, source, destination, timing, even if payload is encrypted, about traffic carrying protected health information, which under HIPAA makes the provider a business associate, requiring a signed Business Associate Agreement, not just a standard commercial contract, specifically because addressable safeguards under HIPAA mean the organization must either implement an equivalent control or document, in writing, why it does not apply, not that the safeguard is optional.
Trade-offs and pitfalls
The most common mistake is treating encrypted traffic as the whole answer to a security review, when the actual exposure is usually metadata and control-plane access, not payload interception. The second is signing a managed SD-WAN contract without a specific audit-rights clause because sales-cycle pressure makes it feel like a minor point; it is the clause an auditor will ask for by name eighteen months later, and it cannot be added retroactively without renegotiating.
Design an overlay-underlay architecture to support multi-tenant hybrid cloud networks using VXLAN/EVPN (or NVGRE). Explain how you would achieve tenant isolation, design IP addressing to avoid collisions, connect the overlay fabric to public cloud provider networks, plan for MTU and fragmentation, and outline operational troubleshooting approaches.
Sample Answer
Direct answer
Give every tenant its own VRF (Virtual Routing and Forwarding instance, an isolated routing table) mapped to its own VNI (VXLAN Network Identifier), run EVPN (Ethernet VPN, a BGP-based control plane) as the source of truth for which MAC and IP addresses live behind which fabric device, carve tenant address space out of a central IPAM (IP Address Management) plan so overlapping tenant ranges never touch a shared segment, and terminate the overlay at the data-center edge rather than trying to extend it natively into a hyperscaler's network, since AWS, Azure, and GCP do not let a customer run their own EVPN control plane inside a VPC or VNet. Recommend VXLAN/EVPN over NVGRE (Network Virtualization using Generic Routing Encapsulation) as the default encapsulation: NVGRE has a smaller encapsulation footprint but materially worse hardware and cloud-provider support, and an interview-grade design should not trade a widely supported standard for a marginal byte saving.
Design
Tenant isolation
One VRF per tenant on every fabric gateway, one VNI per tenant, and EVPN route targets (RT) scoped so tenant A's VRF only imports routes tagged with tenant A's RT. This is the same mechanism MPLS VPNs (Multiprotocol Label Switching Virtual Private Networks) have used for two decades, just running over VXLAN instead of MPLS labels. A shared-services VRF (DNS, logging, a package mirror) gets its own RT that every tenant VRF is allowed to import from, one way, so tenants can reach shared infrastructure without reaching each other.
IP addressing without collisions
Inside its own VRF, a tenant can safely reuse the exact same RFC 1918 (the private-address-space standard) ranges another tenant uses, because the VNI and VRF, not the IP address, are what keep traffic apart on the fabric. The rule that matters is: anything that crosses into a shared segment (shared services, a common egress NAT gateway, or a cloud VPC/VNet more than one tenant's traffic must reach) has to be globally unique, and a central IPAM authority, not each tenant team, owns and allocates that shared-reachable space.
Connecting the overlay to public cloud networks
None of the three major clouds expose an EVPN control plane or arbitrary Layer 2 extension into a customer VPC/VNet. The overlay fabric therefore has to terminate at the edge of the data center or colocation facility, on a pair of gateway devices (hardware VTEPs, VXLAN Tunnel Endpoints, the devices that wrap and unwrap overlay traffic, or software ones such as FRRouting), and from there the design switches to native cloud routing: one Direct Connect or ExpressRoute private virtual interface per tenant VRF, each landing in that tenant's own VPC or VNet, so tenant separation is preserved end to end even though the encapsulation technology changes at the boundary.
MTU and fragmentation
Same arithmetic as a single-tenant VXLAN design (50 bytes of overhead per RFC 7348), but now it has to hold across every hop from tenant VRF to cloud VPC, and NVGRE's overhead (14-byte outer Ethernet, 20-byte outer IP, an 8-byte GRE header carrying the 24-bit Virtual Subnet ID per RFC 7637, about 42 bytes total) does not change that conclusion since the Direct Connect or ExpressRoute leg beyond the gateway carries native, unencapsulated traffic anyway. Set the internal fabric to jumbo frames (9000 bytes is common) so the 42-to-50-byte tax is invisible, and clamp TCP MSS at the point where traffic crosses onto a path that is not fully controlled, such as an internet-routed VPN backup link.
Worked example: sizing the shared-reachable address pool
graph TD
T1[Tenant A: VRF, VNI 1001] --> VTEP1[Fabric VTEP]
T2[Tenant B: VRF, VNI 1002] --> VTEP1
VTEP1 -->|EVPN control plane over BGP| VTEP2[DC Edge VTEP]
VTEP2 -->|Direct Connect, tenant A VIF| VPCA[Tenant A VPC]
VTEP2 -->|Direct Connect, tenant B VIF| VPCB[Tenant B VPC]
For the shared-reachable IPAM pool (not the per-tenant reused space), carving a 10.0.0.0/8 supernet into one /16 per tenant gives:
N=2p2−p1=216−8=28=256256 non-overlapping /16 blocks, so this scheme supports up to 256 tenants that need a shared-reachable range before the central pool has to be resized, with the first three blocks being 10.0.0.0/16, 10.1.0.0/16, and 10.2.0.0/16, and the last being 10.255.0.0/16. Anything a tenant keeps entirely inside its own VRF does not consume from this pool at all, since it never has to be globally unique.
Trade-offs and pitfalls
The design's biggest risk is not technical, it is operational: a single misconfigured route-target import on one gateway silently grants one tenant reachability into another tenant's VRF, which is a security incident, not an outage, so it usually is not noticed quickly. Treat the RT-to-VRF mapping as the single source of truth, keep it in version control, and alert on any gateway whose configured imports drift from that source. The second pitfall is assuming multicast will handle overlay flooding between on-prem and cloud the way it might inside a single data center: public cloud VPCs and VNets do not support multicast, so BUM (Broadcast, Unknown-unicast, and Multicast) traffic has to use ingress replication (unicast head-end replication) or, better, be mostly eliminated by relying on EVPN's control-plane learning instead of flood-and-learn. Finally, resist running NVGRE just because Azure's own internal SDN historically used it: that is an internal implementation detail, not something exposed for customer-built overlays, and VXLAN/EVPN remains the better-supported choice across hardware, hypervisors, and cloud interconnect equipment.
Explain what SD-WAN is and how it changes the design of hybrid connectivity between branch offices, on-prem data centers, and cloud provider networks. Cover central control vs local forwarding, policy-based path selection, encryption, and common deployment models (managed appliance, virtual edge, overlay vs underlay awareness).
Sample Answer
Direct answer
SD-WAN (Software-Defined Wide Area Network) separates the control plane, deciding which path a packet should take, from the data plane, actually forwarding it, and centralizes that control plane so policy is written once and pushed to every site, instead of every branch router being configured by hand. That single change is what lets a hybrid network use several cheap underlay links (broadband, LTE) intelligently instead of relying on one expensive, purpose-built circuit (MPLS, Multiprotocol Label Switching) for everything.
Structured elaboration
Central control, local forwarding
A central controller, or controller cluster, holds the policy: which application goes on which path, what happens when a path degrades, what gets encrypted. Each site's SD-WAN edge appliance still forwards packets locally, at line rate, using the policy it was last given, so a brief loss of contact with the controller does not stop traffic from moving, it just means the site cannot receive a policy update until contact is restored.
Policy-based path selection
Instead of one static route toward the data center, an SD-WAN edge continuously measures loss, latency, and jitter across every available underlay and steers each application's traffic to whichever path currently meets that application's requirement, re-evaluating in near real time rather than only failing over when a link goes fully down.
Encryption
Traffic is encrypted, typically IPsec (IP Security, a protocol suite that encrypts and authenticates traffic between two endpoints), between SD-WAN edges by default, which is what makes it safe to route real branch traffic over ordinary business broadband instead of a private circuit: the link itself does not need to be trusted if every packet crossing it is encrypted and authenticated.
Deployment models
| Model | What it looks like |
|---|---|
| Managed appliance | A vendor-supplied hardware box at each site, typically bundled with a managed service contract |
| Virtual edge | The same software running as a VM (Virtual Machine) or container, at a branch, inside a cloud VPC/VNet, or both |
| Overlay-underlay awareness | The SD-WAN overlay actively measures the real underlay links beneath it (broadband, LTE, occasionally MPLS) rather than treating the underlay as an opaque, always-available pipe |
Worked example
A branch has two underlays: a business broadband circuit and an LTE backup. Under a purely static design, all traffic uses broadband until it fails completely, then everything moves to LTE. Under SD-WAN, a bulk file transfer stays on broadband, cheap, high bandwidth, latency does not matter much, while a voice call is steered onto whichever link is currently measuring under roughly 30 ms of jitter and 1 percent loss, which might mean voice moves to LTE for a few minutes while broadband is congested, then moves back once broadband's measured jitter drops again, all without a human intervening or the file transfer even noticing.
Trade-offs and pitfalls
SD-WAN does not remove the underlying physics of the internet: a broadband link with genuinely poor upstream connectivity is still a poor link, SD-WAN just detects that faster and reacts to it, it does not make the link better. It also introduces a new dependency, the controller and its reachability, so a design has to consider what a site does during an extended controller outage; in practice, keep forwarding on last-known policy, which is why that property is a requirement, not a nice-to-have, when selecting a platform. Finally, overlay-underlay awareness only helps if the SD-WAN platform is actually measuring the real underlay characteristics continuously; a platform that treats every link as equivalent until it is completely down is not meaningfully different from the static routing SD-WAN is supposed to replace.
Unlock Full Question Bank
Get access to all Multi-Cloud and Hybrid Cloud Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.