Cloud Networking and VPC Design Questions
Designing networks inside a cloud provider: VPC/VNet topology, subnets, route tables, gateways, NAT, and peering, plus private connectivity through VPC endpoints and cloud load balancers. Covers segmentation, security groups and network ACLs, hybrid connectivity to on-premises data centers over VPN or dedicated links like Direct Connect and ExpressRoute, IP address planning across many VPCs and accounts, and how cloud network design differs from traditional data-center networking.
Design a multi-account VPC architecture for a large enterprise (50+ accounts) that needs centralized shared services (monitoring, logging, AD), workload isolation, low-latency intra-VPC connectivity, and centralized egress control. Provide a topology, recommended connectivity primitives (TGW/peering/DX), IP allocation approach, and governance controls to prevent accidental exposure.
Sample Answer
Direct answer
At 50-plus accounts, the shape that scales is a hub-and-spoke topology built on a Transit Gateway (TGW): one shared-services hub VPC (Virtual Private Cloud) holding centralized logging, monitoring, Active Directory (AD), and CI/CD (continuous integration and continuous delivery) tooling, with every workload account's VPC attached as a spoke, all inter-VPC and on-premises traffic riding the TGW instead of a mesh of point-to-point peerings. Isolation between workloads and between environments (staging versus production) is enforced with separate TGW route tables, not with separate physical topologies, and all outbound internet traffic is forced through a shared inspection point in the hub rather than being allowed to exit from each spoke independently.
Structured elaboration
Topology. A dedicated network/shared-services account owns the TGW, a hub VPC with the shared logging, monitoring, and directory services, and a Direct Connect (DX) or VPN (Virtual Private Network) attachment for on-premises connectivity. Every workload account attaches its VPC to the TGW as a spoke. Peering does not scale here: 50 VPCs peered pairwise would need up to 1,225 individual peering connections, each with its own route table entries, versus 50 single attachments to one TGW.
flowchart TB
DX[Direct Connect or VPN to on-prem] --> HUB
subgraph HUB[Shared-services account: hub VPC]
TGW[Transit Gateway]
NAT[Central NAT and egress inspection]
DNS[Shared DNS resolver]
LOG[Central logging]
CICD[Shared CI/CD tooling]
end
TGW --- SPOKE1
TGW --- SPOKE2
TGW --- SPOKE3
subgraph SPOKE1[Prod account A]
VPCA[Workload VPC]
end
subgraph SPOKE2[Prod account B]
VPCB[Workload VPC]
end
subgraph SPOKE3[Staging account]
VPCC[Workload VPC]
end
Connectivity primitives. TGW for the hub-and-spoke fabric itself (attachments per TGW default to a quota in the thousands, comfortably covering 50-plus accounts with room to grow); a Direct Connect Gateway (or a Site-to-Site VPN as a lower-committment or failover path) attached to the TGW for on-premises reach, rather than a separate virtual private gateway per VPC; VPC peering reserved only for the rare case of two spokes needing a private connection the hub shouldn't see.
IP allocation. Carve one supernet (for example 10.0.0.0/12) into fixed-size per-account blocks up front so growth never collides. A /12 splits cleanly into 512 blocks of /21 (2,048 addresses each), which comfortably covers 50-plus accounts today with headroom for hundreds more without ever having to resize an existing VPC's CIDR (Classless Inter-Domain Routing, the notation for an IP address range, like the /21 blocks above) after the fact.
Governance controls to prevent accidental exposure.
- Service control policies (SCPs) at the organizational-unit level deny the creation of internet gateways or public-facing load balancers in workload accounts, so any internet-bound path is forced through the hub's controlled egress.
- The hub is the only account that can attach a Direct Connect Gateway or manage the TGW's route tables; workload account administrators get attach permission only, via AWS Resource Access Manager (RAM), not route-table edit rights.
- Tag-enforced IP address management (IPAM), using AWS IPAM or an equivalent CIDR-tracking system, so a new account can only request a CIDR block from the pre-carved pool, never pick an arbitrary range.
- Automated drift detection (AWS Config rules or a scheduled Lambda) alerting when a spoke account's route table gains a route that bypasses the hub's inspection VPC.
Isolation angle: shared-services hub and staging-versus-production separation. The shared-services hub itself typically holds three functions worth naming explicitly: a logging sink (centralized CloudWatch Logs or an S3-based log lake that every spoke ships to, so no individual account can tamper with its own audit trail), a shared CI/CD hub (build and deploy tooling with cross-account IAM roles into each spoke, rather than duplicating pipeline infrastructure per account), and shared DNS (a Route 53 Resolver setup that every spoke inherits). Staging and production get separate TGW route tables associated with their respective spoke attachments: the production route table only propagates routes to other production spokes and the hub's shared services, and the staging route table is scoped the same way for staging, so a compromised staging workload has no route to a production spoke even though both ride the same physical TGW. Both route tables still send 0.0.0.0/0 through the same central egress-inspection VPC (typically a Gateway Load Balancer fronting a third-party or AWS Network Firewall appliance), so all outbound traffic, staging or production, gets the same inspection regardless of which environment it came from.
Worked example
Using the supernet above, 10.0.0.0/12 yields 10.0.0.0/21, 10.0.8.0/21, 10.0.16.0/21, 10.0.24.0/21, and so on, one per account, computed as 2(21−12)=512 available blocks. Account 1 (shared services) takes 10.0.0.0/21; account 2 (a production spoke) takes 10.0.8.0/21; account 3 (a staging spoke) takes 10.0.16.0/21. Because every block comes from the same never-touched allocation table, no two accounts can ever be assigned the same range even as new accounts are onboarded automatically by a script that just takes the next unused /21.
Trade-offs and pitfalls
The TGW hub becomes a single blast-radius point: a misconfigured route table there can affect every spoke at once, so changes to it need a stricter change process than changes inside any individual spoke. Centralizing egress inspection adds a real latency and throughput cost (every spoke's internet-bound packet now makes an extra hop through the inspection VPC), which is the right trade for compliance and visibility but is a genuine tax worth measuring, not assuming away. Finally, resist the temptation to let "just this one account" skip the hub and peer directly with another spoke for a special use case: every exception erodes the governance guarantee that all traffic, at least on paper, flows through a point you can audit.
Compare VPC Peering, Transit Gateway (TGW), and PrivateLink (VPC Endpoints) as connectivity options. For each option, list scale limits, typical traffic flow patterns, security boundaries, routing complexity, and scenarios where it should be preferred (e.g., shared services, SaaS integrations, hub-and-spoke networks).
Sample Answer
Direct answer
VPC Peering, Transit Gateway, and PrivateLink (VPC endpoints) are not three competing answers to the same question; they solve three different connectivity shapes. Peering gives two specific VPCs full, non-transitive network-layer reachability to each other. Transit Gateway gives many VPCs (and on-premises connections) a shared, centrally routed hub so any attached network can reach any other attached network according to its route tables. PrivateLink gives a consumer private access to one specific published service, with no network-layer reachability into the provider's VPC at all. Choosing between them is really choosing how much of the provider's network you want the consumer to be able to see: everything (Peering, Transit Gateway) or nothing but the one thing they asked for (PrivateLink).
Structured elaboration
| Dimension | VPC Peering | Transit Gateway | PrivateLink (VPC Endpoints) |
|---|---|---|---|
| Scale limit | Each VPC has a quota on active peering connections (commonly around 50 by default, adjustable to roughly 125); practically limited well before that by the quadratic growth of a full mesh | A hard cap of 5,000 attachments per Transit Gateway (adjustable via a quota increase), and a per-Transit-Gateway route quota of 10,000 combined dynamic and static routes across its route tables (this specific route quota is not self-service adjustable; raising it requires contacting AWS Support or your account team) | Scales per consumer independently; the meaningful limits are the number of interface endpoints a single consumer VPC can hold and the number of consumers a single published service can serve, unrelated to how many other services or consumers exist elsewhere |
| Traffic flow pattern | Direct, VPC to VPC, no intermediate hop | Hub and spoke, through the Transit Gateway's own routing | Consumer to a private IP inside its own subnet that forwards to the published service; the consumer never sees the provider's network topology at all |
| Security boundary | Full network-layer trust between the two peered VPCs' address spaces, narrowed only by security groups and NACLs on each side | Full network-layer trust across whatever a route table permits between attachments | Strong isolation: the consumer can reach only the published service, and the provider gets no inbound network visibility into the consumer's VPC either |
| Routing complexity | Manual, per-connection route entries; grows unmanageable as connection count grows | Centralized in the Transit Gateway's route tables, but those route tables themselves become the thing you must carefully govern as attachment count grows | Minimal: no route-table coordination between consumer and provider is needed at all, only an endpoint service definition and, per consumer, one interface endpoint |
| Best-fit scenario | A small, stable number of VPC-to-VPC relationships that genuinely need full network reachability | Centralizing shared network-layer resources (a security-inspection VPC, a shared Direct Connect gateway, i.e. a dedicated physical network link back to an on-premises data center rather than a connection over the public internet) that many VPCs need routed access to | Shared services and software-as-a-service (SaaS) style integrations: one team or vendor publishes a specific API or capability that many consumers call, without those consumers needing or wanting broader network access |
Shared services. A shared-services VPC hosting things like centralized logging ingestion or an internal artifact repository is a natural Transit Gateway case if many other VPCs genuinely need routed, low-level network access to it (for example, an on-host logging agent shipping over a raw TCP protocol that expects direct IP reachability). If the same shared capability is exposed as a well-defined API instead, PrivateLink is usually the better fit, since the consuming VPCs never need broader network reachability into the shared-services VPC at all, only access to that one API.
Software-as-a-service (SaaS) integrations. This is PrivateLink's clearest use case: a SaaS vendor publishes their service as an endpoint service, and each customer creates an interface endpoint in their own VPC to reach it privately, with the vendor never gaining any network visibility into the customer's VPC and the customer never gaining any visibility into the vendor's internal architecture beyond the one published service. Neither Peering nor Transit Gateway is appropriate here, since both would require some form of network-layer trust relationship between a customer and a vendor that neither party actually wants.
Hub-and-spoke networks. This is the label most directly describing Transit Gateway's own architecture: many spoke VPCs attach to one hub, and the hub's route tables decide reachability. It becomes the clear choice over a peering mesh the moment the organization needs centralized policy over which spokes can reach which other spokes or shared resources, since that policy lives in one place (the Transit Gateway's route tables) rather than being scattered across dozens of individually managed peering connections.
Worked example
Peering lifecycle, operationally. Even a small, deliberately peering-based design has an operational lifecycle worth naming: a peering connection is created and sits in a pending-acceptance state until the target VPC's owner (which may be a different account) explicitly accepts it, after which both sides must independently add route-table entries pointing at the connection, since accepting it does not automatically add routes. When a peering relationship is no longer needed, deleting the connection immediately breaks connectivity, but the route-table entries referencing it on both sides are left behind as stale, dangling routes that need to be cleaned up separately, which is an easy step to forget and a common source of confusing, half-working configurations discovered much later.
Multi-account hybrid disaster-recovery (DR) example. A financial services company runs its primary environment on-premises with a cloud-based DR environment spread across several AWS accounts. Transit Gateway connects those accounts' VPCs together and to a Direct Connect gateway carrying the on-premises link, giving DR failover the full routed reachability it needs between the recovery environment's tiers and back to on-premises during a failover event. Separately, a small number of especially sensitive internal services (a secrets-management API, an internal certificate authority) are exposed to the DR environment specifically via PrivateLink rather than being reachable over the same Transit Gateway routing as everything else, which means that even in a full DR failover scenario, those sensitive services are reachable only by the specific consumers that were explicitly granted an interface endpoint, not by anything else that happens to be attached to the same Transit Gateway. Avoiding the public internet or a shared, broadly-routed hop for that sensitive traffic, by keeping it on a dedicated private path instead, is the qualitative benefit worth designing for; treat any specific percentage improvement you might see quoted for this kind of change as something to measure in your own environment with real traceroutes and load tests, not a number to design around in advance.
Service mesh as a related, non-competing idea. For east-west traffic between microservices within a single VPC or a tightly coupled set of them, a service mesh (handling service discovery, mutual TLS, and traffic policy at the application layer) is solving a different problem than any of these three options: Peering, Transit Gateway, and PrivateLink all operate below the application layer, deciding whether a network path exists at all, while a service mesh assumes the network path already exists and adds application-aware routing and security on top of it. The two are complementary, not substitutes: a service mesh commonly runs on top of a Transit-Gateway-connected or PrivateLink-exposed set of services, not instead of one of them.
Trade-offs and pitfalls
The most common mistake is picking Transit Gateway for a case that is really a PrivateLink case, most often because it feels like the more "complete" or future-proof answer; this over-grants network reachability where a scoped service endpoint would have done the job with a smaller security footprint and less route-table governance overhead. A second pitfall, already noted above, is forgetting to clean up stale route-table entries after deleting a peering connection, which leaves a route pointing at a connection ID that no longer exists, silently harmless until someone reuses that ID space or spends time debugging why traffic to that destination behaves unexpectedly. Finally, resist quoting a specific performance improvement number for moving traffic off the public internet onto a private path unless you have actually measured it in the environment in question; the qualitative direction (lower and more consistent latency, no exposure to public internet routing variability) is reliable, the exact magnitude is not something to state as a fact without measurement.
Design a hybrid connectivity solution for an enterprise datacenter requiring sustained 10 Gbps throughput and 99.99% availability. Compare options: multiple HA VPN tunnels with BGP versus a dedicated provider connection (Direct Connect/ExpressRoute) with VPN fallback. Discuss encryption, BGP failover, redundancy zones, performance, and operational costs.
Sample Answer
Direct answer
For an enterprise datacenter needing a sustained 10 Gbps and 99.99% availability, the real decision is whether to reach that bar with multiple Site-to-Site VPN tunnels aggregated by Border Gateway Protocol (BGP) and equal-cost multi-path (ECMP) routing over the public internet, or with a dedicated private circuit (Direct Connect or ExpressRoute) as primary and VPN kept only as a lower-throughput fallback. The multi-tunnel VPN approach can reach 10 Gbps on paper by aggregating enough tunnels, but every one of those tunnels still crosses the uncontrolled public internet, so it cannot deliver the consistent latency or the vendor-backed reliability commitment that 99.99% availability at sustained high throughput realistically demands; the dedicated-circuit-primary design is the recommended path, with multi-tunnel VPN reserved as the resilient fallback, not the primary transport.
Structured elaboration
Throughput math, concretely. A single standard Site-to-Site VPN tunnel is capped at roughly 1.25 Gbps of throughput (a large-bandwidth tunnel variant reaches roughly 5 Gbps where supported). Reaching 10 Gbps purely through VPN therefore requires aggregating multiple tunnels using ECMP with dynamic (BGP) routing, for example eight or more standard tunnels, or two to three large-bandwidth tunnels, load-shared across them. This is achievable on paper, but it comes with real caveats: ECMP load-sharing across tunnels is typically per-flow, not per-packet, so a single large flow (one big transfer) is still capped at one tunnel's throughput rather than the aggregate, and the actual sustained aggregate throughput depends on having enough distinct flows to spread evenly across all the tunnels, which is not guaranteed for every traffic pattern. A dedicated circuit, by contrast, is provisioned as a single port at the target speed (a 10 Gbps port, for instance) with no per-flow ceiling of this kind.
| Dimension | Multiple HA VPN tunnels with BGP/ECMP | Dedicated circuit (Direct Connect/ExpressRoute) with VPN fallback |
|---|---|---|
| Encryption | Native: every tunnel is IPsec-encrypted (IPsec, Internet Protocol Security, encrypts and authenticates each packet) by design | Not encrypted by default; the circuit itself is private but not automatically encrypted in transit, so an overlay (MACsec, a link-layer encryption standard for the physical connection itself, where supported, or an application/VPN-layer encryption on top) is needed if encryption-in-transit is a hard requirement |
| BGP failover | Failover is tunnel-to-tunnel within the same underlying transport (the public internet); if the internet path itself degrades broadly, every tunnel is affected simultaneously | Failover is transport-to-transport: BGP shifts from the dedicated circuit to the VPN fallback, which rides a genuinely different path, so a problem specific to the dedicated circuit's provider or facility does not also take down the fallback |
| Redundancy zones | Achieved by terminating tunnels on redundant customer-gateway devices and, ideally, redundant internet connections on the on-premises side | Achieved by provisioning the dedicated circuit itself across multiple devices and, for the higher resiliency tiers, multiple physical locations, on top of which the VPN fallback adds a second, independent transport entirely |
| Performance (latency, consistency) | Subject to public internet routing and congestion for every tunnel; throughput is also subject to the per-flow ECMP ceiling described above | Consistent, low latency on the provider's backbone; not subject to public internet congestion at all |
| Reliability commitment | No cloud-provider service-level agreement (SLA) for the internet portion of any tunnel's path, regardless of how many tunnels you aggregate | Backed by an SLA when deployed in a resilient, multi-device, multi-location configuration; a single non-redundant circuit still carries no SLA |
| Operational cost | Lower base cost (no dedicated port or cross-connect), but the engineering cost of correctly configuring and validating multi-tunnel ECMP aggregation at scale is nontrivial | Higher recurring cost (port-hours, cross-connect, often colocation), plus a materially longer provisioning lead time (commonly weeks) before the dedicated circuit exists at all |
Worked example
An enterprise choosing the dedicated-circuit-primary design provisions two Direct Connect connections, each sized to fully cover the 10 Gbps requirement on its own (not two 5 Gbps connections that only reach 10 Gbps combined, since that would turn any single connection's failure into an immediate capacity shortfall), terminating on separate routers, ideally at two physically separate colocation facilities, which is the resiliency configuration that actually qualifies for an SLA. Both connections use BGP with an explicit local-preference or AS-path configuration so traffic prefers whichever of the two is healthiest, giving device- and facility-level redundancy within the dedicated-circuit tier itself, entirely independent of the VPN fallback. A Site-to-Site VPN, using two tunnels for its own internal redundancy, is layered on top as the fallback transport, sized not to match the full 10 Gbps (accepting that a Direct Connect outage means running in a temporarily degraded-throughput state) but to comfortably carry whatever the business has defined as the minimum acceptable throughput during a primary-path outage. BGP is configured so both dedicated circuits are preferred over the VPN under normal conditions, and only if both dedicated connections are down simultaneously does traffic fail over to the VPN path.
Why the availability math favors this over VPN-only, even before counting SLAs. 99.99% availability allows for roughly 52 minutes of downtime per year. A design whose only transport is multiple VPN tunnels over the public internet is fundamentally exposed to internet-wide routing events (a major transit provider issue, a regional internet disruption) that can degrade every tunnel at once, since they all ultimately traverse the same uncontrolled medium; no amount of tunnel aggregation removes that shared-fate risk. A design with a genuinely separate transport as the fallback (VPN, when the primary is a dedicated circuit riding a completely different physical and provider path) does not share that single point of common-mode failure, which is the structural reason it is the safer way to reach a 99.99% target, independent of any specific SLA percentage either transport happens to carry.
Trade-offs and pitfalls
The most common mistake is sizing multi-tunnel VPN aggregation to hit 10 Gbps in aggregate and treating that as equivalent to a dedicated circuit's 10 Gbps, without accounting for the per-flow ECMP ceiling: a workload dominated by a small number of large, sustained flows (a database replication stream, a large backup job) will not actually see aggregate throughput, since each such flow still rides a single tunnel. A second pitfall is provisioning two dedicated circuits but terminating both on the same router or in the same facility "to simplify operations," which quietly forfeits the SLA-qualifying resiliency tier and reintroduces the single point of failure the second circuit was meant to remove. Finally, teams sometimes size the VPN fallback to match the dedicated circuit's full throughput "to be safe," which is rarely necessary and adds ongoing cost for a path that, by design, only needs to carry traffic during the rare window the primary is down; size it to the minimum acceptable degraded throughput instead, and validate that figure against what the business actually tolerates, not against the primary path's full capacity.
Compare three patterns for exposing internal platform services across accounts: (1) a shared-services VPC (hub) with peering/TGW spokes, (2) cross-account private service endpoints (PrivateLink), and (3) a federated service mesh spanning accounts. Evaluate each pattern for security boundary strength, operational overhead, scalability, and service failure isolation.
Sample Answer
Direct answer
These three patterns trade security-boundary strength against operational overhead in almost a straight line: a shared-services hub VPC with peering or Transit Gateway (TGW) spokes is the cheapest to operate but has the weakest boundary (network-level trust extends across the whole hub), cross-account PrivateLink gives each consumer a narrow, individually-revocable connection with the strongest boundary but the highest per-consumer setup cost, and a federated service mesh spanning accounts sits in between on cost but introduces a different, and often underestimated, category of overhead: the mesh's own control plane, the centralized system that issues certificates, tracks service identity, and pushes out policy to every proxy, becomes a new shared dependency and a new thing that can fail across every account it spans.
Structured elaboration
| Pattern | Security boundary strength | Operational overhead | Scalability | Service failure isolation |
|---|---|---|---|---|
| Shared-services hub VPC, peering or TGW spokes | Weakest: any spoke attached to the hub can generally reach any service exposed there, unless painstakingly restricted with per-spoke route tables and security groups | Lowest ongoing overhead once built: adding a new consuming account is mostly "attach to the TGW" | Scales well to hundreds of spokes with a TGW (peering scales far worse, quadratically in connection count) | Weak: a failure or a security incident in one service on the hub is reachable by, and potentially affects, every attached spoke, since they all share the same network fabric |
| Cross-account private service endpoints (PrivateLink) | Strongest: each consumer gets a narrow, individually-provisioned and individually-revocable connection to one specific service, with no broader network reachability into your account at all | Highest per-consumer overhead: each new consumer requires its own endpoint creation, acceptance, and (if using a custom domain) DNS (Domain Name System) verification | Scales in consumer count well (each one is independent and doesn't affect the others), but the per-consumer setup cost means it doesn't reduce the provider's administrative burden as consumer count grows | Strongest: an incident or outage in one exposed service has no path to affect a different service exposed via a different endpoint, since there's no shared network fabric between them |
| Federated service mesh spanning accounts | Moderate: service-to-service authentication and authorization (via mutual TLS, or mTLS, and the mesh's own policy layer) can be quite strong, but it depends entirely on every participating account correctly running and configuring the mesh's control plane and sidecars | Moderate to build, but introduces an ongoing shared dependency: the mesh's control plane (certificate issuance, service discovery, policy distribution) now spans every account, and a control-plane issue can affect service-to-service communication everywhere the mesh reaches | Scales reasonably in service count within a single mesh, but cross-account meshes add real complexity in certificate trust and service-discovery federation as account count grows | Moderate: individual service failures are isolated by the mesh's own circuit-breaking and retry policies, but a mesh control-plane failure is itself a new, cross-account failure mode the other two patterns simply don't have |
Reading the comparison. The hub-VPC pattern is the right choice when the "consumers" are actually all part of the same trust domain (different accounts within one organization, for operational rather than security reasons), because its weak boundary is an acceptable trade for its low overhead in that context. Cross-account PrivateLink is the right choice specifically at a true external trust boundary, exposing a service to genuinely separate organizations or to internal consumers who should have no broader network access, because the setup overhead buys a boundary strong enough to survive that kind of relationship. A federated service mesh earns its place mainly when the actual requirement is fine-grained, per-service authentication and policy (not just "can this account reach this network"), typically inside a single, sophisticated platform team's remit across their own multiple accounts, because taking on a shared control-plane dependency is a real cost that isn't justified by a simple reachability requirement alone.
Worked example
A platform team evaluates all three for exposing an internal "feature flag" service to 30 other teams' accounts within the same company. The hub-VPC pattern would work operationally (all 30 teams already peer into the shared-services hub for other reasons) but means any of the 30 teams' compromised workloads could potentially reach the feature-flag service's full network surface, not just the one endpoint it actually needs. PrivateLink would give each of the 30 teams a narrow, revocable connection to just that one service, but building and maintaining 30 individual endpoint-acceptance relationships for an internal, low-risk service is disproportionate overhead for the actual trust boundary being crossed (all 30 teams are already inside the same company's TGW). A federated service mesh, if the company doesn't already run one, is the most overhead of all three to introduce just for this one use case. The team picks the hub-VPC pattern, accepting its weaker boundary because the actual risk (an internal team over-reaching into a low-sensitivity feature-flag service) doesn't justify PrivateLink's overhead, but layers a dedicated security group on the feature-flag service in the hub restricting which spoke CIDR (Classless Inter-Domain Routing) ranges can reach it, narrowing the hub pattern's default weakness without the full cost of switching patterns entirely.
Trade-offs and pitfalls
The most common mistake is picking a pattern based on its reputation (PrivateLink "sounds more secure," so use it everywhere) rather than on whether the actual trust boundary being crossed justifies its overhead; applying PrivateLink's per-consumer setup cost to 30 internal teams that already share a trust domain is pure overhead with no corresponding security gain. The reverse mistake, using a shared hub VPC to expose a service across a genuine external trust boundary because it was cheaper to stand up, is the more dangerous direction to get wrong, since it under-protects exactly the boundary that most needs the stronger isolation PrivateLink provides.
Design a resilient Direct Connect (or ExpressRoute/Cloud Interconnect) solution that provides predictable bandwidth and high availability for enterprise workloads across two regions. Discuss using LAGs (link aggregation), redundant virtual interfaces, a Direct Connect Gateway (or equivalent), partner interconnects, use of VPN for failover, and routing and BGP configuration choices.
Sample Answer
Direct answer
Redundancy for a dedicated circuit like Direct Connect (DX) has to be built at two independent levels: multiple physical connections aggregated into a Link Aggregation Group (LAG) for capacity and single-link failure tolerance, and a second, geographically separate connection (ideally through a different DX location and even a different network provider) so a single facility failure doesn't take down connectivity entirely, with VPN (Virtual Private Network) retained purely as a last-resort failover path rather than the primary redundancy mechanism. Getting predictable bandwidth and low latency at real scale also depends on tuning the path all the way to the instances consuming it, not just the circuit itself.
Structured elaboration
Link Aggregation Groups. A LAG bundles multiple physical Direct Connect connections at the same location into what BGP (Border Gateway Protocol, the routing protocol Direct Connect uses to exchange routes) sees as a single logical link, combining their bandwidth and surviving the loss of any one physical connection without a BGP session drop. AWS caps how many physical connections can go into one LAG depending on port speed (fewer connections allowed at higher per-port speeds), so a LAG raises capacity and tolerates a single physical failure, but does not protect against the whole facility going down.
Redundant virtual interfaces and a Direct Connect Gateway. Each LAG terminates in private virtual interfaces (VIFs) that attach to a Direct Connect Gateway, which in turn can reach VPCs (Virtual Private Clouds) in multiple Regions from a single gateway, which is what makes a two-Region design practical without needing a completely separate DX setup per Region. For true resilience, provision two independent LAGs at two different physical DX locations, ideally via two different network providers on the local-loop side, so a fiber cut or facility outage affecting one location doesn't take out both paths.
Partner interconnects. Where a direct cross-connect into an AWS Direct Connect location isn't practical, a Direct Connect Partner provides the physical last-mile connection into that location on your behalf; using two different partners for the two redundant paths adds another layer of independence beyond just using two different physical locations.
VPN as failover, not primary redundancy. A Site-to-Site VPN over the internet, configured as a lower-preference route (via BGP attributes, discussed below) advertised alongside the Direct Connect paths, gives you a path that survives even a total loss of both DX paths simultaneously, at a much lower and less predictable bandwidth ceiling; it exists for the rare case where both physical paths are down, not as a substitute for the second physical path itself.
BGP configuration choices. With two DX paths (and a VPN backup), BGP path-selection attributes decide which path traffic prefers under normal conditions and how quickly it fails over: AS-path prepending (artificially repeating your own network's ID in the route's advertised path so it looks longer and less attractive to other routers) or local-preference values (a BGP setting that tells your own routers which available path to favor, regardless of path length) make the VPN path deliberately less attractive than either DX path so it's only used when both DX paths are actually down, and enabling Bidirectional Forwarding Detection (BFD) alongside BGP lets a path failure be detected and failed over in a small number of seconds rather than waiting on BGP's own, much slower, default hold-timer behavior.
Performance tuning beyond the circuit itself, for a genuinely high-bandwidth, low-latency design. The circuit and BGP configuration only get traffic to the VPC boundary; realizing the bandwidth on the compute side needs its own tuning. Use instances with the Elastic Network Adapter (ENA) for enhanced networking throughput. Configure jumbo frames (an MTU, or maximum transmission unit, of 9001 bytes instead of the default 1500) on the private virtual interface and on the instances terminating that traffic, since Direct Connect private VIFs support jumbo frames while VPN paths and public VIFs typically don't, meaning the two paths in this design may need to negotiate down to the lower MTU when failed over to VPN. Place latency-sensitive instances that consume the DX traffic in a cluster placement group (a placement strategy that packs instances physically close together on the same low-latency network hardware) to minimize inter-instance latency and maximize the network throughput between them. Where many instances need to share the inbound flow, a Network Load Balancer (NLB) distributes it at Layer 4 without adding meaningful latency of its own.
flowchart TB
subgraph OnPrem[On-premises data center]
R1[Router A]
R2[Router B]
end
R1 -->|Connection 1| LAG1[LAG at DX Location 1]
R2 -->|Connection 2| LAG2[LAG at DX Location 2, different provider]
LAG1 --> DXGW[Direct Connect Gateway]
LAG2 --> DXGW
DXGW --> VIF1[Private VIF to Region A VPC]
DXGW --> VIF2[Private VIF to Region B VPC]
R1 -->|BGP-deprioritized VPN failover| VPNGW[Site-to-Site VPN]
VPNGW --> DXGW
Worked example
An enterprise needs predictable, high-throughput, low-latency connectivity from an on-premises data center to workloads in two AWS Regions. They provision a two-connection LAG at DX location 1 and a second two-connection LAG at DX location 2 (a different physical facility, contracted through a different Direct Connect Partner for the local loop), both terminating in private VIFs attached to a single Direct Connect Gateway that reaches both Regions' VPCs. BGP is configured with local-preference values making both DX paths preferred over a Site-to-Site VPN path, which is present purely as a last-resort backup with BFD enabled on the DX BGP sessions for fast failure detection. On the compute side, the fleet consuming this traffic runs ENA-enabled instances in a cluster placement group with jumbo frames enabled end to end across the private VIFs, and an NLB in front distributes inbound flows across the fleet.
Trade-offs and pitfalls
Two LAGs at two different locations cost meaningfully more than one LAG, both in circuit fees and in the operational overhead of managing two physical relationships (potentially with two different partners); that cost is the direct price of surviving a facility-level failure, and skipping it (one LAG, however many connections bundled into it) leaves a single point of physical failure no amount of BGP tuning can route around. On the tuning side, jumbo frames only help if every hop in the path supports them consistently; enabling a 9001-byte MTU on the instances and the VIF while leaving an intermediate hop at 1500 bytes causes silent fragmentation or drops rather than an obvious error, so MTU has to be verified end to end, not just configured at the endpoints.
Unlock Full Question Bank
Get access to all 45 Cloud Networking and VPC Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.