Cloud Networking and VPC Design Questions
Designing networks inside a cloud provider: VPC/VNet topology, subnets, route tables, gateways, NAT, and peering, plus private connectivity through VPC endpoints and cloud load balancers. Covers segmentation, security groups and network ACLs, hybrid connectivity to on-premises data centers over VPN or dedicated links like Direct Connect and ExpressRoute, IP address planning across many VPCs and accounts, and how cloud network design differs from traditional data-center networking.
Explain how route tables work inside a VPC: how route tables are associated to subnets, the concept of local routes, longest-prefix-match (route precedence), propagating routes from virtual gateways (BGP), and how you would handle overlapping routes or conflicting routes when connecting to multiple external networks.
Sample Answer
Direct answer
A route table is an ordered set of rules, one per subnet association, that tells the network where to send traffic based on its destination IP address. Every route table gets one route it cannot delete, the "local" route for the VPC's own address range, which is why resources in the same VPC can always reach each other by default regardless of what else is configured. When more than one route could match a given destination, the network uses longest-prefix match: the most specific (largest prefix length) route wins, not the first one listed or the most recently added one.
Structured elaboration
Association. Each subnet is associated with exactly one route table (a route table can serve multiple subnets, but not the reverse). If a subnet has no explicit association, it uses the VPC's main route table by default. This is a common source of surprise: a newly created subnet silently inherits whatever the main route table says, which may not be what you intended for it.
The local route. The moment you create a VPC with, say, a 10.0.0.0/16 CIDR (Classless Inter-Domain Routing, the notation for writing an IP address range and its size) block, every route table in that VPC gets an implicit 10.0.0.0/16 -> local route that cannot be edited or removed. This is what makes intra-VPC communication work without any explicit configuration and is also why you cannot use a VPC's own address range for something else internally without addressing conflicts.
Longest-prefix match (route precedence). When a packet's destination matches more than one route in the table, the router picks the route with the longest (most specific) prefix, independent of the order the routes were created in. For example, if a table has both 0.0.0.0/0 -> internet gateway (catch-all, prefix length 0) and 10.0.0.0/16 -> local (prefix length 16), traffic to an address inside 10.0.0.0/16 matches the more specific /16 route even though the /0 route also technically matches everything. This is also how you can carve exceptions into a broad NAT-bound default route: add a more specific route for one destination CIDR that should instead go over a VPN or peering connection, and it wins over the general 0.0.0.0/0 route without touching that catch-all rule.
Route propagation from virtual gateways (BGP). When a Virtual Private Gateway or Transit Gateway is attached to a VPC and used with a Border Gateway Protocol (BGP) session over a VPN or a dedicated circuit, the on-premises router advertises its reachable networks over that BGP session, and if you enable route propagation on the route table, those advertised networks are added to the route table automatically instead of you hand-entering static routes. This matters operationally: if the on-premises network adds a new subnet, propagation means your VPC route table updates itself the next time BGP re-converges, with no manual change needed on the cloud side.
Overlapping or conflicting routes from multiple external networks. If you connect to two external networks (say, two data centers, or a data center and a partner network over separate VPN connections) and their advertised CIDR blocks overlap, or if a statically entered route and a BGP-propagated route both cover the same destination, you have to resolve the ambiguity explicitly, because longest-prefix match only helps when the prefixes actually differ in specificity:
- If one advertised range is a strict subset of the other, longest-prefix match naturally sends traffic for the subset to whichever path advertised the more specific range, so no conflict actually exists even though it looks like one on paper.
- If the two ranges genuinely overlap at the same specificity (both advertise the same
/24, for instance), most cloud route tables will simply reject adding the second, identical-specificity route, or, on platforms that support route priority, you set an explicit static route with a fixed destination as the tie-breaker. The durable fix, though, is address planning: two networks that both intend to be reachable from the same VPC should never have been assigned overlapping CIDR blocks in the first place, since even where the platform lets you work around it, only one of the two networks is actually reachable at overlapping addresses at any moment. - For BGP-specific tie-breaking (both paths are valid but you want a preference, not an outright conflict), use attributes like AS-path prepending (making a path look longer and therefore less preferred) or a local-preference / route-priority setting supported by the platform, rather than fighting it with static routes that will silently go stale.
Worked example
A VPC with CIDR 10.0.0.0/16, an Internet Gateway (IGW), and a Virtual Private Gateway (VGW) with route propagation enabled to an on-premises network at 192.168.0.0/16:
| Destination | Target | Where it came from | Why it wins for matching traffic |
|---|---|---|---|
10.0.0.0/16 | local | Implicit, undeletable | Always most specific for intra-VPC traffic |
192.168.0.0/16 | vgw-xxxx | BGP-propagated from on-prem | Longest match for that on-prem range |
192.168.5.0/24 | vgw-xxxx (different tunnel) | Static, entered manually to force a specific tunnel | Longer prefix than the propagated /16 above, so this specific subnet always takes this path even if the general /16 route pointed elsewhere |
0.0.0.0/0 | igw-xxxx | Static, manual | Catch-all; only matches what nothing more specific matches |
Here, the manually entered /24 route deliberately overrides part of the broader propagated /16 route for one subnet, a legitimate and common pattern for routing one sensitive on-premises subnet over a dedicated path while everything else in that address range uses the general one.
Internet Gateway vs Virtual Private Gateway, since they are easy to conflate: an Internet Gateway is a horizontally scaled, highly available AWS-managed component that gives a VPC a path to and from the public internet; it has no concept of BGP or on-premises networks. A Virtual Private Gateway is the VPN/Direct Connect-side endpoint on the AWS side of a hybrid connection; it terminates VPN tunnels or a Direct Connect (AWS's dedicated, private network circuit, as opposed to a VPN's tunnel over the public internet) virtual interface and is what participates in BGP route propagation with an on-premises router. They serve two entirely different paths (public internet vs private hybrid connectivity) and a VPC commonly has both at once, each referenced by different routes in the same route table.
Trade-offs and pitfalls
The most common mistake is assuming route order in a table or console listing matters, then getting confused when a "later" broad route does not override an earlier specific one; only prefix length decides precedence, never entry order. A second pitfall is enabling route propagation without auditing what gets propagated: an on-premises network that advertises a default route (0.0.0.0/0) over BGP will silently start competing with your Internet Gateway's own default route for precedence (resolved by longest-prefix match rules that still apply, but the practical effect, if the propagated route wins for some traffic class, can be egress traffic unexpectedly hairpinning through on-premises). Finally, when connecting to multiple external networks, resolve address-space overlap at design time through CIDR planning, not at incident time through static-route workarounds; a static override fixes today's symptom and quietly becomes tomorrow's undocumented special case.
Recommend an approach for organizing Security Groups at scale: per-application, per-tier, or per-environment. For each approach, explain pros/cons regarding manageability, least privilege, rule explosion, automation, and operational impacts in a large organization.
Sample Answer
Direct answer
Organize security groups (SGs) around the application tier they protect, not around the environment or the individual application, and layer host-based firewalls and a Web Application Firewall (WAF) in front of the public tier as defense in depth rather than relying on security groups alone; a per-tier scheme is the one that scales in a large organization without either drowning teams in rule sprawl or collapsing least privilege into one giant permissive group.
Structured elaboration
| Approach | Manageability | Least privilege | Rule explosion risk | Automation fit | Operational impact at scale |
|---|---|---|---|---|---|
| Per-application | Clear ownership, easy to reason about one app at a time | Strong, since each SG only allows what that specific app needs | High: hundreds of applications means hundreds of near-duplicate SGs, and shared infrastructure (a common database tier, a shared cache) ends up referenced by every one of them | Works well if SG creation is fully templated per app, but the sheer count strains tooling and review at scale | Every new application is a new SG to review and approve, and a security audit has to walk hundreds of near-identical rule sets one at a time to confirm none has quietly drifted |
| Per-tier (web, application, data, management) | Moderate: a smaller, stable number of SGs that map to a mental model everyone shares | Strong, if each tier's SG only allows the specific ports needed from the tier immediately in front of it | Low to moderate: the number of SGs stays roughly constant as the number of applications grows, since new apps reuse the existing tier structure | Best fit for automation: a new application in the "application tier" just gets attached to the existing application-tier SG rather than needing a bespoke one | Onboarding a new application is a one-line attach, not a new review, and a security audit only has to walk a handful of stable, well-understood rule sets regardless of how many applications sit behind them |
| Per-environment (dev, staging, prod) | Simple at a glance, but conflates unrelated concerns | Weak on its own: an environment-wide SG tends toward "anything in this environment can talk to anything else in this environment" | Low SG count, but each SG becomes broad and permissive, which is the opposite of least privilege | Easy to automate, but automating a bad access-control shape doesn't make it a good one | Low day-to-day overhead, but that low overhead is a symptom of the SG doing almost no real access control; any incident inside one environment has a wide blast radius across every application sharing that environment's SG |
Recommendation and reasoning. Per-tier is the right default at scale because it decouples the number of security groups from the number of applications, which is the dimension that actually grows uncontrolled in a large organization; per-application least-privilege intentions are good but the SG count grows linearly with every new service, and per-environment collapses tiers together in a way that violates least privilege as soon as you ask "should the web tier really be able to reach the data tier's admin port." In practice, combine per-tier SGs as the backbone (web, application, data, management/bastion) with narrow, per-application SGs layered on top only where a specific application genuinely needs an exception the tier-level rule doesn't cover, rather than choosing one pattern exclusively.
Host-based firewalls as defense in depth. Security groups operate at the ENI (elastic network interface) level and only see IP, port, and protocol; a host-based firewall (iptables/nftables on Linux, Windows Firewall) running on the instance itself adds a second, independent enforcement layer that still protects the instance even if a security group is ever misconfigured too permissively, and can enforce finer-grained policy (per-process rules, for instance) that security groups have no visibility into at all. This isn't redundant work, it's the standard defense-in-depth principle: two independent layers fail independently, so a single mistake in one layer doesn't fully expose the instance.
WAF in front of the public tier. For any tier actually facing the internet, a Web Application Firewall sitting in front of the load balancer adds a layer that security groups structurally cannot provide: security groups only see IP, port, and protocol, never the content of a request, while a WAF inspects HTTP-layer payloads for attack patterns (SQL injection, cross-site scripting) and can rate-limit or block based on request content and behavior. The public-tier security group still does its job (restricting which ports are even reachable), and the WAF does a job no security group configuration could ever do, inspecting what's actually inside the allowed traffic.
Worked example
A large organization with 200 microservices organizes security groups into four tier-level groups: sg-web (allows 443 inbound from the internet or from a WAF-fronted load balancer only), sg-app (allows the application port inbound only from sg-web, referenced by security group ID rather than a CIDR range), sg-data (allows the database port inbound only from sg-app), and sg-mgmt (allows administrative access only from the session-broker service, not from the internet). A new microservice being onboarded is simply attached to sg-app and, if it talks to the shared data tier, inherits access through sg-data's existing rule referencing sg-app, with no new security group needed for the common case; the team only creates a bespoke, narrower SG on top if this particular service needs an exception (say, a webhook receiver that needs a different inbound port than the rest of the application tier).
Trade-offs and pitfalls
The per-tier model's main risk is over-trusting the tier boundary itself: referencing sg-app as an allowed source in sg-data's rule is much better than allowing a broad CIDR range, but it still means every application in the tier can reach the data tier's port, so a compromised low-value service in the application tier has a path toward the data tier that a stricter, per-application rule would have prevented. The mitigation is exactly the "layer per-application exceptions on top" and "add host-based firewalls and, on the public edge, a WAF" guidance above: no single layer, including a well-organized security group scheme, is sufficient on its own at this scale.
For a compliance-heavy environment, describe how you would restrict and monitor outbound egress traffic from private subnets, including the use of NAT gateway, centralized proxy, firewall rules, and logging. Explain pros/cons of forcing egress through a single inspection point.
Sample Answer
Direct answer
For a compliance-heavy environment, the requirement is usually not just "block bad outbound traffic" but "prove every outbound connection was inspected and can be reconstructed later," which means forcing all egress from private subnets through one controlled, logged inspection point rather than letting each subnet or account exit independently. The trade-off you're explicitly accepting is a single choke point that adds latency, cost, and a scaling bottleneck, in exchange for the auditability a compliance program actually needs.
Structured elaboration
The building blocks. A NAT (Network Address Translation) gateway still handles the address translation itself, but instead of sitting directly in each workload's own path to the internet, it sits behind (or alongside) a centralized proxy or firewall tier that every subnet's outbound traffic is routed through, typically via a shared inspection VPC (Virtual Private Cloud) reached over a Transit Gateway (TGW). A centralized forward proxy (or a firewall appliance like AWS Network Firewall or a third-party equivalent, fronted by a Gateway Load Balancer for scale) applies domain and IP allowlisting or denylisting, TLS (Transport Layer Security) inspection where policy requires it, and logs every connection attempt, allowed or blocked. Flow logs and, where applicable, proxy access logs are shipped to a retained, tamper-evident log store, since "we had a firewall" is a much weaker compliance answer than "here is the log of every connection this workload made in the last 12 months."
Pros of forcing all egress through one inspection point. A single, well-audited enforcement point is dramatically easier to certify and reason about than dozens of independent NAT gateways each with their own, potentially drifting, rule set; a compliance auditor can review one policy and one log stream rather than reconciling configuration across every account. It also makes detection easier: a single place to watch for anomalous destinations or volumes across the entire estate, rather than needing to aggregate signals from many independent egress points after the fact.
Cons of forcing all egress through one inspection point. Every workload's outbound traffic now takes an extra hop, adding latency that's real even if often small, and concentrating both the traffic volume and the cost (data-processing charges accrue at the centralized NAT/firewall tier, at a scale proportional to the whole estate's egress, not any one workload's). It also creates a single point that must be sized and made highly available correctly, because an outage there doesn't just degrade one workload's egress, it degrades every workload's egress across the estate simultaneously; and it becomes an organizational bottleneck if every new outbound destination a team needs requires a change request against a shared, centrally-owned firewall rule set, which can slow legitimate work if the change process isn't kept fast.
What a compliance program typically requires beyond the mechanism itself. Documented change control on the firewall/proxy rule set (who approved adding a new allowed destination, and why), retained logs meeting the specific regulation's retention window, and periodic access review confirming the rule set still reflects only what's currently needed, since egress rules tend to accumulate permissive exceptions over time if nobody prunes them.
Worked example
A financial services company subject to a regulatory requirement to log and control all outbound network connections routes every workload VPC's 0.0.0.0/0 traffic through a shared inspection VPC via TGW. The inspection VPC runs a proxy that allowlists specific external domains (a payment processor, a small number of SaaS (Software as a Service) vendors, and nothing else by default) and logs every connection attempt with the requesting workload's identity, destination, and outcome. When an application team needs a new external dependency, they submit a change request to add the domain to the allowlist rather than being able to reach arbitrary internet destinations by default; the resulting log, retained for the regulation's required period, becomes the artifact the compliance team hands to an auditor as evidence that outbound connectivity is both controlled and observable.
Trade-offs and pitfalls
The single inspection point is the correct design for the stated compliance goal, but treating it as a "set it once" control invites rule-set rot: allowlist entries added under time pressure during an incident and never revisited accumulate into a broad, effectively-uncontrolled allowlist that technically satisfies "we have an egress control" while defeating its purpose. Sizing the shared inspection tier for peak aggregate load across the whole estate, not just its current load, matters too: because every workload depends on the same choke point, under-provisioning it turns a routine traffic spike in one team's workload into a shared outage for every team behind it.
You need to explain the core components of a VPC to a junior admin: subnets, route tables, the internet gateway, the NAT gateway, security groups, and network ACLs. For each one, give a one-sentence description and a simple rule of thumb for when they'd need to change it.
Sample Answer
Direct Answer
Six pieces, each doing one job: subnets carve a Virtual Private Cloud (VPC), your own isolated slice of network address space in the cloud, into smaller ranges tied to one Availability Zone; route tables decide where each subnet's traffic is allowed to go; the internet gateway is the one door between a subnet and the public internet; the NAT gateway lets private resources reach out to the internet without letting the internet reach in; security groups are a per-instance list of who may talk to that instance; and network ACLs are a per-subnet list that applies to everyone in that subnet, regardless of instance.
The Six Pieces, One Sentence and One Rule of Thumb Each
Subnets. A subnet is a range of IP addresses, written in CIDR (Classless Inter-Domain Routing) notation such as 10.0.1.0/24, carved out of the VPC's overall range and pinned to one Availability Zone. Change or add a subnet when isolating something into its own Availability Zone for redundancy, or into its own tier with different routing needs, not just to organize instances cosmetically.
Route tables. A route table is the attached-to-a-subnet list of "if traffic is headed here, send it that way" rules. Change a route table when changing where a subnet's traffic is allowed to go, giving it a path to the internet gateway or to a VPN, not when changing who is allowed to send that traffic; that is a security group or NACL decision.
Internet gateway. The internet gateway is the single door AWS provides between the VPC and the public internet. Only subnets that genuinely need to be reachable from, or need to directly reach, the internet should have a route to it; without that route, nothing in a subnet is directly internet-facing no matter what else is configured.
NAT gateway. A NAT (Network Address Translation) gateway lets instances in a private subnet, one with no route to the internet gateway, initiate outbound connections, downloading a software update, for example, while keeping the internet unable to initiate connections back in. Add or resize a NAT gateway when a private subnet's workloads need outbound internet access, and check it first when an egress bill grows unexpectedly, since it is billed both hourly and by the gigabyte processed.
Security groups. A security group is a stateful, meaning it automatically allows the reply to traffic it already permitted, without a separate rule for the response, allow-list attached to an individual instance. Change a security group when deciding which other instances or IP ranges a specific instance should accept traffic from; this is usually the first and most common place to adjust access.
Network ACLs (NACLs). A NACL is a stateless, meaning the reply is not automatic and needs its own explicit rule, allow-and-deny list attached to a whole subnet and evaluated in numbered order. Reach for a NACL when a blanket rule is needed for an entire subnet regardless of instance, blocking a known-bad IP range for everyone in it, not as the everyday per-instance access control; that role belongs to security groups.
Worked Example: Tracing One Request
A laptop on the internet requests a web page. The request enters through the internet gateway. The destination subnet's route table has a route sending internet-bound traffic (0.0.0.0/0) to the internet gateway, confirming this subnet is meant to be reachable. The subnet's network ACL is checked first, the subnet-wide gate, and allows it. The specific web server instance's security group is checked next, the instance-specific gate, and allows port 443 from anywhere. The web server, sitting one tier back in a private subnet for its own database call, uses a NAT gateway to reach out for a software update, with its own security group allowing only that specific outbound destination.
Trade-offs and Pitfalls
The most common mistake a junior admin makes is assuming a security group alone is enough to keep something private. If the subnet's route table has a path to the internet gateway and the instance has a public IP address, the security group is the only thing standing between it and the internet, so one overly broad rule exposes it directly.
Forgetting that NACLs are stateless causes confusing outages: adding a NACL rule that allows inbound traffic but forgetting the matching outbound rule for the reply, on the ephemeral port range, makes the connection look like it is being silently dropped for no visible reason.
A NAT gateway is often confused with an internet gateway. The simplest way to keep them apart: the internet gateway is a two-way door, in and out, while the NAT gateway is one-way in intent, only outbound-initiated traffic and its replies, with nothing able to newly connect in through it.
Given a three-tier application (web, app, database) inside one VPC, propose the specific security group and network ACL rules that implement least privilege between the tiers. For each tier, specify the ports and traffic direction, and say whether you'd enforce it with a stateful security group or a stateless NACL and why.
Sample Answer
Direct Answer
Enforce least privilege between the three tiers with security groups that reference each other by security-group ID rather than IP range, since that survives autoscaling and re-addressing; keep network ACLs (NACLs) as a coarse, mostly-default backstop at the subnet boundary rather than trying to replicate the same fine-grained per-tier logic in a stateless, rule-numbered format.
Per-Tier Rules
| From (source) | To (destination) | Port / protocol | Direction | Enforced by | Why |
|---|---|---|---|---|---|
| Load balancer or internet | Web tier | 443/TCP (plus 80 for redirect only) | Inbound to web | Security group | Only the entry-point tier needs public or load-balancer-facing exposure |
| Web tier | App tier | Application's listening port, for example 8080/TCP | Outbound from web, inbound to app | Security group | Referencing the app tier's security-group ID, not a CIDR block (a range of IP addresses written like 10.0.0.0/16), means scaling web instances never requires a rule change |
| App tier | Database tier | Database port, for example 5432/TCP for PostgreSQL or 3306/TCP for MySQL | Outbound from app, inbound to database | Security group | Only the app tier, never the web tier and never the internet, may reach the database |
| Database tier | (none) | Default-deny outbound | Outbound | Security group | A database in this design never initiates outbound connections, so there is nothing legitimate for an open outbound rule to permit |
| Whole subnet | Explicitly known-bad CIDR ranges | All | Inbound | Network ACL | A coarse, subnet-wide blocklist backstop, not the primary access control |
Security Groups vs Network ACLs, and Why
Security groups are stateful: allowing an inbound request automatically allows its reply, and they attach per instance (technically per elastic network interface), which is exactly the granularity a three-tier, per-service rule set needs. They are the right layer for the tier-to-tier rules above.
Network ACLs are stateless: a reply to an allowed request needs its own explicit rule, and they attach per subnet, applying to everyone in it regardless of which specific instance sent or received the traffic. Trying to mirror the exact tier-to-tier port matrix in NACLs duplicates the security-group logic in a harder-to-maintain, rule-numbered format, and is a common source of a self-inflicted outage when someone inserts a new rule at the wrong priority and silently blocks a needed reply. NACLs are best reserved for what they are uniquely good at, a subnet-wide deny on a known-bad range, with a default allow for everything else, leaving the actual least-privilege enforcement in the stateful security groups.
If a NACL rule set is used for something finer anyway, for example to satisfy a compliance requirement for an explicit second layer, remember that a client-initiated TCP connection's return traffic arrives on an ephemeral port, typically in the 1024 to 65535 range, and must be explicitly allowed both inbound and outbound, unlike a stateful security group where allowing the request automatically allows the reply.
Trade-offs and Pitfalls
Writing NACL rules that duplicate the security-group matrix doubles the maintenance surface for no additional protection and is a real, recurring source of self-inflicted outages.
Referencing tiers by CIDR block instead of security-group ID is the most common mistake in a rule set like this: a CIDR-based rule breaks the moment the subnet is resized or an instance lands in a different Availability Zone, while a security-group-ID reference keeps working automatically as the fleet scales.
Security groups default to allowing all outbound traffic unless explicitly restricted; it is easy to forget this and leave a tier able to reach anywhere outbound, exactly the kind of gap a compromised instance would exploit.
A stricter default-deny-egress (blocking all outbound traffic by default, allowing only named exceptions) posture on every tier is more secure but adds ongoing maintenance, every new legitimate destination needs an explicit rule. Most teams accept that cost for the database tier, which rarely changes, and relax it somewhat for the app tier if it calls several external services, provided each destination is named explicitly rather than left open to everywhere.
Unlock Full Question Bank
Get access to all 36 Cloud Networking and VPC Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.