Cloud Networking and VPC Design Questions
Designing networks inside a cloud provider: VPC/VNet topology, subnets, route tables, gateways, NAT, and peering, plus private connectivity through VPC endpoints and cloud load balancers. Covers segmentation, security groups and network ACLs, hybrid connectivity to on-premises data centers over VPN or dedicated links like Direct Connect and ExpressRoute, IP address planning across many VPCs and accounts, and how cloud network design differs from traditional data-center networking.
Compare traditional bastion-host (jump box) patterns with agent-based approaches such as AWS Systems Manager Session Manager or Azure Bastion. Discuss differences in attack surface area, auditing and session recording, patching responsibilities, least privilege access, and recommend a secure, operationally maintainable pattern for remote administrator access.
Sample Answer
Direct answer
A traditional bastion host is a hardened jump box sitting in a public subnet with an open inbound port (typically SSH or RDP) that administrators connect to before hopping onward to private resources; it is itself a standing, internet-facing attack surface that has to be patched, monitored, and key-managed forever. An agent-based approach, such as AWS Systems Manager Session Manager or Azure Bastion, replaces that open inbound port with an agent on the target instance that opens an outbound-only connection to the cloud provider's session-broker service, so administrators authenticate through the same identity and access management (IAM) system as everything else and no inbound port needs to exist at all. For most organizations today, the agent-based approach is the right default: it is not just more convenient, it structurally removes an entire class of exposure the bastion pattern cannot avoid.
Structured elaboration
| Dimension | Traditional bastion host | Agent-based (Session Manager / Azure Bastion) |
|---|---|---|
| Attack surface | An always-on, internet-facing host with an open inbound port; a single compromised bastion is a pivot point into every private resource it can reach | No inbound port on any target instance; the connection is always initiated outbound from the instance to the provider's service, so there is no listening port for an attacker to find or exploit |
| Auditing and session recording | Depends entirely on what you build yourself: shell history, a session-recording tool bolted on, or nothing at all if not deliberately configured | Session logging and, optionally, full session recording, are built into the service and centrally stored, tied to the authenticated identity that started the session, with no extra tooling required |
| Patching responsibility | You own the bastion host's operating system, SSH/RDP daemon, and any hardening, indefinitely, as a standing piece of infrastructure | The provider manages the underlying service; you are only responsible for keeping the lightweight agent on each target instance up to date, which is a much smaller and more automatable surface |
| Least privilege | Access control is whatever the bastion's own user accounts and network rules enforce, often coarser than the cloud provider's IAM model and managed separately from it | Access is granted through the same fine-grained IAM policies as everything else in the account, so a specific engineer can be scoped to specific instances, for a specific time window, without a separate credential system to maintain |
| Credential model | Typically SSH keys or shared accounts, which have to be distributed, rotated, and revoked, often manually | Tied to the identity provider already in use; revoking someone's IAM access revokes their session access at the same time, with no separate key to rotate or revoke |
| Operational cost over time | Grows over time: patch cycles, key rotation, monitoring the bastion itself for compromise, capacity planning for a host that is really just a means to an end | Low and mostly fixed: the service itself is managed, and per-instance overhead is a small agent rather than an entire dedicated host |
Worked example
Recommended pattern. For a company running a fleet of private-subnet instances that administrators need occasional shell access to, the recommended pattern is to remove the bastion host entirely and rely on Session Manager (or the equivalent service on other clouds) for all administrative access, layered with the following hardening controls:
- Multi-factor authentication (MFA) enforced at the identity provider, not just a password, before a session can be started at all, since the session broker is only as strong as the identity check gating it.
- Session recording enabled and shipped to a separate, access-controlled log store (not just left in the default retention), so a compromised administrative identity's actions are still reviewable after the fact, and so the recording itself cannot be tampered with by whoever was using the compromised identity.
- Source restricted at the IAM policy level to only the specific instances and only the specific engineers who need them, scoped by tag or resource identifier rather than a blanket "any engineer can reach any instance" policy, which is the least-privilege principle applied to session access specifically rather than assumed to be handled elsewhere.
- No inbound security-group rule for SSH or RDP on any target instance at all, which is the concrete, verifiable proof the bastion's attack surface has actually been removed, not just supplemented; a lingering "just in case" SSH rule left open alongside Session Manager defeats much of the point.
A legacy environment still running a jump box during a transition can apply the same hardening principles to it in the meantime (restrict its source IP range as tightly as possible, require MFA for its own login, ship session logs somewhere durable), but this is explicitly a stopgap, not the target state, since none of that hardening removes the fundamental issue that the bastion is still an always-on, internet-facing host with an open inbound port.
Trade-offs and pitfalls
The most common mistake during a transition is standing up the agent-based approach alongside the existing bastion "just in case" and never actually decommissioning the bastion, which leaves the original attack surface fully intact while adding a second access path to manage and audit; the security benefit only materializes once the old inbound port is actually closed. A second pitfall is treating the agent-based service as a complete solution without also tightening the identity side of it: session-broker access inherits whatever IAM policy grants it, so an overly broad policy ("any authenticated user in the account may start a session to any instance") reproduces much of the bastion's own weak, coarse-grained access model, just without the open port. Finally, a legitimate reason organizations sometimes keep a hardened bastion around longer than expected is a genuine offline or air-gapped requirement, or a compliance framework that has not yet been updated to recognize agent-based session brokering as an equivalent control; in those cases, apply every hardening control above to the bastion itself rather than treating "we still have a bastion" as an excuse to skip them.
Given a three-tier application (web, app, database) inside one VPC, propose the specific security group and network ACL rules that implement least privilege between the tiers. For each tier, specify the ports and traffic direction, and say whether you'd enforce it with a stateful security group or a stateless NACL and why.
Sample Answer
Direct Answer
Enforce least privilege between the three tiers with security groups that reference each other by security-group ID rather than IP range, since that survives autoscaling and re-addressing; keep network ACLs (NACLs) as a coarse, mostly-default backstop at the subnet boundary rather than trying to replicate the same fine-grained per-tier logic in a stateless, rule-numbered format.
Per-Tier Rules
| From (source) | To (destination) | Port / protocol | Direction | Enforced by | Why |
|---|---|---|---|---|---|
| Load balancer or internet | Web tier | 443/TCP (plus 80 for redirect only) | Inbound to web | Security group | Only the entry-point tier needs public or load-balancer-facing exposure |
| Web tier | App tier | Application's listening port, for example 8080/TCP | Outbound from web, inbound to app | Security group | Referencing the app tier's security-group ID, not a CIDR block (a range of IP addresses written like 10.0.0.0/16), means scaling web instances never requires a rule change |
| App tier | Database tier | Database port, for example 5432/TCP for PostgreSQL or 3306/TCP for MySQL | Outbound from app, inbound to database | Security group | Only the app tier, never the web tier and never the internet, may reach the database |
| Database tier | (none) | Default-deny outbound | Outbound | Security group | A database in this design never initiates outbound connections, so there is nothing legitimate for an open outbound rule to permit |
| Whole subnet | Explicitly known-bad CIDR ranges | All | Inbound | Network ACL | A coarse, subnet-wide blocklist backstop, not the primary access control |
Security Groups vs Network ACLs, and Why
Security groups are stateful: allowing an inbound request automatically allows its reply, and they attach per instance (technically per elastic network interface), which is exactly the granularity a three-tier, per-service rule set needs. They are the right layer for the tier-to-tier rules above.
Network ACLs are stateless: a reply to an allowed request needs its own explicit rule, and they attach per subnet, applying to everyone in it regardless of which specific instance sent or received the traffic. Trying to mirror the exact tier-to-tier port matrix in NACLs duplicates the security-group logic in a harder-to-maintain, rule-numbered format, and is a common source of a self-inflicted outage when someone inserts a new rule at the wrong priority and silently blocks a needed reply. NACLs are best reserved for what they are uniquely good at, a subnet-wide deny on a known-bad range, with a default allow for everything else, leaving the actual least-privilege enforcement in the stateful security groups.
If a NACL rule set is used for something finer anyway, for example to satisfy a compliance requirement for an explicit second layer, remember that a client-initiated TCP connection's return traffic arrives on an ephemeral port, typically in the 1024 to 65535 range, and must be explicitly allowed both inbound and outbound, unlike a stateful security group where allowing the request automatically allows the reply.
Trade-offs and Pitfalls
Writing NACL rules that duplicate the security-group matrix doubles the maintenance surface for no additional protection and is a real, recurring source of self-inflicted outages.
Referencing tiers by CIDR block instead of security-group ID is the most common mistake in a rule set like this: a CIDR-based rule breaks the moment the subnet is resized or an instance lands in a different Availability Zone, while a security-group-ID reference keeps working automatically as the fleet scales.
Security groups default to allowing all outbound traffic unless explicitly restricted; it is easy to forget this and leave a tier able to reach anywhere outbound, exactly the kind of gap a compromised instance would exploit.
A stricter default-deny-egress (blocking all outbound traffic by default, allowing only named exceptions) posture on every tier is more secure but adds ongoing maintenance, every new legitimate destination needs an explicit rule. Most teams accept that cost for the database tier, which rarely changes, and relax it somewhat for the app tier if it calls several external services, provided each destination is named explicitly rather than left open to everywhere.
An EC2 in a private subnet with an S3 gateway endpoint is failing to access S3. Describe a systematic troubleshooting process to identify the root cause. Include checks for route tables, endpoint policy, IAM role permissions, NACLs, security groups, DNS resolution, and any cloud provider quirks you would examine.
Sample Answer
Direct answer
Work outward from "does the traffic even take the path I think it does" before touching identity and permissions. A gateway endpoint reaches Amazon S3 (Simple Storage Service) by injecting a special route into specific route tables, not through a DNS (Domain Name System) trick or a network interface, so the single most common root cause is a private subnet whose route table was never associated with the endpoint. Check that first, then work through the endpoint policy, the IAM (Identity and Access Management) role, network ACLs (access control lists), security groups, and DNS resolution in that order: a hang or timeout points at routing or filtering, while a fast, clean rejection points at a permissions layer.
Structured elaboration
- Route tables. A gateway endpoint (unlike an interface endpoint) has no ENI (elastic network interface) of its own. At creation time you explicitly associate it with one or more route tables, and AWS inserts a route whose destination is a managed prefix list (something like
pl-xxxxxxxxrepresenting "S3 in this Region") with the target set to the endpoint ID (vpce-xxxxxxxx). If the instance's subnet uses a different route table than the one the endpoint was associated with, that route simply doesn't exist there, and traffic falls through to whatever the default route is: a NAT (Network Address Translation) gateway if one exists (works, but silently bypasses the endpoint and starts costing data-processing fees), or nowhere if the subnet is fully private. Pull the route table for the exact subnet the instance is in and look for the prefix-list route before doing anything else. - Endpoint policy. The endpoint carries its own resource policy, separate from IAM. The default is "allow all actions on all resources," but many organizations tighten it to specific buckets or actions. If the target bucket was added to a bucket policy's allow list but never added to the endpoint policy (or vice versa), you get an accurate but confusing Access Denied.
- IAM role permissions. Check the instance profile's role for
s3:GetObject/s3:PutObject/s3:ListBucket(as applicable) on the correct resource ARN (Amazon Resource Name), then check for an explicit Deny from a service control policy or permission boundary. Separately, check the S3 bucket policy for anaws:sourceVpcecondition: if it's pinned to a different endpoint ID than the one actually serving this VPC (common once an org has more than one gateway endpoint), requests fail even with correct IAM. - NACLs. These are stateless, so the subnet's NACL needs an outbound rule allowing port 443 to the destination (S3's IP range or the same prefix list) and an inbound rule allowing the ephemeral return ports (commonly 1024 to 65535), evaluated independently in both directions. A NACL that allows the outbound request but not the ephemeral inbound reply produces a classic "request sent, response silently dropped" hang.
- Security groups. Stateful, so only the outbound rule matters (the return traffic is auto-allowed); confirm egress on 443 is actually open, since hardened baseline security groups often remove the default allow-all-outbound rule.
- DNS resolution. This is usually a red herring for a gateway endpoint specifically: the instance still resolves the normal public S3 hostname to its normal public IP range, and the route table's prefix-list route is what silently intercepts and redirects that traffic. Don't spend time here unless you confirm someone has also created an S3 interface endpoint (which does use a private hosted zone to override DNS) alongside the gateway endpoint, or that
enableDnsSupport/enableDnsHostnamesare disabled on the VPC, either of which can make DNS the actual culprit. - Cloud provider quirks.
- Endpoint routes are scoped per route table, not per VPC (Virtual Private Cloud). A new subnet you add later gets no S3 access through the endpoint until its route table is explicitly associated.
- Gateway endpoints are not reachable through VPC peering or a Transit Gateway attachment by default; a peered VPC's traffic to S3 does not automatically ride the endpoint route unless you deliberately route it through an appliance that has one.
- The endpoint only covers S3 in its own Region. A bucket in another Region is unaffected by the endpoint and goes out whatever egress path (NAT, proxy) is available, or fails if none exists.
Worked example
Instance sits in subnet subnet-0a1, associated with route table rtb-priv1. The gateway endpoint vpce-0abc was created and associated with rtb-priv2, a route table used by a different application tier. aws s3 ls from the instance hangs and times out rather than returning a 403. Tracing it: describe-route-tables --route-table-ids rtb-priv1 shows only 10.0.0.0/16 -> local and 0.0.0.0/0 -> nat-0123, no pl-xxxxxxxx -> vpce-0abc row; the same command against rtb-priv2 shows that row present. That confirms the endpoint was wired to the wrong route table. The fix is aws ec2 modify-vpc-endpoint --vpc-endpoint-id vpce-0abc --add-route-table-ids rtb-priv1 (or, in TerraForm, adding rtb-priv1 to the endpoint's route_table_ids list). Because this environment has no live AWS credentials, treat the CLI output above as illustrative of the diagnostic shape, not an executed transcript; the way to confirm the fix for real is to re-run the same describe-route-tables call and see the prefix-list route appear, then re-run the failing S3 call and check VPC Flow Logs for an ACCEPT record on the instance's ENI with a destination matching the S3 IP range.
Trade-offs and pitfalls
The instinct is to check IAM first because that's where errors are loudest, but the symptom should drive the order: a clean, fast Access Denied means IAM, the endpoint policy, or the bucket policy; a hang or timeout means routing, NACLs, or security groups. Also watch for a NAT gateway masking the real bug: if a default route exists, a missing endpoint route doesn't break the instance's S3 access, it just quietly routes it over the internet path instead, so "it works" doesn't mean the endpoint is configured correctly, only that you're now paying NAT data-processing charges for traffic that should have been free and private.
You need to explain the core components of a VPC to a junior admin: subnets, route tables, the internet gateway, the NAT gateway, security groups, and network ACLs. For each one, give a one-sentence description and a simple rule of thumb for when they'd need to change it.
Sample Answer
Direct Answer
Six pieces, each doing one job: subnets carve a Virtual Private Cloud (VPC), your own isolated slice of network address space in the cloud, into smaller ranges tied to one Availability Zone; route tables decide where each subnet's traffic is allowed to go; the internet gateway is the one door between a subnet and the public internet; the NAT gateway lets private resources reach out to the internet without letting the internet reach in; security groups are a per-instance list of who may talk to that instance; and network ACLs are a per-subnet list that applies to everyone in that subnet, regardless of instance.
The Six Pieces, One Sentence and One Rule of Thumb Each
Subnets. A subnet is a range of IP addresses, written in CIDR (Classless Inter-Domain Routing) notation such as 10.0.1.0/24, carved out of the VPC's overall range and pinned to one Availability Zone. Change or add a subnet when isolating something into its own Availability Zone for redundancy, or into its own tier with different routing needs, not just to organize instances cosmetically.
Route tables. A route table is the attached-to-a-subnet list of "if traffic is headed here, send it that way" rules. Change a route table when changing where a subnet's traffic is allowed to go, giving it a path to the internet gateway or to a VPN, not when changing who is allowed to send that traffic; that is a security group or NACL decision.
Internet gateway. The internet gateway is the single door AWS provides between the VPC and the public internet. Only subnets that genuinely need to be reachable from, or need to directly reach, the internet should have a route to it; without that route, nothing in a subnet is directly internet-facing no matter what else is configured.
NAT gateway. A NAT (Network Address Translation) gateway lets instances in a private subnet, one with no route to the internet gateway, initiate outbound connections, downloading a software update, for example, while keeping the internet unable to initiate connections back in. Add or resize a NAT gateway when a private subnet's workloads need outbound internet access, and check it first when an egress bill grows unexpectedly, since it is billed both hourly and by the gigabyte processed.
Security groups. A security group is a stateful, meaning it automatically allows the reply to traffic it already permitted, without a separate rule for the response, allow-list attached to an individual instance. Change a security group when deciding which other instances or IP ranges a specific instance should accept traffic from; this is usually the first and most common place to adjust access.
Network ACLs (NACLs). A NACL is a stateless, meaning the reply is not automatic and needs its own explicit rule, allow-and-deny list attached to a whole subnet and evaluated in numbered order. Reach for a NACL when a blanket rule is needed for an entire subnet regardless of instance, blocking a known-bad IP range for everyone in it, not as the everyday per-instance access control; that role belongs to security groups.
Worked Example: Tracing One Request
A laptop on the internet requests a web page. The request enters through the internet gateway. The destination subnet's route table has a route sending internet-bound traffic (0.0.0.0/0) to the internet gateway, confirming this subnet is meant to be reachable. The subnet's network ACL is checked first, the subnet-wide gate, and allows it. The specific web server instance's security group is checked next, the instance-specific gate, and allows port 443 from anywhere. The web server, sitting one tier back in a private subnet for its own database call, uses a NAT gateway to reach out for a software update, with its own security group allowing only that specific outbound destination.
Trade-offs and Pitfalls
The most common mistake a junior admin makes is assuming a security group alone is enough to keep something private. If the subnet's route table has a path to the internet gateway and the instance has a public IP address, the security group is the only thing standing between it and the internet, so one overly broad rule exposes it directly.
Forgetting that NACLs are stateless causes confusing outages: adding a NACL rule that allows inbound traffic but forgetting the matching outbound rule for the reply, on the ephemeral port range, makes the connection look like it is being silently dropped for no visible reason.
A NAT gateway is often confused with an internet gateway. The simplest way to keep them apart: the internet gateway is a two-way door, in and out, while the NAT gateway is one-way in intent, only outbound-initiated traffic and its replies, with nothing able to newly connect in through it.
What are VPC Flow Logs (or equivalent network flow logging) and why are they important? Explain what fields they record, where to send them for analysis (object store, SIEM), retention considerations, sampling/aggregation limitations, and common uses in troubleshooting, billing, and security monitoring.
Sample Answer
Direct answer
VPC (Virtual Private Cloud) Flow Logs capture metadata about the IP traffic crossing a network interface, subnet, or whole VPC: who talked to whom, on what port and protocol, how much data moved, and whether it was accepted or rejected. They matter because they're the closest thing to a network-level audit trail you get without running your own packet capture: without them, "why can't this instance reach that service" or "who talked to this compromised host" are questions you simply cannot answer after the fact.
Structured elaboration
What fields they record. The default (version 2) record format captures fourteen fields, in order: version, account-id, interface-id, srcaddr and dstaddr (source and destination IP), srcport and dstport, protocol (the IANA protocol number, 6 for TCP and 17 for UDP being the common ones), packets and bytes transferred, start and end timestamps for the aggregation window, action (ACCEPT or REJECT, meaning a security group or NACL, or network access control list, blocked it), and log-status. Later versions add optional fields you can request in a custom format, including the VPC ID, subnet ID, instance ID, TCP flags, and (more recently) the AWS service a source or destination IP belongs to.
Where to send them. Flow logs publish to Amazon CloudWatch Logs, Amazon S3 (Simple Storage Service), or Amazon Data Firehose. CloudWatch Logs suits real-time alerting and metric filters; S3 suits cheap long-term retention and offline analysis (commonly queried later with Amazon Athena); Firehose suits streaming the data into a third-party SIEM (Security Information and Event Management) system.
Retention. Flow logs themselves have no built-in expiry; retention is whatever you configure on the destination, a log group's retention setting in CloudWatch Logs or a lifecycle rule on the S3 bucket. Leaving this unset means logs (and their cost) accumulate indefinitely.
Sampling and aggregation limitations. Flow logs aggregate traffic over a window (the "aggregation interval"), by default up to 10 minutes, with an option to reduce it to 1 minute for finer granularity at higher log volume and cost; instances on the Nitro hypervisor (AWS's underlying virtualization layer, the software that creates and runs virtual machines on the physical host) always get an interval of 1 minute or less regardless of what you request. This means flow logs describe flows, not individual packets, so very short-lived connections can be summarized rather than itemized. Flow logs also don't capture everything: traffic to the Amazon-provided DNS resolver, DHCP (Dynamic Host Configuration Protocol) traffic, traffic to the instance metadata service at 169.254.169.254, and Windows license-activation traffic are all excluded by design, among a few other narrow exclusions.
Common uses. Troubleshooting connectivity (an unexpected run of REJECT records pinpoints whether a security group or a NACL is the blocker, and which one); security monitoring (unusual destination IPs, unexpected ports, or a spike in rejected connections can indicate scanning or a compromised host); and billing or cost analysis (correlating byte counts against data-transfer charges, especially useful for tracking down unexpectedly high cross-AZ (cross-Availability-Zone) or NAT (Network Address Translation) gateway egress costs).
Worked example
A team investigating a connectivity complaint enables flow logs on the affected subnet and searches the resulting CloudWatch Logs group for REJECT records involving the instance's ENI (elastic network interface) over the last hour. They find several records with action=REJECT, dstport=443, and a source address matching a partner service, dated exactly when the complaint started. Cross-referencing the security group and NACL rule sets shows a recent NACL change added a deny rule for a CIDR range that, as it turns out, also covers the partner's IP: the flow log record is the evidence that pinpoints which control (NACL, not the security group) is responsible, and exactly when the change took effect, since the REJECT records only start appearing after that timestamp.
Trade-offs and pitfalls
Flow logs record metadata, never packet payload, so they can tell you that two hosts talked and how much data moved, but never what was said; they are not a substitute for application-level logging or a payload-inspecting tool when you need to know the content of a request. The other common oversight is enabling flow logs without ever setting a retention policy on the destination, which turns a useful diagnostic tool into a slowly growing, unbounded storage cost that nobody is actually querying.
Unlock Full Question Bank
Get access to all 8 Cloud Networking and VPC Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.