Cloud Networking and VPC Design Questions
Designing networks inside a cloud provider: VPC/VNet topology, subnets, route tables, gateways, NAT, and peering, plus private connectivity through VPC endpoints and cloud load balancers. Covers segmentation, security groups and network ACLs, hybrid connectivity to on-premises data centers over VPN or dedicated links like Direct Connect and ExpressRoute, IP address planning across many VPCs and accounts, and how cloud network design differs from traditional data-center networking.
Compare and contrast instance-level security groups and subnet-level Network ACLs (NACLs) in cloud providers. Explain evaluation order, stateful vs stateless behavior, default rules and limits, and give examples of when to use each for a multi-tier application. Include an example where both are required and explain why.
Sample Answer
Direct answer
Security groups operate at the instance level (attached to the network interface) and are stateful: allow an inbound request and the matching response is automatically permitted back out, with no separate outbound rule needed. Network ACLs (NACLs) operate at the subnet level and are stateless: you must explicitly allow both directions of every conversation, including the ephemeral high-numbered ports used for return traffic, or replies get silently dropped. In a well-designed multi-tier application both layers are typically in play at once, security groups doing the fine-grained, per-tier access control and NACLs providing a coarser, second layer that holds even if a security group is ever misconfigured.
Structured elaboration
| Dimension | Security Group | Network ACL |
|---|---|---|
| Scope | Attached to individual network interfaces (effectively, per instance) | Attached to a subnet; applies to every resource in it |
| State | Stateful: return traffic for an allowed connection is automatically permitted | Stateless: inbound and outbound rules are evaluated independently; return traffic needs its own explicit rule |
| Rule evaluation | All rules are evaluated; there is no "first match wins" concept, only allow rules exist (nothing is explicitly denied, traffic simply isn't allowed if no rule matches) | Rules are evaluated in numbered order, lowest number first; the first rule that matches wins, and an explicit Deny is possible |
| Default behavior | A default security group denies all inbound and allows all outbound until you add rules | A default NACL allows all inbound and outbound traffic; a custom NACL you create denies everything until you add rules |
| Typical limits | Commonly around 60 rules per direction by default (adjustable), and a handful of security groups per network interface | Commonly around 20 rules per direction by default (adjustable up to around 40 per direction), evaluated in strict numeric order |
| Best used for | Fine-grained, per-tier or per-role access control ("only the web tier's security group may reach the app tier's security group on this port") | A coarse, subnet-wide boundary: blocking a known-bad IP range outright, or enforcing a hard compliance requirement that must hold regardless of what any individual security group says |
Why stateful vs stateless matters in practice, concretely. A client outside the VPC opens a connection to a web server: the request arrives from an ephemeral, randomly assigned high-numbered source port on the client side, destined for port 443 on the server, and the server's reply goes from port 443 back to that same ephemeral port on the client. A security group only needs one inbound rule (allow port 443 from the client's range) because it tracks the connection's state and automatically permits the matching reply out, no matter what port the reply uses. A NACL has no such memory: the inbound rule allowing port 443 says nothing about the outbound reply, so the NACL also needs an explicit outbound rule allowing traffic to the entire ephemeral port range (typically 1024 to 65535) back to the client, or every reply will be silently dropped at the subnet boundary even though the security group and the application are both configured correctly. This ephemeral-port gap is the single most common NACL misconfiguration: everything looks right, connections still fail, and the cause is invisible unless you specifically know to check for the missing return-traffic rule.
When to use each for a multi-tier application. Security groups do essentially all of the meaningful access control day to day: web tier's security group allows inbound 443 from the internet, application tier's security group allows inbound only from the web tier's security group, database tier's security group allows inbound only from the application tier's security group. NACLs are usually left at a permissive default in this pattern, and are reached for deliberately, not by default, when you need a control that survives a security-group mistake: for example, a subnet-wide rule blocking a specific IP range known to be malicious regardless of which instance or security group it targets, or a hard compliance boundary (say, "the database subnet must never accept inbound traffic from outside the VPC's own CIDR (its assigned IP address range, e.g. 10.0.0.0/16), full stop") that you want enforced even if someone someday attaches an overly permissive security group to something in that subnet by mistake.
Worked example
A case where both are genuinely required, and why. A financial services company runs its database tier in a private subnet and, separately from its normal tier-to-tier security groups, has a compliance requirement that the database subnet must categorically reject any traffic from outside the VPC's address range, independent of any application-level security group configuration, because a security group misconfiguration must never be sufficient on its own to expose the database externally. They implement this as a custom NACL on the database subnet with an explicit rule denying inbound traffic from any source outside the VPC's CIDR, positioned before (a lower rule number than) a broader allow rule for the VPC's own range, plus the normal outbound ephemeral-port allow rule for return traffic to the application tier. Independently, the database's security group allows inbound only from the application tier's specific security group on the database port, which is the layer actually doing meaningful per-tier access control day to day. The two layers answer different questions: the security group asks "is this specific source explicitly permitted," while the NACL asks "does this traffic even belong in this subnet at all," and having both means a mistake in one (an overly broad security-group rule accidentally added during a migration, for instance) does not by itself defeat the other.
Trade-offs and pitfalls
The most common mistake is trying to use NACLs for fine-grained, frequently changing access control, such as per-application-team rules: because NACL rules are evaluated in strict numeric order and typically capped at a modest number per direction, this quickly becomes an unmanageable, easy-to-misorder rule list, which is precisely the job security groups are built for instead. The second most common mistake, already covered above, is forgetting the ephemeral-port outbound rule when customizing a NACL, which silently breaks return traffic for entirely correctly configured applications and security groups. Finally, teams sometimes treat NACLs as unnecessary given that security groups exist, missing that the value of a NACL is specifically that it is a separate, subnet-wide layer that does not depend on any individual resource's security-group configuration being correct, which is exactly the property that makes it worth the extra maintenance for a genuinely hard boundary like the database example above.
Design the network connectivity for a PCI DSS scoped workload in AWS that must connect to external payment processors. Explain segmentation to minimize scope, options between Direct Connect and VPN for payment traffic, PrivateLink usage, ensuring encryption in transit, centralized logging and monitoring, and practical steps to keep non-PCI systems out of scope.
Sample Answer
Direct Answer
Segment the cardholder data environment (the CDE, the systems that store, process, or transmit primary account numbers) into its own subnets or account so the smallest possible set of systems carries PCI DSS (Payment Card Industry Data Security Standard) scope, reach the external payment processor over a Direct Connect circuit with IPsec (Internet Protocol Security, a protocol suite that encrypts and authenticates traffic between two network endpoints) or MACsec on top (or a plain Site-to-Site VPN if the processor only exposes a public endpoint), use PrivateLink (private AWS network endpoints that reach one specific service without touching the public internet) for any AWS-internal hop inside that boundary, encrypt everything in transit at the network layer AND the application layer, and centralize logging so you can prove, not just assert, that non-CDE systems cannot reach the CDE.
Segmentation and Access Design
CDE-subnet isolation. Put only the minimum necessary compute, typically the payment gateway or tokenization service, in dedicated CDE subnets. Everything else (web tier, general app tier that never touches a raw card number) lives in explicitly separate subnets, ideally a separate VPC or a separate AWS account inside an AWS Organization, with one documented, minimal, logged connection between them (usually one direction only: the app tier calls a tokenization API and never sees the card number itself). This is the move that actually shrinks audit scope, because PCI DSS scope is every system that can affect the CDE's security, not just systems that store card data.
Route tables and security groups as the enforcement layer. The CDE subnet's route table should have routes only to: the local VPC, the specific PrivateLink endpoints it needs, and the Direct Connect/VPN path to the processor. No route from any non-CDE subnet should point at the CDE subnet. Security groups on CDE instances should be default-deny with explicit, named allow rules; network ACLs at the CDE subnet boundary are a cheap secondary layer but are not, by themselves, what a Qualified Security Assessor (QSA) will accept as segmentation evidence if the CDE and non-CDE workloads still share a flat VPC and a permissive security group.
Bastion access without a bastion. A classic SSH jump box is itself an inbound path into the CDE and a standing credential a PCI DSS Requirement 7/8 assessor will scrutinize. Replace it with AWS Systems Manager Session Manager: no inbound security group rule, no public IP on the target, every session logged to CloudWatch or S3 as compliance evidence, and access gated by an IAM policy that only allows a temporary, just-in-time (JIT) role assumption scoped to an approved change window, rather than a permanent credential sitting on a bastion host.
Direct Connect vs VPN for Payment Traffic
Direct Connect (DX) is a private, dedicated circuit to AWS, but it is not encrypted by default. To satisfy PCI DSS's encryption-in-transit requirement over a network path outside your direct control, add either MACsec (available on certain dedicated DX connections, IEEE-standard data confidentiality and integrity at the link layer) or an IPsec VPN running over the DX virtual interface. DX buys predictable latency and bandwidth, which matters if the processor requires low-jitter, high-volume real-time authorization traffic, and many processors offer colocation at the same facilities as AWS Direct Connect locations.
A plain Site-to-Site VPN over the internet is IPsec-encrypted by default (typically IKEv2, the protocol that negotiates and sets up the encrypted session, paired with AES-GCM, the algorithm that then does the actual encryption), quick to provision, and the right choice when the processor only exposes a public API endpoint with no private-connectivity option. Its throughput is capped per tunnel (1.25 Gbps standard, up to 5 Gbps with Large Bandwidth Tunnels), and its latency is less predictable than a dedicated circuit.
Recommendation: for a processor with steady, high-volume traffic and a colocation or partner presence, commit to Direct Connect with IPsec or MACsec. For a processor reachable only over a public HTTPS endpoint, or for lower, less latency-sensitive volume, a Site-to-Site VPN is proportionate and faster to stand up. What would flip the choice: if the workload later needs sub-50ms, high-throughput authorization at sustained volume, the added operational cost of DX becomes worth it; for an occasional batch settlement file, it usually is not.
PrivateLink Usage
For any hop that stays inside AWS, calling a fraud-scoring service in another account, or reaching Amazon S3 for a tokenized batch file, use interface VPC endpoints (AWS PrivateLink) so the traffic never touches the public internet or a NAT gateway. This shrinks the network attack surface that counts toward PCI scope. If the payment processor itself publishes a PrivateLink endpoint service (some processors partner with AWS to offer exactly this), prefer it over any public path, since it removes the internet hop and public IP addressing entirely.
Encryption, Logging, and Keeping Non-PCI Systems Out of Scope
Encryption in transit happens at two layers that are easy to conflate: the network layer (IPsec over VPN, MACsec over DX) and the application layer (TLS 1.2 minimum, TLS 1.3 preferred, per PCI DSS 4.0). PrivateLink is private connectivity, not an encryption guarantee by itself, so terminate TLS at the application layer even over a PrivateLink path.
Centralized logging and monitoring means VPC Flow Logs enabled on every elastic network interface (ENI) and subnet in the CDE, shipped to a centralized log-archive account, retained to meet PCI DSS Requirement 10 (a full year, with the most recent three months immediately queryable). This is literal compliance evidence for the QSA, not just an operational nicety. Pair it with a Gateway Load Balancer (GWLB)-fronted intrusion detection appliance at the CDE boundary and centrally aggregated CloudTrail and GuardDuty findings.
Keeping non-PCI systems out of scope in practice means: no shared security group or NACL between CDE and non-CDE subnets, no route from non-CDE subnets into the CDE, dedicated IAM roles scoped by account or tag, a network diagram that stays current as part of every assessment, and periodic connectivity tests (a scan from a non-CDE subnet to the CDE should fail) run alongside quarterly Approved Scanning Vendor (ASV) scans so scope drift is caught rather than assumed away.
Worked Example
flowchart LR
subgraph NonCDE[Non-CDE VPC/account]
Web[Web tier 10.0.1.0/24]
App[App tier 10.0.5.0/24]
end
subgraph CDE[CDE subnet 10.0.100.0/24, no IGW route]
Gateway[Tokenization service]
end
LogAcct[Centralized log account: Flow Logs, CloudTrail]
Processor[Payment processor]
App -->|one documented API call, TLS| Gateway
Gateway -->|DX + IPsec/MACsec or VPN| Processor
Gateway -.->|Flow Logs, one-way| LogAcct
Web -->|no route exists| CDE
A VPC of 10.0.0.0/16 carries CDE subnets at 10.0.100.0/24 with no route to an internet gateway. Non-CDE subnets, 10.0.1.0/24 through 10.0.9.0/24, have zero route table entries pointing at 10.0.100.0/24, so the "no route in" claim is checkable directly against the route tables, not just asserted in a diagram. The CDE's only egress routes are to the PrivateLink endpoints it uses and to the Direct Connect or VPN attachment reaching the processor.
Trade-offs and Pitfalls
Putting the CDE in the same flat VPC as everything else and relying only on security groups is the most common failure. QSAs generally will not accept security-group-only isolation as PCI-grade segmentation, because a single misconfigured rule reopens the whole VPC; separate subnets, separate route tables, and ideally a separate account are what actually holds up.
Assuming PrivateLink traffic does not need TLS because "it's private" is a real and recurring finding. PCI DSS 4.0 requires strong cryptography for cardholder data in transit over any network, trusted or not.
A separate AWS account for the CDE gives the cleanest scope boundary (the IAM and Organizations boundary does more work than any network control) but adds operational overhead: cross-account networking, duplicated shared services. A single account with strict subnet and route isolation is cheaper to run and harder to defend to a QSA. Choose based on how much that operational cost is worth against how clean the scope story needs to be.
MACsec is the strongest transit control for Direct Connect but only works on specific dedicated connection speeds with a compatible customer router; verify availability before designing around it as the sole encryption layer.
Design a hardened bastion/access solution that eliminates inbound SSH from the internet, supports audit and session recording, and allows emergency access for on-call engineers. Compare options: AWS Systems Manager Session Manager, Azure Bastion, traditional bastion hosts with just-in-time (JIT) access, and third-party jump hosts. Describe IAM policies, ephemeral credentials, MFA, session logging, and a migration plan to roll out the safest option.
Sample Answer
Direct answer
Eliminate inbound SSH (Secure Shell) or RDP (Remote Desktop Protocol) entirely rather than trying to secure a bastion host that still has an open inbound port: use an agent-based, outbound-only connection broker (AWS Systems Manager Session Manager, Azure Bastion, or a comparable managed service) so the engineer authenticates through identity and access management (IAM) rather than a network path, and the target instance never has a listening port exposed to anything, including the "bastion" itself in the traditional sense.
Structured elaboration
| Option | Inbound port required | Session recording | Credential model | Emergency/on-call fit |
|---|---|---|---|---|
| AWS Systems Manager Session Manager | None (agent makes an outbound connection to the Systems Manager service) | Native: sessions can be logged to CloudWatch Logs or S3 | IAM policy grants session start on specific instances; no SSH key or password ever exists on the instance | Good: access is granted by adding an IAM policy, revocable instantly, no key distribution |
| Azure Bastion | None from the internet; Bastion is a managed PaaS (platform as a service) reachable only via the Azure portal or CLI over HTTPS | Native session logging available | Azure AD (Microsoft's identity and access management service, recently renamed Entra ID) authentication and role-based access control | Good: similarly identity-driven, no VM-level open port |
| Traditional bastion host with just-in-time (JIT) access | Yes, but only opened for a short, approved window per request | Depends on tooling layered on top (session recording usually bolted on via script or a proxy) | Often SSH keys or short-lived certificates issued per request | Workable, but the JIT approval step itself becomes a dependency during an incident |
| Third-party jump host / VPN appliance | Yes, generally always-open to authorized networks | Varies by vendor | Often a mix of shared credentials and MFA (multi-factor authentication), harder to keep fully per-user | Weakest fit here: usually the most operational overhead to keep current and audited |
IAM policies. Access is granted as an IAM policy statement scoping which instances (by tag or resource ARN, or Amazon Resource Name) a given role or user may start a session on, which replaces "who has the SSH key" with "who has the IAM permission," a model that's centrally auditable and instantly revocable by removing the policy, with no key rotation or distribution problem.
Ephemeral credentials. No long-lived SSH key ever needs to exist on the instance at all; the session broker (Systems Manager) authenticates the human via IAM and establishes the connection without a static credential ever being provisioned. This directly closes the most common real-world bastion failure mode: a leaked or never-rotated SSH private key granting standing access indefinitely.
Multi-factor authentication (MFA). Because access is gated by IAM sign-in rather than possession of a key, requiring MFA on the underlying identity (via an IAM policy condition or the organization's identity provider) applies uniformly to every session-broker connection, without needing a separate MFA integration bolted onto the bastion host itself.
Session logging. Every keystroke and output of a session can be streamed to CloudWatch Logs or an S3 bucket, giving a full audit trail per session, tied to the IAM identity that started it, which is materially stronger than a traditional bastion's SSH access log, which typically only shows that a connection happened, not what was done inside it.
Migration plan. Roll out in phases rather than a single cutover: install and validate the Systems Manager agent on a subset of instances first, grant IAM session-start permissions to a pilot group of engineers, and run the new path in parallel with the existing bastion for a defined period; once the pilot group confirms full workflow coverage (including any tooling that assumed direct SSH, like certain deployment scripts), remove inbound SSH/RDP security group rules for the migrated instances, and only then decommission the legacy bastion host itself, since removing it too early, before every workflow has an equivalent, forces engineers back to workarounds that reopen inbound access informally.
flowchart LR
Eng[On-call engineer] -->|IAM auth plus MFA| SSM[Session broker: Systems Manager Session Manager]
SSM -->|outbound-only agent connection, no open inbound port| Instance[Private EC2 instance]
SSM -->|session transcript| Log[CloudWatch Logs or S3]
SSM -->|session start and stop events| Audit[CloudTrail audit trail]
Worked example
An on-call engineer needs emergency access to a production database host at 2 a.m. With the session-broker model, they authenticate to the AWS console or CLI with their existing IAM identity (already requiring MFA), which already carries a break-glass IAM policy scoped to session-start on production instances tagged role=oncall-emergency, and start a session directly; no key to locate, no bastion IP to remember, no separate VPN client. The full session transcript lands in CloudWatch Logs automatically, and the security team's next-morning review shows exactly which commands were run, by whom, and for how long, without needing to correlate a bastion's SSH log against a separate change ticket.
Trade-offs and pitfalls
The session broker becomes a critical dependency itself: if the managed service or its agent has an outage, emergency access through that path is unavailable, so a genuine break-glass fallback (a tightly controlled, alarmed, rarely-used traditional path) is still worth keeping for that narrow case, rather than assuming the managed service is infallible. A second pitfall in migration: cutting over the humans but forgetting the automation, deployment scripts, configuration management tools, or monitoring agents that were quietly relying on direct SSH access, which breaks in ways that look unrelated to the bastion project until someone traces the failure back to the removed inbound rule.
You need to capture packet-level data for a short window to debug a production high-throughput network problem without impacting performance or incurring excessive storage cost. Propose an architecture using sampling, BPF-based filtering, traffic mirroring, ephemeral capture targets, and automation to trigger captures on detected anomalies. Explain how to secure captured data and automate lifecycle (start, stop, store, purge).
Sample Answer
Direct Answer
Keep a continuous, cheap sampled baseline running at all times using flow sampling and Berkeley Packet Filter (BPF)-based filters, so there is always some signal, then automatically trigger a short, full-fidelity Traffic Mirroring burst to an ephemeral capture target only when an anomaly detector fires, so full packet detail is available exactly when it is needed without the storage and processing cost of capturing everything all the time.
The Always-On Baseline: Sampling
Continuous, low-overhead flow sampling, VPC Flow Logs at minimum, or sampled flow export from an appliance or mesh that supports it, runs at all times. This is what the anomaly detector actually watches: a cheap, always-available signal for "something changed" without the cost of full packet capture running around the clock.
Narrowing What Gets Captured: BPF-Based Filtering
Whenever a capture is running, the continuous baseline or a triggered burst, a BPF filter, the same filter language standard packet-capture tools use and expressible as a Traffic Mirroring filter, narrows capture to exactly the flows implicated by the anomaly: a specific five-tuple (the combination of source IP, destination IP, source port, destination port, and protocol that identifies one connection), a specific port range, a specific instance's network interface, rather than the whole host's traffic. This keeps both the processing cost of capture and the resulting file volume proportional to the actual problem.
The Triggered Mechanism: Traffic Mirroring
On anomaly detection, programmatically create a Traffic Mirroring session, source the implicated elastic network interface, filter the narrowed rule from the previous step, target an ephemeral capture instance, via the API, rather than keeping a mirror session running permanently. This reuses the same primitive used for manual debugging, just automated and scoped to a detected event instead of a human-initiated one.
Ephemeral Capture Targets
The capture instance, or fleet, launches on demand from a pre-baked image specifically for the capture window and terminates afterward, rather than running a capture fleet continuously for an event that might happen rarely. This is the core cost lever that makes "high-throughput, short window" capture affordable, since the capture instance is paid for only during the anomaly, not continuously.
Automation to Trigger Captures on Detected Anomalies
An anomaly detector, an alarm on a flow-log-derived metric, or a custom detector watching for a retransmit-rate or error-rate spike, invokes an orchestrated workflow (useful here specifically because the process has several sequential, stateful steps rather than a single action) that: launches the ephemeral capture target; creates the Traffic Mirroring session pointed at it; waits a bounded window, long enough to catch a few cycles of the anomaly, short enough to bound cost, for example five to fifteen minutes; deletes the mirror session; triggers upload of the resulting capture files to object storage; and terminates the ephemeral target.
Securing Captured Data
Packet captures can contain unencrypted payloads or session tokens, so treat the destination storage as sensitive by default: default server-side encryption with a dedicated key rather than a shared default key, so access can be scoped and audited independently; a bucket policy that denies access outside a specific incident-response role; and object versioning or a write-once policy if a capture might become evidence for a security investigation that must not be alterable.
Automating the Lifecycle: Start, Stop, Store, Purge
The same orchestrated workflow owns start and stop. Storage goes to the encrypted, access-restricted bucket described above. Purge is the step most manual pipelines forget: a lifecycle rule that expires, or transitions to cold storage and then expires, raw capture files after a short, deliberately chosen retention window, long enough to complete the investigation, short enough to bound both cost and the exposure of holding sensitive captured payloads, kept separate from any sanitized findings an analyst chooses to retain longer.
Worked Example
flowchart LR
A[VPC Flow Logs baseline] --> B[Anomaly alarm]
B -->|fires| C[Orchestrated workflow]
C --> D[Launch ephemeral capture instance]
C --> E[Create Traffic Mirroring session: source=order-service ENI, filter=narrowed BPF rule]
D --> F[Bounded capture window, e.g. 10 min]
E --> F
F --> G[Delete mirror session]
F --> H[Upload capture to encrypted storage]
H --> I[Terminate ephemeral capture instance]
H --> J[Lifecycle rule purges raw capture after retention window]
A retransmit-rate anomaly on an order-service instance's network interface fires the alarm. The workflow launches a small ephemeral capture instance and a Traffic Mirroring session scoped only to that one interface, filtered to its TCP traffic rather than the whole subnet, captures for a bounded ten-minute window (long enough to catch several retransmit cycles without an open-ended cost commitment), tears down the mirror session and the instance, and lands the encrypted capture in storage, where a lifecycle rule purges it automatically after, say, fourteen days unless an analyst has already pulled it into a longer-lived investigation record.
Trade-offs and Pitfalls
Setting the anomaly threshold too sensitive triggers frequent, needless capture cycles, reintroducing exactly the cost problem this design exists to avoid; tune the detector against the flow-log baseline's normal variance first, do not guess a threshold.
Forgetting the purge step is the most common gap: a pipeline that automates start and stop but leaves storage manual tends to accumulate sensitive captures indefinitely, the opposite of the short-window, minimal-footprint goal.
An ephemeral, on-demand capture target adds a launch delay between anomaly detection and capture start, so the very first moment of a fast, transient anomaly might be missed; the always-on flow-sampling baseline is what covers that gap, since it does not depend on any launch.
A BPF or mirroring filter that is too broad reintroduces both cost and sensitive-data exposure; too narrow and it can filter out the very packets that explain the anomaly. Scope it from the same signal that triggered the detector, do not guess it independently.
Given a three-tier application (web, app, database) inside one VPC, propose the specific security group and network ACL rules that implement least privilege between the tiers. For each tier, specify the ports and traffic direction, and say whether you'd enforce it with a stateful security group or a stateless NACL and why.
Sample Answer
Direct Answer
Enforce least privilege between the three tiers with security groups that reference each other by security-group ID rather than IP range, since that survives autoscaling and re-addressing; keep network ACLs (NACLs) as a coarse, mostly-default backstop at the subnet boundary rather than trying to replicate the same fine-grained per-tier logic in a stateless, rule-numbered format.
Per-Tier Rules
| From (source) | To (destination) | Port / protocol | Direction | Enforced by | Why |
|---|---|---|---|---|---|
| Load balancer or internet | Web tier | 443/TCP (plus 80 for redirect only) | Inbound to web | Security group | Only the entry-point tier needs public or load-balancer-facing exposure |
| Web tier | App tier | Application's listening port, for example 8080/TCP | Outbound from web, inbound to app | Security group | Referencing the app tier's security-group ID, not a CIDR block (a range of IP addresses written like 10.0.0.0/16), means scaling web instances never requires a rule change |
| App tier | Database tier | Database port, for example 5432/TCP for PostgreSQL or 3306/TCP for MySQL | Outbound from app, inbound to database | Security group | Only the app tier, never the web tier and never the internet, may reach the database |
| Database tier | (none) | Default-deny outbound | Outbound | Security group | A database in this design never initiates outbound connections, so there is nothing legitimate for an open outbound rule to permit |
| Whole subnet | Explicitly known-bad CIDR ranges | All | Inbound | Network ACL | A coarse, subnet-wide blocklist backstop, not the primary access control |
Security Groups vs Network ACLs, and Why
Security groups are stateful: allowing an inbound request automatically allows its reply, and they attach per instance (technically per elastic network interface), which is exactly the granularity a three-tier, per-service rule set needs. They are the right layer for the tier-to-tier rules above.
Network ACLs are stateless: a reply to an allowed request needs its own explicit rule, and they attach per subnet, applying to everyone in it regardless of which specific instance sent or received the traffic. Trying to mirror the exact tier-to-tier port matrix in NACLs duplicates the security-group logic in a harder-to-maintain, rule-numbered format, and is a common source of a self-inflicted outage when someone inserts a new rule at the wrong priority and silently blocks a needed reply. NACLs are best reserved for what they are uniquely good at, a subnet-wide deny on a known-bad range, with a default allow for everything else, leaving the actual least-privilege enforcement in the stateful security groups.
If a NACL rule set is used for something finer anyway, for example to satisfy a compliance requirement for an explicit second layer, remember that a client-initiated TCP connection's return traffic arrives on an ephemeral port, typically in the 1024 to 65535 range, and must be explicitly allowed both inbound and outbound, unlike a stateful security group where allowing the request automatically allows the reply.
Trade-offs and Pitfalls
Writing NACL rules that duplicate the security-group matrix doubles the maintenance surface for no additional protection and is a real, recurring source of self-inflicted outages.
Referencing tiers by CIDR block instead of security-group ID is the most common mistake in a rule set like this: a CIDR-based rule breaks the moment the subnet is resized or an instance lands in a different Availability Zone, while a security-group-ID reference keeps working automatically as the fleet scales.
Security groups default to allowing all outbound traffic unless explicitly restricted; it is easy to forget this and leave a tier able to reach anywhere outbound, exactly the kind of gap a compromised instance would exploit.
A stricter default-deny-egress (blocking all outbound traffic by default, allowing only named exceptions) posture on every tier is more secure but adds ongoing maintenance, every new legitimate destination needs an explicit rule. Most teams accept that cost for the database tier, which rarely changes, and relax it somewhat for the app tier if it calls several external services, provided each destination is named explicitly rather than left open to everywhere.
Unlock Full Question Bank
Get access to all 16 Cloud Networking and VPC Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.