Cloud Networking and VPC Design Questions
Designing networks inside a cloud provider: VPC/VNet topology, subnets, route tables, gateways, NAT, and peering, plus private connectivity through VPC endpoints and cloud load balancers. Covers segmentation, security groups and network ACLs, hybrid connectivity to on-premises data centers over VPN or dedicated links like Direct Connect and ExpressRoute, IP address planning across many VPCs and accounts, and how cloud network design differs from traditional data-center networking.
Design a hardened bastion/access solution that eliminates inbound SSH from the internet, supports audit and session recording, and allows emergency access for on-call engineers. Compare options: AWS Systems Manager Session Manager, Azure Bastion, traditional bastion hosts with just-in-time (JIT) access, and third-party jump hosts. Describe IAM policies, ephemeral credentials, MFA, session logging, and a migration plan to roll out the safest option.
Sample Answer
Direct answer
Eliminate inbound SSH (Secure Shell) or RDP (Remote Desktop Protocol) entirely rather than trying to secure a bastion host that still has an open inbound port: use an agent-based, outbound-only connection broker (AWS Systems Manager Session Manager, Azure Bastion, or a comparable managed service) so the engineer authenticates through identity and access management (IAM) rather than a network path, and the target instance never has a listening port exposed to anything, including the "bastion" itself in the traditional sense.
Structured elaboration
| Option | Inbound port required | Session recording | Credential model | Emergency/on-call fit |
|---|---|---|---|---|
| AWS Systems Manager Session Manager | None (agent makes an outbound connection to the Systems Manager service) | Native: sessions can be logged to CloudWatch Logs or S3 | IAM policy grants session start on specific instances; no SSH key or password ever exists on the instance | Good: access is granted by adding an IAM policy, revocable instantly, no key distribution |
| Azure Bastion | None from the internet; Bastion is a managed PaaS (platform as a service) reachable only via the Azure portal or CLI over HTTPS | Native session logging available | Azure AD (Microsoft's identity and access management service, recently renamed Entra ID) authentication and role-based access control | Good: similarly identity-driven, no VM-level open port |
| Traditional bastion host with just-in-time (JIT) access | Yes, but only opened for a short, approved window per request | Depends on tooling layered on top (session recording usually bolted on via script or a proxy) | Often SSH keys or short-lived certificates issued per request | Workable, but the JIT approval step itself becomes a dependency during an incident |
| Third-party jump host / VPN appliance | Yes, generally always-open to authorized networks | Varies by vendor | Often a mix of shared credentials and MFA (multi-factor authentication), harder to keep fully per-user | Weakest fit here: usually the most operational overhead to keep current and audited |
IAM policies. Access is granted as an IAM policy statement scoping which instances (by tag or resource ARN, or Amazon Resource Name) a given role or user may start a session on, which replaces "who has the SSH key" with "who has the IAM permission," a model that's centrally auditable and instantly revocable by removing the policy, with no key rotation or distribution problem.
Ephemeral credentials. No long-lived SSH key ever needs to exist on the instance at all; the session broker (Systems Manager) authenticates the human via IAM and establishes the connection without a static credential ever being provisioned. This directly closes the most common real-world bastion failure mode: a leaked or never-rotated SSH private key granting standing access indefinitely.
Multi-factor authentication (MFA). Because access is gated by IAM sign-in rather than possession of a key, requiring MFA on the underlying identity (via an IAM policy condition or the organization's identity provider) applies uniformly to every session-broker connection, without needing a separate MFA integration bolted onto the bastion host itself.
Session logging. Every keystroke and output of a session can be streamed to CloudWatch Logs or an S3 bucket, giving a full audit trail per session, tied to the IAM identity that started it, which is materially stronger than a traditional bastion's SSH access log, which typically only shows that a connection happened, not what was done inside it.
Migration plan. Roll out in phases rather than a single cutover: install and validate the Systems Manager agent on a subset of instances first, grant IAM session-start permissions to a pilot group of engineers, and run the new path in parallel with the existing bastion for a defined period; once the pilot group confirms full workflow coverage (including any tooling that assumed direct SSH, like certain deployment scripts), remove inbound SSH/RDP security group rules for the migrated instances, and only then decommission the legacy bastion host itself, since removing it too early, before every workflow has an equivalent, forces engineers back to workarounds that reopen inbound access informally.
flowchart LR
Eng[On-call engineer] -->|IAM auth plus MFA| SSM[Session broker: Systems Manager Session Manager]
SSM -->|outbound-only agent connection, no open inbound port| Instance[Private EC2 instance]
SSM -->|session transcript| Log[CloudWatch Logs or S3]
SSM -->|session start and stop events| Audit[CloudTrail audit trail]
Worked example
An on-call engineer needs emergency access to a production database host at 2 a.m. With the session-broker model, they authenticate to the AWS console or CLI with their existing IAM identity (already requiring MFA), which already carries a break-glass IAM policy scoped to session-start on production instances tagged role=oncall-emergency, and start a session directly; no key to locate, no bastion IP to remember, no separate VPN client. The full session transcript lands in CloudWatch Logs automatically, and the security team's next-morning review shows exactly which commands were run, by whom, and for how long, without needing to correlate a bastion's SSH log against a separate change ticket.
Trade-offs and pitfalls
The session broker becomes a critical dependency itself: if the managed service or its agent has an outage, emergency access through that path is unavailable, so a genuine break-glass fallback (a tightly controlled, alarmed, rarely-used traditional path) is still worth keeping for that narrow case, rather than assuming the managed service is infallible. A second pitfall in migration: cutting over the humans but forgetting the automation, deployment scripts, configuration management tools, or monitoring agents that were quietly relying on direct SSH access, which breaks in ways that look unrelated to the bastion project until someone traces the failure back to the removed inbound rule.
Design the network connectivity for a PCI DSS scoped workload in AWS that must connect to external payment processors. Explain segmentation to minimize scope, options between Direct Connect and VPN for payment traffic, PrivateLink usage, ensuring encryption in transit, centralized logging and monitoring, and practical steps to keep non-PCI systems out of scope.
Sample Answer
Direct Answer
Segment the cardholder data environment (the CDE, the systems that store, process, or transmit primary account numbers) into its own subnets or account so the smallest possible set of systems carries PCI DSS (Payment Card Industry Data Security Standard) scope, reach the external payment processor over a Direct Connect circuit with IPsec (Internet Protocol Security, a protocol suite that encrypts and authenticates traffic between two network endpoints) or MACsec on top (or a plain Site-to-Site VPN if the processor only exposes a public endpoint), use PrivateLink (private AWS network endpoints that reach one specific service without touching the public internet) for any AWS-internal hop inside that boundary, encrypt everything in transit at the network layer AND the application layer, and centralize logging so you can prove, not just assert, that non-CDE systems cannot reach the CDE.
Segmentation and Access Design
CDE-subnet isolation. Put only the minimum necessary compute, typically the payment gateway or tokenization service, in dedicated CDE subnets. Everything else (web tier, general app tier that never touches a raw card number) lives in explicitly separate subnets, ideally a separate VPC or a separate AWS account inside an AWS Organization, with one documented, minimal, logged connection between them (usually one direction only: the app tier calls a tokenization API and never sees the card number itself). This is the move that actually shrinks audit scope, because PCI DSS scope is every system that can affect the CDE's security, not just systems that store card data.
Route tables and security groups as the enforcement layer. The CDE subnet's route table should have routes only to: the local VPC, the specific PrivateLink endpoints it needs, and the Direct Connect/VPN path to the processor. No route from any non-CDE subnet should point at the CDE subnet. Security groups on CDE instances should be default-deny with explicit, named allow rules; network ACLs at the CDE subnet boundary are a cheap secondary layer but are not, by themselves, what a Qualified Security Assessor (QSA) will accept as segmentation evidence if the CDE and non-CDE workloads still share a flat VPC and a permissive security group.
Bastion access without a bastion. A classic SSH jump box is itself an inbound path into the CDE and a standing credential a PCI DSS Requirement 7/8 assessor will scrutinize. Replace it with AWS Systems Manager Session Manager: no inbound security group rule, no public IP on the target, every session logged to CloudWatch or S3 as compliance evidence, and access gated by an IAM policy that only allows a temporary, just-in-time (JIT) role assumption scoped to an approved change window, rather than a permanent credential sitting on a bastion host.
Direct Connect vs VPN for Payment Traffic
Direct Connect (DX) is a private, dedicated circuit to AWS, but it is not encrypted by default. To satisfy PCI DSS's encryption-in-transit requirement over a network path outside your direct control, add either MACsec (available on certain dedicated DX connections, IEEE-standard data confidentiality and integrity at the link layer) or an IPsec VPN running over the DX virtual interface. DX buys predictable latency and bandwidth, which matters if the processor requires low-jitter, high-volume real-time authorization traffic, and many processors offer colocation at the same facilities as AWS Direct Connect locations.
A plain Site-to-Site VPN over the internet is IPsec-encrypted by default (typically IKEv2, the protocol that negotiates and sets up the encrypted session, paired with AES-GCM, the algorithm that then does the actual encryption), quick to provision, and the right choice when the processor only exposes a public API endpoint with no private-connectivity option. Its throughput is capped per tunnel (1.25 Gbps standard, up to 5 Gbps with Large Bandwidth Tunnels), and its latency is less predictable than a dedicated circuit.
Recommendation: for a processor with steady, high-volume traffic and a colocation or partner presence, commit to Direct Connect with IPsec or MACsec. For a processor reachable only over a public HTTPS endpoint, or for lower, less latency-sensitive volume, a Site-to-Site VPN is proportionate and faster to stand up. What would flip the choice: if the workload later needs sub-50ms, high-throughput authorization at sustained volume, the added operational cost of DX becomes worth it; for an occasional batch settlement file, it usually is not.
PrivateLink Usage
For any hop that stays inside AWS, calling a fraud-scoring service in another account, or reaching Amazon S3 for a tokenized batch file, use interface VPC endpoints (AWS PrivateLink) so the traffic never touches the public internet or a NAT gateway. This shrinks the network attack surface that counts toward PCI scope. If the payment processor itself publishes a PrivateLink endpoint service (some processors partner with AWS to offer exactly this), prefer it over any public path, since it removes the internet hop and public IP addressing entirely.
Encryption, Logging, and Keeping Non-PCI Systems Out of Scope
Encryption in transit happens at two layers that are easy to conflate: the network layer (IPsec over VPN, MACsec over DX) and the application layer (TLS 1.2 minimum, TLS 1.3 preferred, per PCI DSS 4.0). PrivateLink is private connectivity, not an encryption guarantee by itself, so terminate TLS at the application layer even over a PrivateLink path.
Centralized logging and monitoring means VPC Flow Logs enabled on every elastic network interface (ENI) and subnet in the CDE, shipped to a centralized log-archive account, retained to meet PCI DSS Requirement 10 (a full year, with the most recent three months immediately queryable). This is literal compliance evidence for the QSA, not just an operational nicety. Pair it with a Gateway Load Balancer (GWLB)-fronted intrusion detection appliance at the CDE boundary and centrally aggregated CloudTrail and GuardDuty findings.
Keeping non-PCI systems out of scope in practice means: no shared security group or NACL between CDE and non-CDE subnets, no route from non-CDE subnets into the CDE, dedicated IAM roles scoped by account or tag, a network diagram that stays current as part of every assessment, and periodic connectivity tests (a scan from a non-CDE subnet to the CDE should fail) run alongside quarterly Approved Scanning Vendor (ASV) scans so scope drift is caught rather than assumed away.
Worked Example
flowchart LR
subgraph NonCDE[Non-CDE VPC/account]
Web[Web tier 10.0.1.0/24]
App[App tier 10.0.5.0/24]
end
subgraph CDE[CDE subnet 10.0.100.0/24, no IGW route]
Gateway[Tokenization service]
end
LogAcct[Centralized log account: Flow Logs, CloudTrail]
Processor[Payment processor]
App -->|one documented API call, TLS| Gateway
Gateway -->|DX + IPsec/MACsec or VPN| Processor
Gateway -.->|Flow Logs, one-way| LogAcct
Web -->|no route exists| CDE
A VPC of 10.0.0.0/16 carries CDE subnets at 10.0.100.0/24 with no route to an internet gateway. Non-CDE subnets, 10.0.1.0/24 through 10.0.9.0/24, have zero route table entries pointing at 10.0.100.0/24, so the "no route in" claim is checkable directly against the route tables, not just asserted in a diagram. The CDE's only egress routes are to the PrivateLink endpoints it uses and to the Direct Connect or VPN attachment reaching the processor.
Trade-offs and Pitfalls
Putting the CDE in the same flat VPC as everything else and relying only on security groups is the most common failure. QSAs generally will not accept security-group-only isolation as PCI-grade segmentation, because a single misconfigured rule reopens the whole VPC; separate subnets, separate route tables, and ideally a separate account are what actually holds up.
Assuming PrivateLink traffic does not need TLS because "it's private" is a real and recurring finding. PCI DSS 4.0 requires strong cryptography for cardholder data in transit over any network, trusted or not.
A separate AWS account for the CDE gives the cleanest scope boundary (the IAM and Organizations boundary does more work than any network control) but adds operational overhead: cross-account networking, duplicated shared services. A single account with strict subnet and route isolation is cheaper to run and harder to defend to a QSA. Choose based on how much that operational cost is worth against how clean the scope story needs to be.
MACsec is the strongest transit control for Direct Connect but only works on specific dedicated connection speeds with a compatible customer router; verify availability before designing around it as the sole encryption layer.
Explain how to set up packet capture in AWS for debugging intermittent network issues using VPC Traffic Mirroring. Include selecting mirror sources, creating mirror sessions and filters, choosing mirror targets (appliances or capture instances), expected performance impacts, and how to pipeline stored PCAPs to analysis tools without overloading storage.
Sample Answer
Direct Answer
Pick the specific elastic network interface (ENI) showing the intermittent problem as the mirror source, write a narrow filter so you capture only the traffic you actually need, point the session at a right-sized target, one capture instance or a fleet behind a load balancer, and keep the whole thing running only as long as the investigation does, so cost and stored data stay proportionate to the problem.
Setting Up Traffic Mirroring
Mirror sources. Any supported ENI can be a source. Choose the specific instance or instances exhibiting the issue rather than mirroring an entire fleet; this keeps both cost and the exposure of potentially sensitive captured data small.
Mirror sessions and filters. A session binds one source to one target through a filter. A filter is an ordered list of rules, protocol, source and destination CIDR (Classless Inter-Domain Routing, the notation for writing an IP address range, like 10.0.0.0/16), port range, accept or reject, evaluated top to bottom similarly to a network ACL, that lets you mirror, say, only TCP port 443 traffic instead of everything on the ENI. Filters also support packet truncation, capturing only the first N bytes of each packet, which is often enough for header-level troubleshooting and cuts both processing and storage cost.
Mirror targets. A single dedicated capture instance running standard packet-capture tooling works for low to moderate volume. For higher volume, a fleet of capture instances behind a Network Load Balancer, or a Gateway Load Balancer with a UDP listener, spreads the mirrored traffic instead of overwhelming one box.
Expected Performance Impact
Mirroring happens alongside the real traffic at the hypervisor layer, so the impact on the source instance's own network performance is minimal by design, that is the point of an off-box copy. The impact to actually manage is on the target side: a busy source ENI can mirror more data than a small target instance's network interface can absorb, so size the target, or the fleet, to the expected mirrored volume, and lean on a tight filter and packet truncation to cut that volume before it ever reaches the target.
Pipelining Captures Without Overloading Storage
Do not let a capture instance buffer indefinitely to local block storage. Rotate captures into small, time-boxed files, for example every one to five minutes, and ship each closed file to object storage immediately rather than accumulating it locally. Run analysis, packet-inspection tooling or a custom parser, either on the capture instance before deletion or as a downstream job triggered by the new file landing in storage. Apply a lifecycle rule that expires raw capture files after a short retention window measured in days, not months, since raw packet captures can contain sensitive payloads and unrestricted retention is both a cost problem and a compliance liability.
Worked Example
A team reports an intermittent two- to three-second stall between service A and service B every few hours. The setup: create a mirror filter that accepts only TCP traffic on the specific port the two services use, reject everything else, and point the session at one small capture instance in the same Availability Zone as the source (avoiding an unnecessary cross-Availability-Zone hop for the mirrored copy). Let the session run for a bounded window, a few hours, spanning at least one expected occurrence of the stall, then pull the resulting capture files from storage and inspect them for retransmissions or duplicate acknowledgments, exactly the signature an intermittent stall like this usually leaves behind.
Trade-offs and Pitfalls
Mirroring one hundred percent of a busy production ENI's traffic "to be safe" multiplies both the data-transfer cost of the mirrored copy and the risk of overwhelming the target. Start narrow and widen the filter only if the first pass misses the signal.
Leaving a mirror session running after the incident is resolved is a common, silent cost: it bills hourly per source ENI regardless of whether anyone is looking at the captured data.
A single capture instance is simpler and cheaper for low or moderate traffic; a load-balanced fleet becomes necessary once mirrored volume approaches one instance type's network capacity, at the cost of more moving parts to operate.
Not every EC2 instance family supports Traffic Mirroring as a source. Confirm support for the specific instance type before designing a debugging plan around it.
Compare and contrast instance-level security groups and subnet-level Network ACLs (NACLs) in cloud providers. Explain evaluation order, stateful vs stateless behavior, default rules and limits, and give examples of when to use each for a multi-tier application. Include an example where both are required and explain why.
Sample Answer
Direct answer
Security groups operate at the instance level (attached to the network interface) and are stateful: allow an inbound request and the matching response is automatically permitted back out, with no separate outbound rule needed. Network ACLs (NACLs) operate at the subnet level and are stateless: you must explicitly allow both directions of every conversation, including the ephemeral high-numbered ports used for return traffic, or replies get silently dropped. In a well-designed multi-tier application both layers are typically in play at once, security groups doing the fine-grained, per-tier access control and NACLs providing a coarser, second layer that holds even if a security group is ever misconfigured.
Structured elaboration
| Dimension | Security Group | Network ACL |
|---|---|---|
| Scope | Attached to individual network interfaces (effectively, per instance) | Attached to a subnet; applies to every resource in it |
| State | Stateful: return traffic for an allowed connection is automatically permitted | Stateless: inbound and outbound rules are evaluated independently; return traffic needs its own explicit rule |
| Rule evaluation | All rules are evaluated; there is no "first match wins" concept, only allow rules exist (nothing is explicitly denied, traffic simply isn't allowed if no rule matches) | Rules are evaluated in numbered order, lowest number first; the first rule that matches wins, and an explicit Deny is possible |
| Default behavior | A default security group denies all inbound and allows all outbound until you add rules | A default NACL allows all inbound and outbound traffic; a custom NACL you create denies everything until you add rules |
| Typical limits | Commonly around 60 rules per direction by default (adjustable), and a handful of security groups per network interface | Commonly around 20 rules per direction by default (adjustable up to around 40 per direction), evaluated in strict numeric order |
| Best used for | Fine-grained, per-tier or per-role access control ("only the web tier's security group may reach the app tier's security group on this port") | A coarse, subnet-wide boundary: blocking a known-bad IP range outright, or enforcing a hard compliance requirement that must hold regardless of what any individual security group says |
Why stateful vs stateless matters in practice, concretely. A client outside the VPC opens a connection to a web server: the request arrives from an ephemeral, randomly assigned high-numbered source port on the client side, destined for port 443 on the server, and the server's reply goes from port 443 back to that same ephemeral port on the client. A security group only needs one inbound rule (allow port 443 from the client's range) because it tracks the connection's state and automatically permits the matching reply out, no matter what port the reply uses. A NACL has no such memory: the inbound rule allowing port 443 says nothing about the outbound reply, so the NACL also needs an explicit outbound rule allowing traffic to the entire ephemeral port range (typically 1024 to 65535) back to the client, or every reply will be silently dropped at the subnet boundary even though the security group and the application are both configured correctly. This ephemeral-port gap is the single most common NACL misconfiguration: everything looks right, connections still fail, and the cause is invisible unless you specifically know to check for the missing return-traffic rule.
When to use each for a multi-tier application. Security groups do essentially all of the meaningful access control day to day: web tier's security group allows inbound 443 from the internet, application tier's security group allows inbound only from the web tier's security group, database tier's security group allows inbound only from the application tier's security group. NACLs are usually left at a permissive default in this pattern, and are reached for deliberately, not by default, when you need a control that survives a security-group mistake: for example, a subnet-wide rule blocking a specific IP range known to be malicious regardless of which instance or security group it targets, or a hard compliance boundary (say, "the database subnet must never accept inbound traffic from outside the VPC's own CIDR (its assigned IP address range, e.g. 10.0.0.0/16), full stop") that you want enforced even if someone someday attaches an overly permissive security group to something in that subnet by mistake.
Worked example
A case where both are genuinely required, and why. A financial services company runs its database tier in a private subnet and, separately from its normal tier-to-tier security groups, has a compliance requirement that the database subnet must categorically reject any traffic from outside the VPC's address range, independent of any application-level security group configuration, because a security group misconfiguration must never be sufficient on its own to expose the database externally. They implement this as a custom NACL on the database subnet with an explicit rule denying inbound traffic from any source outside the VPC's CIDR, positioned before (a lower rule number than) a broader allow rule for the VPC's own range, plus the normal outbound ephemeral-port allow rule for return traffic to the application tier. Independently, the database's security group allows inbound only from the application tier's specific security group on the database port, which is the layer actually doing meaningful per-tier access control day to day. The two layers answer different questions: the security group asks "is this specific source explicitly permitted," while the NACL asks "does this traffic even belong in this subnet at all," and having both means a mistake in one (an overly broad security-group rule accidentally added during a migration, for instance) does not by itself defeat the other.
Trade-offs and pitfalls
The most common mistake is trying to use NACLs for fine-grained, frequently changing access control, such as per-application-team rules: because NACL rules are evaluated in strict numeric order and typically capped at a modest number per direction, this quickly becomes an unmanageable, easy-to-misorder rule list, which is precisely the job security groups are built for instead. The second most common mistake, already covered above, is forgetting the ephemeral-port outbound rule when customizing a NACL, which silently breaks return traffic for entirely correctly configured applications and security groups. Finally, teams sometimes treat NACLs as unnecessary given that security groups exist, missing that the value of a NACL is specifically that it is a separate, subnet-wide layer that does not depend on any individual resource's security-group configuration being correct, which is exactly the property that makes it worth the extra maintenance for a genuinely hard boundary like the database example above.
A fintech customer requires microsegmentation to limit lateral movement between services. Describe how you would implement microsegmentation in a cloud environment using a combination of security groups, NACLs, host-based firewalls, service mesh policies, and cloud firewall appliances. Explain enforcement points, policy lifecycle (authoring, testing, rollout), performance impact, and a migration strategy from a flat network.
Sample Answer
Direct Answer
Layer four distinct enforcement points, security groups, network ACLs, host-based firewalls, and service mesh policy, plus a centralized cloud firewall appliance for anything crossing a trust boundary, so a compromised service is stopped by more than one control. Treat policy authoring, testing, and rollout as a reviewed, staged pipeline with automated drift detection rather than a one-time firewall change, because in a fintech environment the rules that get hand-added during an incident are exactly the ones that quietly outlive the incident.
Enforcement Points and What Each One Actually Stops
- Security groups (SGs): stateful, attached per elastic network interface (ENI). The primary AWS-native microsegmentation primitive between service tiers; reference other tiers by security-group ID, not IP range, so autoscaling doesn't break the rule.
- Network ACLs (NACLs): stateless, attached per subnet. Good for a coarse, subnet-wide blocklist; not fine-grained enough alone for service-to-service policy.
- Host-based firewalls (iptables, nftables, or a fleet-managed agent): enforce policy even if the workload itself is compromised and tries to open an unexpected listener, and are the only layer here that can see the actual process or user making the connection, not just the packet.
- Service mesh policies (mutual TLS plus layer-7 authorization, Istio- or Linkerd-style): enforce identity-based access rather than IP-based access, which matters once workloads are autoscaled or containerized and IP addresses churn faster than SG rules can safely be automated.
- Cloud firewall appliances (for example a Gateway Load Balancer fronting a next-generation firewall): the centralized chokepoint for anything crossing a VPC, account, or internet boundary that the layers above do not cover.
The layers compose rather than duplicate: SG and NACL narrow the network path, the host firewall narrows what a compromised host can originate even off that path, mesh mTLS narrows who is allowed to call what regardless of network path, and the appliance is the checkpoint at the actual trust boundary.
Policy Lifecycle: Authoring, Testing, Rollout, and Drift
Authoring. Define policy as code (Terraform for SG and NACL rules, an authorization-policy definition for the mesh), reviewed by pull request, mapped against a real service-dependency graph rather than a guess at who calls whom.
Testing. Apply a new rule in an audit or log-only mode first (most service meshes and next-generation firewalls support this) against real production-shaped traffic before it can deny anything, and canary a new deny rule against a single instance or a single availability zone before a fleet-wide rollout.
Rollout. Stage the rollout (canary, then a percentage, then everywhere) with an automated rollback if an error-rate or health-check service-level objective (SLO) regresses.
Rule drift and cross-account audit. Rules drift because an emergency exception, "just open this port for the incident", gets added and never removed. Three controls close this: policy-as-code as the only allowed change path, enforced by denying direct console-level security-group edits outside the CI role; a scheduled drift-detection job that diffs the live SG and NACL state against the git-declared state and pages on divergence; and a recurring cross-account audit (an AWS Config aggregator, or a scheduled reachability check across every account in the Organization) so an exception opened in one account does not silently persist for months unnoticed by anyone outside that account. Every manual exception should carry an automatic expiry.
Performance Impact
Security groups and NACLs are enforced in the hypervisor's fast path, so their overhead is effectively negligible. Host-based firewalls add a small, fixed per-packet cost that is usually dwarfed by TLS termination cost. A service mesh sidecar proxy is the layer that adds the most measurable overhead: an extra userspace hop per call, whose dominant cost is proxy processing and serialization rather than the mutual-TLS handshake itself, since sessions are reused across calls. Mitigate with sidecar resource sizing, connection pooling, and meshing only the boundaries that genuinely need identity-based policy rather than every internal call. Avoid asserting a specific millisecond number here; it depends on payload size, proxy implementation, and instance type, and should be validated with a load test against the real environment before it goes into an SLA (service-level agreement, a contractual performance commitment made to customers).
Migration Strategy from a Flat Network
- Map the real dependency graph first, from VPC Flow Log analysis or existing mesh telemetry. Do not guess the call graph.
- Turn on monitor-only (log, don't enforce) mode across every layer at once, to build confidence in the mapped graph without breaking production.
- Segment by blast-radius priority, not alphabetically: isolate the service handling the most sensitive data, in a fintech context typically anything touching payment or customer PII (personally identifiable information), first.
- Introduce one enforcement layer at a time in production, starting with SG default-deny (cheapest and most reversible), verify, then add mesh mTLS, then host firewalls, then the appliance.
- Keep a documented, time-boxed break-glass path, an emergency access rule that requires an approval workflow to enable, so a migration incident does not turn into "we opened the network to everything for a week and forgot."
Worked Example
A payments call graph: web calls orders-api, which calls ledger-service, which calls both the database and a separate fraud-service. Under the flat-network starting point, all four sit in one subnet with one broad security group. The migration puts each service behind its own security group referencing the caller's SG by ID (orders-api's SG allows inbound only from web's SG on its listening port; ledger-service's SG allows inbound only from orders-api's SG), adds a mesh identity policy so fraud-service only accepts calls whose mesh-issued certificate identifies them as ledger-service, and puts a GWLB-fronted firewall at the VPC boundary so nothing in this graph can be called from outside the VPC at all except through that one inspected path. Each of these was rolled out in log-only mode first against the real, already-flowing traffic, then flipped to enforce, tier by tier, over the migration.
Trade-offs and Pitfalls
Relying on security groups alone in an autoscaled or containerized environment, where pod IP addresses are ephemeral, forces a choice between rules too broad (allowing an entire subnet CIDR) or unmaintainable (constantly editing IP-based rules). Identity-based mesh policy is what actually solves this; IP-based rules alone do not.
Log-only mode that never gets flipped to enforce is a common failure. Without an owner and a deadline for the flip, monitor-only policy is a dashboard, not a control.
A service mesh adds real operational weight: a control plane, certificate rotation, sidecar upgrades, versus the near-zero maintenance of security groups and NACLs. Take it on when ephemeral-IP, identity-based segmentation is a genuine requirement, not because it is fashionable.
A centralized firewall appliance concentrates policy and inspection but also becomes a scaling and failure-domain dependency; reserve it for the actual trust-boundary chokepoints rather than east-west traffic the mesh already covers.
Unlock Full Question Bank
Get access to all 14 Cloud Networking and VPC Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.