Amazon Senior Network Engineer Interview Preparation Guide
Amazon's interview process for Senior Network Engineer typically consists of 1 recruiter screening round, 2 technical phone screens, and 4 onsite technical and behavioral interview rounds. The process emphasizes deep technical expertise in network architecture and design, troubleshooting complex infrastructure problems, scaling systems to support millions of users, and demonstrating Amazon's Leadership Principles. Candidates should expect scenario-based questions, system design discussions, and behavioral assessments focused on decision-making with incomplete information and handling ambiguity.
Interview Rounds
Recruiter Screening
What to Expect
Initial conversation with Amazon recruiter covering background, career goals, role expectations, and preliminary fit assessment. May include a follow-up technical conversation with a recruiter to verify baseline networking knowledge. This round confirms your interest in the role, discusses relocation if necessary, compensation expectations, and notice period. The recruiter will also explain the interview process and timeline.
Tips & Advice
Be enthusiastic about Amazon's mission and the specific team. Have a clear narrative about your career progression and why senior level is appropriate for your experience. Ask insightful questions about the team, infrastructure challenges, and career growth opportunities. Be honest about your salary expectations and availability. Mention any AWS certifications or experience with large-scale infrastructure.
Focus Topics
Understanding of Amazon's Leadership Principles
Familiarity with Amazon's 16 Leadership Principles and ability to relate your experience to principles like Ownership, Customer Obsession, and Think Big
Practice Interview
Study Questions
Motivation for Role & Company
Clear articulation of why you're interested in this specific role, team, and Amazon as an organization
Practice Interview
Study Questions
Career Narrative & Progression
Articulate your career journey from previous roles to senior network engineering, highlighting progression in responsibilities, complexity of systems managed, and technical depth gained
Practice Interview
Study Questions
Technical Phone Screen 1: Networking Fundamentals & Troubleshooting
What to Expect
Technical interview with a network engineer from Amazon conducted over video/phone. Focuses on your ability to diagnose and troubleshoot complex network connectivity issues using a systematic, hypothesis-driven approach. Expect scenario-based questions where you must isolate problems (DNS vs routing vs firewall), demonstrate knowledge of networking tools (netstat, ss, dig, curl, tcpdump), and explain your troubleshooting methodology. Questions may include port connectivity issues, VLAN routing problems, or application performance degradation. The interviewer will evaluate your problem-solving process, tool knowledge, and ability to separate different network layers.
Tips & Advice
Always start by confirming the problem scope and gathering information before proposing solutions. Use a layered troubleshooting approach (OSI model): verify physical connectivity, IP layer (routing/interfaces), DNS resolution, port accessibility, and application-layer issues in systematic order. Explain your commands and what they validate. For example, when addressing a connectivity issue, demonstrate knowledge that you'd check IP routing with 'ip route', verify the service is listening with 'ss -lntp', check firewall rules, then test with 'nc' or 'curl'. Avoid making multiple changes simultaneously; isolate variables. If you don't know a tool or command, say so and explain how you'd research it. For senior level, interviewers expect you to not just solve the problem but explain architectural decisions and trade-offs.
Focus Topics
Application-Layer Network Issues vs Infrastructure Issues
Ability to isolate whether slow performance is due to network MTU issues, proxy settings, TCP retransmissions, QoS/traffic shaping, or application-level rate limiting
Practice Interview
Study Questions
DNS Troubleshooting Scenarios
Ability to differentiate DNS resolution failures from routing problems; understanding resolver configuration files (/etc/nsswitch.conf), DNS server reachability, and application-specific DNS issues
Practice Interview
Study Questions
Port Connectivity & Service Configuration
Verification that services listen on correct ports, understand binding to 0.0.0.0 vs localhost, firewall rule impact, and NAT/port-forwarding for external access
Practice Interview
Study Questions
VLAN Configuration & Inter-VLAN Routing
Understanding VLAN segmentation, trunk port configuration, access modes, inter-VLAN routing requirements, ACLs between VLANs, and subnet mask/gateway configuration
Practice Interview
Study Questions
Linux Network Diagnostic Tools
Proficiency with tools including ss/netstat (port listening), ip (address, route, neighbor), dig/nslookup (DNS), tcpdump/wireshark (packet analysis), traceroute, nc/curl (connectivity testing), iptables/nftables (firewall rules)
Practice Interview
Study Questions
Network Troubleshooting Methodology
Systematic approach to isolating connectivity issues using the OSI model: physical layer, IP layer (routing, ARP, interfaces), DNS resolution, transport layer (ports, firewalls), and application layer
Practice Interview
Study Questions
Technical Phone Screen 2: Network Design & Architecture
What to Expect
Technical interview focused on network architecture design and strategic decision-making. You'll be presented with infrastructure requirements (e.g., 'Design a network for a service supporting 100K concurrent users across 3 AWS regions') and must propose solutions considering redundancy, security, performance, and scalability. Expect questions about routing protocols (BGP, OSPF), load balancing strategies, firewall architecture, VPN design, redundancy models, and trade-offs between solutions. This round assesses whether you can think architecturally, understand the business requirements behind technical decisions, and communicate design rationale clearly. You'll need to ask clarifying questions, state assumptions, and explain how your design handles failures and scales.
Tips & Advice
Start by asking clarifying questions about requirements (scale, geographic distribution, failure tolerance, latency requirements, security constraints, budget). State your assumptions explicitly. Propose a design broken into layers (edge/CDN, load balancing, core network, data center internal network, security). For a senior role, explain trade-offs: why BGP over OSPF, why that redundancy model, cost implications of design choices. Use AWS concepts if relevant (Route 53 health checks, VPC design, AWS Global Accelerator). Explain how your design handles cascading failures and graceful degradation. Draw diagrams or describe topology clearly. For a service with millions of users, discuss geographic distribution, multi-region failover, and edge location strategy. Mention monitoring and visibility into the network. Address security at each layer (firewalls, ACLs, DDoS mitigation). Be prepared to adapt your design based on interviewer feedback.
Focus Topics
Trade-Off Analysis & Justification
Ability to articulate design trade-offs (complexity vs resilience, cost vs performance, security vs usability) and justify architectural choices based on requirements
Practice Interview
Study Questions
Load Balancing Architecture & Strategies
Load balancing layers (L4 vs L7), algorithms (round-robin, least connections, content-based), health checking, session persistence, and multi-region load balancing strategies
Practice Interview
Study Questions
Multi-Region & High Availability Design
Designing fault-tolerant systems across availability zones and regions, implementing cross-region failover, data replication strategies, and graceful degradation during outages
Practice Interview
Study Questions
Network Security Architecture & Segmentation
Firewall architecture, DMZ design, VPC/subnet segmentation strategies, ACL rules, DDoS mitigation, intrusion detection/prevention, and security policy implementation
Practice Interview
Study Questions
Routing Protocols & Strategic Selection
Understanding BGP (exterior), OSPF/IS-IS (interior), route convergence, path selection, failover behavior, and when to use each protocol based on scale and requirements
Practice Interview
Study Questions
Large-Scale Network Architecture Design
Design principles for networks supporting millions of concurrent users: multi-region deployment, geographic distribution, redundancy models, failover strategies, and traffic management across availability zones
Practice Interview
Study Questions
Onsite Round 1: Technical Deep Dive - Network Architecture for Enterprise Scale
What to Expect
Full-day onsite with first technical interview focused on designing complex network infrastructure for enterprise or service provider scale. You'll receive a detailed scenario describing a business requirement (e.g., 'Design a global network supporting multiple data centers, high-traffic services, and international expansion') and must design the entire infrastructure. Expect 45-60 minutes of discussion where you propose topology, address redundancy, security, operational management, and monitoring. The interviewer will challenge your assumptions, ask follow-up questions about failure scenarios, and push you to think through operational details. This round assesses architectural thinking, depth of networking knowledge, and ability to balance multiple conflicting requirements.
Tips & Advice
Take 5-10 minutes to understand requirements fully before proposing design. Draw a clear topology diagram. Structure your design by layers and explain each layer's purpose. For enterprise scale, discuss: core network redundancy (dual core, CLOS fabric, or spine-leaf), data center interconnect (dark fiber, MPLS, or cloud-based WAN), edge distribution, and security zones. Explain operational aspects: how you monitor the network, manage configuration changes, perform maintenance without downtime. Address cost implications of design choices. Discuss SDN/automation potential for large-scale networks. When challenged, be willing to refine your design—show that you can adapt based on new information. Use standards and industry best practices (RFC specifications, architectural patterns). For a senior role, demonstrate knowledge of modern network architectures (segment routing, network slicing, cloud-native networking).
Focus Topics
Security Architecture for Distributed Networks
Implementing security zones, firewalling strategies for multi-layer networks, DDoS mitigation at different layers, encryption in transit between data centers, and compliance requirements
Practice Interview
Study Questions
Network Monitoring & Observability Architecture
Designing telemetry collection (sFlow, NetFlow, SNMP), metrics aggregation, alerting strategies, and dashboarding for large networks; understanding traffic engineering and capacity planning
Practice Interview
Study Questions
Data Center Interconnect (DCI) Design
Strategies for connecting multiple data centers: dark fiber, MPLS/SD-WAN, cloud-based WAN, throughput requirements, latency optimization, and active-active vs active-passive modes
Practice Interview
Study Questions
Network Redundancy & Failover Design
Designing redundancy at core, distribution, and access layers; fast convergence techniques, hitless failover mechanisms, and protecting against cascading failures
Practice Interview
Study Questions
Spine-Leaf & CLOS Network Architecture
Modern data center fabric architecture providing non-blocking, predictable bandwidth; understanding leaf switches, spine switches, and traffic patterns in CLOS topologies
Practice Interview
Study Questions
Onsite Round 2: Technical Deep Dive - Operations, Troubleshooting & Performance Optimization
What to Expect
Technical interview focused on operational excellence and performance optimization in large-scale networks. You'll discuss how you monitor, troubleshoot, and optimize network performance in production. Expect scenarios like 'Latency suddenly increased in a critical service' or 'How would you diagnose and resolve packet loss affecting 5% of traffic?' Interviewers will evaluate your depth of operational knowledge, including traffic engineering, capacity planning, performance metrics, and how you'd mentor junior engineers through complex troubleshooting. This round assesses whether you can translate design into reliable operations and continuously improve network performance. Discussions may cover buffering, congestion control, QoS implementation, and handling extreme traffic scenarios.
Tips & Advice
Use a hypothesis-driven approach even for operational scenarios. For performance issues, explain your monitoring strategy first—what metrics would you track? How would you detect issues? When troubleshooting, walk through your systematic approach to isolating the root cause. Discuss packet capture analysis, flow analysis with NetFlow/sFlow, and correlation of metrics from different layers. For optimization, explain trade-offs: improving performance may increase cost or complexity. Demonstrate knowledge of TCP behavior, packet loss handling, buffering strategies, and QoS policies. For a senior role, discuss how you'd design monitoring and alerting to catch issues before they impact customers. Address how you'd handle extreme scenarios (traffic spike during Prime Day, DDoS attack, partial infrastructure failure). Explain how you'd document issues and create runbooks for junior team members.
Focus Topics
Quality of Service (QoS) Implementation
QoS policies for prioritizing traffic, rate limiting, traffic shaping, protecting critical services during congestion, and handling graceful degradation under extreme load
Practice Interview
Study Questions
TCP Behavior & Congestion Control
Understanding TCP congestion algorithms (Cubic, BBR), window scaling, retransmission behavior, timeout handling, and how congestion control impacts application performance
Practice Interview
Study Questions
Traffic Engineering & Capacity Planning
Predicting future capacity needs, planning upgrades, understanding traffic patterns across time/geography, load balancing for optimal utilization, and handling traffic spikes
Practice Interview
Study Questions
Handling Extreme Scale & Traffic Spikes
Network design decisions for handling millions of concurrent connections, preventing cascading failures, implementing circuit breakers, graceful degradation, and testing infrastructure limits
Practice Interview
Study Questions
Production Network Monitoring & Metrics
Key performance indicators (bandwidth utilization, latency, jitter, packet loss), telemetry collection methods (NetFlow, sFlow, SNMP), alert thresholds, and anomaly detection strategies
Practice Interview
Study Questions
Complex Troubleshooting in Production
Systematic approach to performance issues and packet loss; packet capture analysis, flow-level debugging, correlation of metrics across infrastructure layers, and identifying root causes of latency
Practice Interview
Study Questions
Onsite Round 3: System Design - Large-Scale Distributed Network Infrastructure
What to Expect
Extended system design interview (60-75 minutes) focused on designing complete network infrastructure for a complex scenario with multiple conflicting requirements. Example: 'Design a network for a social media platform serving 1 billion users across 6 continents with requirements for sub-100ms latency, high availability, and support for emerging features.' You must propose end-to-end architecture including edge CDN strategy, inter-data center connectivity, cloud integration (AWS VPC design if relevant), security architecture, operations model, and capability for future growth. The interviewer will introduce constraints (budget limitations, regulatory requirements, geographic restrictions) and push you to adapt your design. This round assesses strategic thinking, understanding of trade-offs at massive scale, and ability to balance engineering and business requirements.
Tips & Advice
Ask clarifying questions about scale (users, data volume, geographic distribution, latency targets, availability targets, security requirements). Propose a layered architecture: edge/CDN layer for proximity, access layer for aggregation, core for backbone, security throughout. Discuss technology choices and justify them (why this routing protocol, why this failover mechanism). For AWS context, discuss VPC architecture, Route 53 for health checks and failover, multi-region deployment, and private connectivity options. Address non-functional requirements explicitly: what's your latency strategy (anycast, GeoDNS, CDN caching)? How do you ensure sub-100ms global latency? How do you handle regional failures without cascading to other regions? Explain cost implications of your design. Discuss operational model: how do you deploy changes across regions? How do you handle maintenance windows? What's your disaster recovery strategy? For a senior role, discuss modern technologies (segment routing for traffic engineering, intent-based networking, cloud-native architectures) and show awareness of emerging challenges (IPv6 migration, carrier-grade NAT). Be prepared to iterate on your design based on interviewer feedback.
Focus Topics
Cost Optimization & Resource Efficiency
Balancing performance, reliability, and cost; understanding bandwidth costs, equipment costs, and designing networks that meet requirements without over-provisioning
Practice Interview
Study Questions
Evolution & Future-Proofing of Design
Designing networks with capacity for 10x growth, planning for emerging technologies, maintaining flexibility to adapt to changing requirements without complete redesigns
Practice Interview
Study Questions
Cloud Integration Architecture (AWS VPC & Services)
Designing hybrid/cloud networks: VPC architecture, subnet strategy for multi-AZ deployment, private connectivity (VPN, Direct Connect equivalents), and integrating cloud services with on-premises infrastructure
Practice Interview
Study Questions
Network Automation & Infrastructure as Code
Designing networks for automation-first operations; templating network configurations, automated deployment/rollback, GitOps principles, and reducing manual operational toil
Practice Interview
Study Questions
Global Content Distribution & Edge Caching Strategy
Design of CDN-like architecture with regional caches, geographic routing (anycast, GeoDNS), cache coherence, and minimizing latency for global users
Practice Interview
Study Questions
Multi-Region Architecture & Fault Isolation
Designing systems where failure in one region doesn't cascade to others; independent data replication, async communication patterns, idempotency in distributed systems
Practice Interview
Study Questions
Onsite Round 4: Amazon Leadership Principles & Behavioral Interview
What to Expect
Behavioral interview assessing cultural fit and alignment with Amazon's Leadership Principles. You'll be asked about past experiences demonstrating principles like Ownership, Customer Obsession, Invent and Simplify, Are Right, A Lot, Earn Trust, Think Big, and Bias for Action. Expect 4-6 scenario-based questions asking how you've handled difficult situations: 'Tell me about a time you had to make a difficult technical decision with incomplete information,' 'Describe a situation where you disagreed with a colleague and how you resolved it,' 'Give an example of when you failed and what you learned,' 'How do you balance customer needs with technical constraints?' This round evaluates leadership potential, decision-making ability, conflict resolution, and whether you thrive in Amazon's culture. For a senior role, interviewers assess whether you can mentor others, drive culture, and lead without formal authority.
Tips & Advice
Prepare 8-10 detailed stories using the STAR format (Situation, Task, Action, Result) that demonstrate Amazon Leadership Principles. For a senior network engineer, prioritize stories showing: 1) Ownership of large projects (taking responsibility beyond your immediate scope), 2) Customer Obsession (understanding how network decisions impact end users), 3) Bias for Action (making decisions with 70% information rather than waiting for perfect data), 4) Invent and Simplify (proposing novel solutions to complex problems, removing unnecessary complexity), 5) Are Right, A Lot (sound judgment in technical decisions), 6) Earn Trust (gaining credibility through competence and reliability). Emphasize mentoring and developing junior team members for a senior role. When answering, focus on your personal contribution, not just team achievements. Explain your reasoning and thought process. Address failures honestly and discuss what you learned. Show awareness of customer impact and business objectives, not just technical elegance.
Focus Topics
Amazon Leadership Principle: Invent and Simplify
Proposing novel solutions to complex problems, removing unnecessary complexity from designs, and driving simplification that improves maintainability and reduces operational burden
Practice Interview
Study Questions
Difficult Decision-Making with Incomplete Information
Examples of technical decisions made under uncertainty; how you gathered information, weighed trade-offs, made the decision, and handled outcomes
Practice Interview
Study Questions
Mentoring & Developing Technical Team Members
Examples of mentoring junior engineers, helping them grow technically, creating psychological safety for learning, and developing the next generation of network leaders
Practice Interview
Study Questions
Amazon Leadership Principle: Bias for Action
Making decisions with incomplete information (70% data), taking calculated risks, learning from outcomes, and preferring speed to excessive analysis
Practice Interview
Study Questions
Amazon Leadership Principle: Customer Obsession
Understanding how network decisions impact customers; optimizing for customer experience over internal convenience; gathering customer feedback and acting on it
Practice Interview
Study Questions
Amazon Leadership Principle: Ownership
Taking end-to-end responsibility for projects or systems; thinking long-term, going beyond immediate scope, and being accountable for outcomes even when things go wrong
Practice Interview
Study Questions
Frequently Asked Network Engineer Interview Questions
Tell me about a time you diagnosed and resolved a network outage caused by a VLAN or switching misconfiguration. Using the STAR method, describe the Situation, the Task you were responsible for, the Actions you took (tools/commands used), the Result, and the preventive measures you implemented after the incident.
Sample Answer
Direct answer
A strong STAR (Situation, Task, Action, Result) answer is one specific incident told in order, with the commands you ran and what each told you, then what you changed so it cannot repeat. Below is a complete worked example built on a classic switching mistake; replace the details with your own incident and quote your own measured figures from the ticket (duration, users affected), not estimates.
Situation
During a planned change window I was adding a new VLAN (a virtual LAN, a separate broadcast domain) for a finance team to an existing trunk between an access switch and the distribution switch. A trunk is one link that carries many VLANs at once, and its allowed VLAN list is the set of VLANs it is permitted to carry. Minutes after the change, users on several other VLANs on that floor lost connectivity, and monitoring showed the access switch's management address no longer answering from the network side, so I worked from the console.
Task
I was the engineer who made the change, so I owned the restoration and the root cause. The aim was to go straight to the change I had just made, restore service, check for collateral damage, and then stop it recurring.
Actions
- Went to the change I had just made. The change record said the only item touched was the trunk, so I looked at it before anything else:
show interfaces trunkon the access switch. The allowed VLAN list contained only the new VLAN, VLAN 30: the output has a section per trunk port listing the VLANs allowed, and where the design called for the finance VLAN plus every existing user and management VLAN, it showed just 30. That also explains the symptoms: the management VLAN was no longer carried, so the switch could not be reached over the trunk, and every other user VLAN had no path out of the floor. I had typedswitchport trunk allowed vlan 30with a bare VLAN list, which sets the list to exactly what is typed, instead ofswitchport trunk allowed vlan add 30, which appends. Every existing VLAN had been removed from the trunk. - Restored service. I ran
switchport trunk allowed vlan addwith the missing VLANs, taking the list from the VLAN design record and the last saved configuration, because I had not captured the trunk's list myself before the change, then verified withshow interfaces trunkthat every VLAN in the design was in the allowed list and in the forwarding state (forwarding means spanning tree lets the port pass traffic, as opposed to blocking, where it only listens). - Checked for collateral damage. Spanning tree is the protocol that blocks redundant links so Layer 2 frames cannot circle forever.
show spanning-tree detailshowed a burst of topology changes at the time of the change (a topology change is spanning tree reconverging after a port changed state; it makes switches age out or delete learned MAC entries so stale ones are dropped and relearned), andshow mac address-tableconfirmed that hosts were relearned on the right ports. I also confirmed the native VLAN (the one VLAN whose frames cross a trunk untagged) matched on both ends, because Cisco warns that a mismatch can cause spanning-tree loops: the two ends disagree about which VLAN the untagged frames, including spanning tree's own messages, belong to. - Told people. I updated the incident channel as soon as I knew the cause and the fix, instead of waiting for the post-mortem.
Result
Connectivity came back as soon as the trunk list was restored, and every VLAN was verified end to end. State the real duration and affected user count from your incident record. The honest summary is that the outage came from my command, and I owned that in the post-incident review.
Preventive measures
- Changes to trunk allowed lists now use the
addandremoveforms by policy, and the change template includes a pre-changeshow interfaces trunkcapture and a post-change diff. - Pre-change configuration backups are automatic, so the restore is a copy, not a recollection.
- A post-change checklist verifies at least one host per VLAN (ping the gateway) before the window is closed.
- Changes of this type are peer reviewed and, where the tooling allows, pushed from a template that renders the full intended list.
What I would do differently
I would have captured the allowed list before the change, tested the command on a lab switch, and kept a rollback timer ready (a safety net that reverts the change automatically, for example a scheduled reload to the saved configuration that you cancel once the change is proven), so that restoration would have been a single step rather than a diagnosis.
Explain Bandwidth-Delay Product (BDP) and its implications for TCP throughput on long-haul hybrid WAN links. Show how to compute BDP (bandwidth * RTT) and how you'd choose TCP window sizes or buffering to saturate a 1 Gbps transcontinental link with a 150 ms RTT. What other techniques or appliances can help improve effective throughput?
Sample Answer
Direct answer
Bandwidth-Delay Product (BDP) is the amount of data that can be "in flight" on a link at any instant, unacknowledged, before the sender has to stop and wait: it equals the link's bandwidth multiplied by its round-trip time (RTT). It matters because TCP will not send more unacknowledged data than the receiver's advertised window allows, so if that window is smaller than the BDP, the sender stalls waiting for acknowledgments (ACKs) even though the link itself has spare capacity. On a 1 Gbps link with a 150 millisecond (ms) RTT, the BDP works out to about 18.75 megabytes (MB), far larger than the roughly 64 kilobyte (KB) window a default TCP stack negotiates without extensions, which is why long-haul hybrid wide-area network (WAN) links routinely max out at a small fraction of their rated bandwidth unless the stack is explicitly tuned.
Structured elaboration
Computing BDP:
BDPbits=B×RTT BDPbits=1×109 bit/s×0.150 s=1.5×108 bit BDPbytes=81.5×108=1.875×107 bytes≈18.75 MB (17.9 MiB)Why the default window can't fill that pipe. The TCP header's window field is 16 bits, so without extensions the maximum advertised window is 65,535 bytes, about 0.35% of the 18.75 MB this link needs to stay full. Once the sender has 65,535 bytes outstanding it must stop and wait for an ACK, and on a 150 ms RTT link that means the connection can only push one window's worth of data roughly every 150 ms, which caps effective throughput at:
0.150 s65,535 bytes×8≈3.5 MbpsThat is well under 1% of the link's 1 Gbps rating, which is the classic long-fat-network (LFN) problem: high bandwidth, high delay, small window.
Fixing it with TCP window scaling (RFC 1323). Window scaling multiplies the advertised 16-bit window value by a power of two, up to a scale factor of 14 (a maximum receive window just over 1 GiB), so both endpoints must support and negotiate it (a SYN option exchanged during the handshake). To saturate this link, the effective window needs to be at least the BDP:
65,535×2n≥18,750,000 2n≥286.1⇒n=9(29=512) 65,535×512=33,553,920 bytes≈32 MiBA scale factor of 9 gives roughly 32 MiB of addressable window, about 1.8 times the raw BDP, which is deliberate headroom: packet loss and retransmission mean the connection needs room to keep sending while waiting for a retransmit timer, and modern congestion-control algorithms probe above the pipe's capacity before backing off. On Linux this is exposed as net.ipv4.tcp_window_scaling (on by default in any current kernel) plus buffer sizing via net.core.rmem_max / net.core.wmem_max and net.ipv4.tcp_rmem / net.ipv4.tcp_wmem, which should be set to at least the 32 to 64 MiB range for this link so the kernel's auto-tuning has room to grow the socket buffer up to the scaled window.
Other techniques and appliances for long-haul throughput:
- Congestion-control algorithm choice. Classic loss-based algorithms (Reno, and to a lesser extent CUBIC) interpret any packet loss as congestion and cut the window aggressively, which is punishing on a long-haul link where a single lost packet takes a full RTT to detect and recover from. A delay/bandwidth-aware algorithm such as BBR probes the path's actual bandwidth and RTT directly and holds throughput closer to the link's real capacity even with occasional non-congestive loss, which is common on satellite or long transcontinental paths.
- Parallel TCP streams. Splitting one transfer across multiple TCP connections (what tools like
bbcp, GridFTP, or a multi-threadedrsync/aws s3 cp --recursiveeffectively do) sidesteps a single connection's window ceiling by multiplying the number of independent windows in flight, at the cost of more connection and fairness overhead. - WAN optimization / acceleration appliances (for example Riverbed SteelHead or Silver Peak, now part of Aruba/HPE EdgeConnect). These typically do TCP proxying or "spoofing" at each end of the link, terminating the WAN-facing TCP connection locally so the application never has to negotiate a window across the full RTT itself, combined with deduplication and compression to cut the actual bytes that cross the link.
- Jumbo frames, where every hop on the path supports them, reduce per-packet header overhead and interrupt/processing load, which helps overall throughput but does not by itself fix the window-size problem.
Worked example
Take the 1 Gbps, 150 ms RTT link from the question. With window scaling disabled (or a stack that negotiates only the default window), the connection is capped near 3.5 Mbps as computed above, meaning a 10 gigabyte (GB) nightly transfer would take on the order of 6.5 hours purely from the window ceiling, independent of how fast either endpoint's disk or network interface actually is. Enabling window scaling with a scale factor of 9 (about 32 MiB effective window, comfortably above the 18.75 MB BDP) removes that ceiling; the achievable throughput then becomes a function of the link's actual available bandwidth and loss rate rather than the window, so the same transfer approaches the link's real 1 Gbps capacity, and a well-tuned single stream on a low-loss path can get close to line rate. On a lossier or more congested path, adding BBR or splitting the transfer into 4 to 8 parallel streams recovers most of the remaining gap, because a single window-scaled TCP connection is still just one flow reacting to loss as a single unit.
Trade-offs & pitfalls
- Window scaling only helps if both endpoints, and every middlebox on the path, pass the RFC 1323 SYN options through unmodified; a firewall or older load balancer that strips TCP options silently reintroduces the 64 KB ceiling with no obvious error.
- Oversizing buffers well beyond the BDP causes bufferbloat: packets queue for a long time in an oversized buffer under congestion, inflating latency for every flow sharing that link, including latency-sensitive traffic that has nothing to do with the bulk transfer. Size buffers to the BDP plus a deliberate, bounded margin, not "as large as possible."
- Parallel streams and WAN accelerators both trade fairness and complexity for throughput: parallel streams can starve other flows sharing the same bottleneck, and a proxying WAN accelerator terminates TCP semantics end to end, which complicates anything that depends on true end-to-end acknowledgment (some replication or transactional protocols assume that).
- BBR is not a universal upgrade; on paths with very high genuine congestion it can behave less fairly toward classic loss-based flows sharing the same bottleneck, so validate the choice against the actual traffic mix rather than assuming it as a default.
Describe the tcpdump command to capture only DNS (UDP port 53) traffic to/from host 10.1.1.5 on interface eth0, rotate captures to avoid filling disk (ring buffer), and explain how to interpret DNS transaction IDs and flags in the pcap.
Sample Answer
Direct answer
tcpdump -i eth0 -n 'udp port 53 and host 10.1.1.5' -w /var/log/dns_capture.pcap -C 100 -W 5
-n skips reverse DNS lookups on captured addresses (avoids the tool itself generating more DNS traffic while you are capturing DNS traffic, and speeds up display). The filter expression restricts capture to DNS traffic (UDP port 53) to or from the specific host. -w writes to a capture file instead of printing decoded packets to the terminal. -C 100 rotates to a new file every time the current one reaches 100 million bytes, and -W 5 caps the ring buffer at 5 files total, so the oldest file is overwritten once that limit is reached rather than filling the disk indefinitely; this produces dns_capture.pcap0 through dns_capture.pcap4 cycling in place, a fixed roughly-500-megabyte disk footprint instead of an unbounded one.
Reading the transaction ID and flags in the capture
Every DNS message carries a 16-bit transaction ID, chosen by the client on the query and echoed back unchanged by the server on the response, which is how a client matches a response to the query that triggered it, especially important over UDP, since UDP itself carries no connection state to associate the two automatically. Alongside the ID, the DNS header carries a small set of flags: QR (0 for query, 1 for response), Opcode (almost always standard query in practice), AA (authoritative answer, set when the responding server is authoritative for that zone rather than just relaying a cached answer), TC (truncated, meaning the response was too large for UDP and the client should retry over TCP), RD (recursion desired, set by the client to ask the server to chase the answer itself rather than just returning what it already has), RA (recursion available, set by the server to say it is willing to do that), and RCODE (the response status, 0 for no error, 3 for NXDOMAIN, meaning the name does not exist, among others).
tcpdump decodes DNS natively without needing a hex dump: running it live (or reading back the capture with tcpdump -r dns_capture.pcap0 -n) prints lines like:
12:00:01.123456 IP 10.1.1.5.54321 > 8.8.8.8.53: 34521+ A? example.com. (29)
12:00:01.145678 IP 8.8.8.8.53 > 10.1.1.5.54321: 34521 1/0/0 A 93.184.216.34 (45)
34521 is the transaction ID, matching between the query and response lines, confirming they belong to the same exchange. The + after the ID on the query line indicates recursion desired (RD) was set. 1/0/0 on the response means one answer record, zero authority records, zero additional records. For deeper inspection than tcpdump's compact one-line summaries give you, such as fully decoded flag bits or malformed packets, reading the same capture in a protocol-aware analyzer like tshark or Wireshark against the saved .pcap file is the practical next step, since the capture format itself is standard and portable between tools.
Trade-offs and pitfalls
- Forgetting
-non a busy host generates a meaningful amount of extra DNS traffic just fromtcpdump's own reverse-lookup attempts on every captured address, which is a self-inflicted noise source specifically counterproductive while debugging DNS itself. -Cand-Wbound disk usage but do not bound time: on a very high-traffic host, a 5-file, 100MB-each ring buffer can rotate through its entire history in minutes, so if you are trying to catch an intermittent issue that only happens every few hours, this configuration will have already overwritten the relevant capture by the time you go looking; size the ring buffer to the actual traffic rate and the window of time you need to retain, not an arbitrary round number.- A filter this specific (single host, single port) is good practice for keeping capture size manageable and signal-to-noise high, but if the actual problem turns out to involve a different DNS server or a client the filter excludes, you will see nothing at all and can mistake an overly narrow filter for "there is no DNS traffic happening," which is a different conclusion than the correct one.
- Running
tcpdumpwith-wrequires elevated privileges (typically root, or a capability grant likeCAP_NET_RAWon the binary) and a filesystem location with enough free space for the configured ring buffer; on a host with tight partition sizing, writing the capture to/var/logalongside application logs risks filling that partition and affecting unrelated services, so a dedicated capture directory or partition is worth considering before running this on production infrastructure.
In a network where OSPF and BGP are both used and routes are redistributed in both directions on the same router, explain how redistribution can cause routing loops or feedback. Propose a configuration approach using route tagging, route-maps, and distribute-lists to prevent re-injection of the same routes into the origin protocol and show the logical steps you'd follow to implement it.
Sample Answer
Direct answer
Mutual redistribution fails because each protocol discards what the other needs to recognise a route's origin: OSPF carries no BGP attributes and BGP carries no OSPF route type or metric type. A prefix that left one protocol can therefore re-enter it, look like a fresh route, and then outlive its real source. The fix is to stamp every route with its origin as it crosses (an OSPF route tag in the BGP-to-OSPF direction, a BGP community in the OSPF-to-BGP direction), deny stamped routes in the opposite direction with route maps, and use a distribute-list to keep OSPF's copy of BGP-origin routes out of the routing tables of the border routers.
Two terms first. Redistribution means copying routes learned by one routing protocol into another. Administrative distance (AD) is the trust ranking a router uses when two protocols offer the same prefix: the lower number wins whatever the metrics are (Cisco defaults: eBGP 20, OSPF 110, iBGP 200). The core rule below is stamp on entry, deny on re-entry; the distribute-list is a local safety net on top.
How the feedback loop forms
Take two border routers R1 and R2, both running OSPF 1 and BGP AS 65000 and peering iBGP with each other, with R1 learning 203.0.113.0/24 from its ISP over eBGP.
- R1 redistributes BGP into OSPF, so the prefix becomes an OSPF external route.
- R2 learns it from OSPF and, also redistributing OSPF into BGP, re-originates it in BGP and advertises it back to R1 over iBGP.
- Both routers now hold two copies of the prefix, but the copies are compared differently on each router. On R2 the copies come from different protocols, so administrative distance decides: the iBGP copy from R1 (AD 200) loses to the OSPF external copy (AD 110), and R2 forwards by R1's OSPF advertisement, not by the real BGP route. On R1 both copies are BGP paths (the eBGP path from the ISP and the iBGP path from R2), so AD plays no part and the BGP best-path steps decide. A route that R2 originates into BGP carries an empty AS_PATH when sent to an internal peer (RFC 4271 section 5.1.2), so with equal weight and local preference R1 compares AS_PATH length and prefers R2's copy, whose path is empty, over the ISP's path, unless local preference is set to favour the ISP routes. The iBGP copy R2 sends to R1 was built from R1's own advertisement.
- Once R2's copy is best at R1 (or when the ISP withdraws the prefix and R2's copy is R1's only remaining BGP path), R1 sends traffic for the prefix to R2. R2's route still points back to R1 through OSPF, since R1's OSPF advertisement has not yet been withdrawn. Packets bounce R1 to R2 to R1. By default Cisco IOS does not redistribute iBGP routes into an IGP, so R1 then withdraws the OSPF copy, R2 loses its OSPF route, withdraws its BGP re-origination, and the loop clears after a delay. If the ISP route is still present, it becomes R1's best path again, is redistributed into OSPF again, and the cycle can repeat, so the prefix flaps. If someone enables
bgp redistribute-internal(which lets iBGP-learned routes be redistributed into an IGP) to make the setup "work", R1 re-injects R2's iBGP route into OSPF and the loop sustains itself: the prefix never ages out and packets loop until their TTL expires. - The same mechanism can also leak a route learned from ISP B out to ISP A as transit, because redistribution erases where the route came from.
Configuration approach (Cisco IOS)
Line by line: FROM-OSPF is the community that marks OSPF-origin routes. BGP-TO-OSPF refuses anything carrying that community and lets the rest into OSPF. OSPF-TO-BGP refuses anything carrying OSPF tag 65000 (a route that came from BGP) and stamps the rest with the community. BLOCK-BGP-ORIGIN finds OSPF copies carrying tag 65000 so the border router can refuse to install them. Under router ospf, tag 65000 stamps routes entering OSPF, metric-type 2 makes the external cost a fixed number that does not grow with internal link costs, and subnets includes subnets, not just classful networks (whole class A, B or C networks).
ip community-list standard FROM-OSPF permit 65000:100
!
! BGP into OSPF: skip anything that came from OSPF, mark the rest
route-map BGP-TO-OSPF deny 10
match community FROM-OSPF
route-map BGP-TO-OSPF permit 20
!
! OSPF into BGP: skip anything that came from BGP, mark the rest
route-map OSPF-TO-BGP deny 10
match tag 65000
route-map OSPF-TO-BGP permit 20
set community 65000:100
!
! on a border router, never install OSPF copies of BGP-origin routes
route-map BLOCK-BGP-ORIGIN deny 10
match tag 65000
route-map BLOCK-BGP-ORIGIN permit 20
!
router ospf 1
redistribute bgp 65000 subnets metric 100 metric-type 2 tag 65000 route-map BGP-TO-OSPF
distribute-list route-map BLOCK-BGP-ORIGIN in
!
router bgp 65000
address-family ipv4 unicast
redistribute ospf 1 match internal external 1 external 2 route-map OSPF-TO-BGP
neighbor 10.0.0.2 next-hop-self
neighbor 10.0.0.2 send-community
(10.0.0.2 stands for the other border router's iBGP address.)
Logical steps to implement
- List every route source and every crossing point. Here: OSPF to BGP and BGP to OSPF on R1 and R2.
- Choose one marker per direction that survives the crossing. The OSPF route tag is carried inside the external LSA (it appears as the External Route Tag field, a 32-bit label carried in the external LSA, in the OSPF database). BGP has no route tag, so a community is used there.
- Stamp on entry.
tag 65000onredistribute bgp;set community 65000:100inOSPF-TO-BGP. The community is only propagated to iBGP peers ifsend-communityis configured on that neighbor, which is why that line is in the BGP block. - Deny on re-entry.
match tag 65000denies BGP-origin routes going back into BGP;match community FROM-OSPFdenies OSPF-origin routes going back into OSPF. The community deny does most of its work when iBGP routes are redistributed into OSPF (bgp redistribute-internal) or when OSPF-origin routes arrive back from another BGP speaker. - Keep OSPF's BGP-origin copy out of the border routers' routing tables.
distribute-list route-map ... inblocks installation in the routing table of the router it is configured on and does not remove the LSA from the OSPF database, so the LSA is still flooded and internal OSPF routers still receive the route. The border routers keep using the real BGP route, which removes the 110-versus-200 distance trap. That route must be usable, so each border router setsnext-hop-selftoward the other (in the configuration above): without it the iBGP copy keeps the ISP's address as its next hop, and a router with no route to that address cannot install it, leaving the prefix with no route once the OSPF copy is blocked. Check whichmatchtypes your IOS release accepts in an OSPF inbound distribute-list route-map before relying onmatch tagthere; if it is not accepted, match the BGP-origin prefixes with a prefix-list inBLOCK-BGP-ORIGINinstead. - Set metrics deliberately.
metric 100andmetric-type 2are given so the external cost does not depend on a default; note that when no metric is specified OSPF assigns 20 to redistributed routes from most sources and 1 to BGP routes, and without thesubnetskeyword only classful networks are redistributed. - Optionally add a prefix-list backstop to either route map so only expected prefixes can ever cross.
Verification
show ip ospf database external 203.0.113.0on any internal router lists External Route Tag 65000.show ip bgp 10.20.0.0/16, for an OSPF-origin prefix such as that one, shows community 65000:100 on both border routers.- Fail the ISP link in a maintenance window and confirm the prefix disappears from R1, R2 and the OSPF database within the expected convergence time and that no traceroute oscillates between the border routers.
Pitfalls
- Tags must be unique per direction and per border-router pair, or a tag reused for another purpose silently drops legitimate routes.
- Single-point redistribution (only one border router redistributes each direction) removes the loop class at the cost of redundancy; the tag design is what makes two-router redistribution safe.
- A distribute-list filters by what is installed locally; it does not stop the neighbor from advertising the route onward, so it is a local safety net, not a replacement for the route maps.
You must connect three regions with expected inter-region traffic 10 Gbps and occasional spikes to 20 Gbps. Compare cloud dedicated interconnect (Direct Connect/ExpressRoute), site-to-site VPN, and public internet peering. Propose an architecture balancing cost, reliability, and security for consistent replication and workload movement.
Sample Answer
Direct answer
Provision dedicated interconnects (a cloud provider's private circuit product, such as AWS
Direct Connect or Azure ExpressRoute) sized for the 20 Gbps burst, not just the 10 Gbps
steady state, using two bonded 10 Gbps ports per critical path (a link aggregation group,
or LAG, which combines multiple physical connections into one logical one) rather than a
single larger port. That gives you 20 Gbps of capacity with a graceful-degradation property
that a single circuit doesn't: if one port fails, the remaining 10 Gbps still fully covers
steady-state traffic, only the burst headroom is lost until the port is repaired.
Sizing the topology
With 3 regions needing to exchange replication and workload-movement traffic, first
establish whether the traffic pattern is full mesh (every region talks meaningfully to every
other region) or hub-and-spoke (one region is the primary, the other two mostly talk to it).
The question's framing (consistent replication and workload movement across three regions)
points toward full mesh being the safer default unless you know the actual traffic pattern
is asymmetric. For a full mesh, that means 3 pairwise links, each sized for the same 10 Gbps
steady / 20 Gbps burst profile unless per-pair traffic is known to differ.
Say out loud which reading of "10 Gbps expected, 20 Gbps spike" you are sizing to, because
the two readings differ by a factor of three in circuit cost and the question does not settle
it. If those figures are the aggregate across all inter-region traffic, then three 20 Gbps
LAGs provision 60 Gbps of capacity for a 20 Gbps worst case, which is a lot of money for
headroom you will never use at once. If they are the per-pair figures, the same build is
correctly sized. Ask. If nobody knows, start from the aggregate reading (a 2x10 Gbps LAG on
the busiest pair, VPN or a single port on the others) and let measured per-pair utilization
pull you toward the full build, rather than buying three full-size circuits on day one.
For each pairwise link: 2x 10 Gbps dedicated ports in a LAG gives 20 Gbps aggregate
capacity. At steady 10 Gbps, the link runs at roughly 50% utilization, leaving headroom for
the stated burst to 20 Gbps without needing to provision a third port. If one port in the
LAG fails, the remaining single 10 Gbps port still fully absorbs the steady 10 Gbps load,
the system degrades from "handles bursts" to "handles steady state only," rather than
failing outright.
Comparing the three connectivity options for this load
Dedicated interconnect (Direct Connect / ExpressRoute). Fits this workload well: the
traffic is continuous, predictable in its steady/burst envelope, and sensitive to jitter for
replication consistency. A reserved circuit gives a bandwidth guarantee and a stable latency
profile that a shared path can't, which matters when the workload is "consistent
replication," not occasional best-effort transfers.
Site-to-site VPN. Could technically carry this traffic, but VPN throughput is bounded by
the encryption/decryption capacity of the endpoint devices doing the tunneling, and at
10-20 Gbps sustained, that endpoint capacity (not the link) is very often the actual
bottleneck; VPN is a better fit as a backup path than as the primary carrier here.
Public internet peering. Cheapest to start, but offers no bandwidth guarantee and
variable latency, both of which are risky for a workload described as needing "consistent"
replication; reserve this as a fallback path for non-critical or lower-priority traffic, not
the primary link for the replication workload itself.
Recommended architecture
Primary: dedicated interconnect (2x10 Gbps LAG per region pair, full mesh) carrying
replication and workload-movement traffic. Secondary: a site-to-site VPN path per pair as an
automatic failover if the dedicated link degrades or is fully down, accepting reduced
throughput during that window rather than losing connectivity entirely. Balance cost by not
over-provisioning: a 20 Gbps LAG per pair matches the stated burst exactly rather than
padding further, and utilization monitoring should trigger a capacity review (adding a third
port, or re-evaluating whether full mesh is actually needed) if sustained traffic trends
toward the burst ceiling rather than staying an occasional spike.
Trade-offs and pitfalls
- Don't size only for the steady state. A single 10 Gbps port per pair meets the average
but leaves zero headroom the moment the described 20 Gbps spike occurs, causing queuing,
increased replication lag, or dropped connections exactly when the system is under the
most load. - A LAG of two same-speed ports is not the same as one port at double the speed
operationally: it adds the resilience property (partial capacity survives a single port
failure) that a bigger single port does not, at a similar aggregate cost, which is usually
worth the design choice. - Full mesh costs more circuits than hub-and-spoke, and the gap widens fast with region
count. At three regions the difference is modest: full mesh needs 3 pairwise circuits
against hub-and-spoke's 2, so 50% more, not three times as many. The reason to settle the
traffic pattern anyway is what happens next: full mesh grows as N(N-1)/2 while
hub-and-spoke grows as N-1, so the same decision at six regions is 15 circuits against 5.
If traffic really is concentrated through one primary region, confirm that before
committing to 3 separate pairwise circuits; building unnecessary full mesh is a real,
ongoing cost for capacity you don't use, and it is the decision that compounds as regions
are added. The counterweight is that hub-and-spoke makes the hub a single point of
failure for region-to-region traffic and adds a hop of latency to every spoke-to-spoke
path, so the honest framing is cost against blast radius, not cost alone.
Tell me about a mentoring relationship that didn't go the way you hoped, one where your mentee didn't improve, or where things ended badly. What would you do differently now?
Sample Answer
Direct answer
A mentoring relationship going badly is rarely one big failure; it's usually a slow accumulation of choices, like taking on too much of the work yourself to protect the outcome, that quietly undercut the mentee's growth. The honest answer names a specific relationship, is candid about what you did (not just what the mentee did), and shows what changed in how you mentor afterward.
What "went badly" usually looks like
- Common patterns: being too directive and doing the hard parts yourself to protect delivery; giving feedback too infrequently or too late to be actionable; misjudging the mentee's actual gap (treating a confidence problem as a skill problem, or the reverse); or disengaging when the relationship got effortful.
- A strong answer picks one specific pattern and owns your part in it, rather than a vague "they weren't a good fit."
What separates a senior answer from a junior one
- Junior answers blame the mentee ("they just weren't receptive") or stay abstract ("communication could have been better"). Senior answers identify a decision you made and trace its actual effect: what you did, what it produced, and why it made sense to you at the time even though it was wrong.
- Senior answers also show what changed structurally afterward, not just an apology or a resolution to "communicate better." Concrete changes: an explicit mentoring agreement up front, checkpoints instead of open-ended availability, deliberately handing over ownership even when it's slower.
How to close it out
- End on what you'd do differently now, stated specifically enough that it's clear you'd actually behave differently in the next relationship, not just that you feel bad about the last one.
Worked example
During a stretch project with a hard deadline, I mentored a junior engineer by taking over the riskiest parts myself rather than coaching them through it, to keep the timeline safe. That worked in the short term, but it meant they never built confidence handling ambiguity or incidents on their own, and toward the end of the project they told me directly that they felt sidelined rather than developed. That was the moment it became clear the relationship hadn't done what I'd intended, even though the project itself shipped fine.
What I changed afterward: instead of stepping in when something got risky, I started requiring myself to narrate my reasoning out loud and have the mentee drive, only taking over if there was a genuine, immediate risk. I also set an explicit checkpoint (a short regular sync, not just "come find me") so growth stalls would surface early instead of only becoming visible at the end of a project. The relationship after that wasn't measured by how smoothly the project went; it was measured by whether the mentee could handle the next similar situation without me in the room, which is a slower thing to build but the actual point of mentoring.
Trade-offs and pitfalls
- The tempting failure mode is optimizing for the deliverable (visible and rewarded) at the expense of the mentee's growth (slower and less visible), especially under deadline pressure.
- Being self-critical is necessary but insufficient; an answer that's all remorse with no concrete process change reads as unreflective in a different way.
- Watch for over-correcting into never stepping in, which just replaces one failure mode (too directive) with another (abandoning someone to a mistake they can't yet recover from alone).
Explain NAT traversal problems for peer-to-peer applications (VoIP, WebRTC) and describe how STUN, TURN, and ICE work together to establish connectivity. Include an explanation of the security considerations for running a TURN server and how you would authenticate and limit abuse.
Sample Answer
Devices behind Network Address Translation (NAT) don't have a stable, publicly reachable address, which breaks peer-to-peer (P2P) protocols like Voice over IP (VoIP) and Web Real-Time Communication (WebRTC) that need two endpoints to connect to each other directly. Session Traversal Utilities for NAT (STUN), Traversal Using Relays around NAT (TURN), and Interactive Connectivity Establishment (ICE) work together to solve this.
How they work together
- STUN lets a client ask a public server "what address and port does the outside world see when I send from here," discovering its own NAT-mapped public address so it can share it with the other peer.
- TURN provides a fallback relay: when direct connectivity genuinely isn't possible, both peers send their traffic through a TURN server, which forwards it between them. This costs bandwidth and adds latency but always works.
- ICE is the framework that ties both together: each peer gathers every candidate address it has (its local address, its STUN-discovered public address, and a TURN-relayed address), exchanges the full candidate list with the other peer, and tries them in priority order (direct paths first, relay last) until one actually connects.
Worked example
Peer A and peer B both sit behind "symmetric" NATs, each of which assigns a different external IP and port mapping for every distinct destination. Peer A discovers a public mapping by talking to the STUN server, but that mapping was created for the STUN server's address specifically; when A later tries to reach B, A's NAT allocates a fresh, different mapping, so the address A advertised from STUN is not one B's packets are allowed to use. The same is true in reverse for B. Because neither side can predict the per-destination port the other's NAT will assign, every server-reflexive (STUN) candidate pair fails its connectivity check in both directions. ICE exhausts the direct candidates and falls back to the TURN-relayed candidate, so media between A and B flows through the TURN server, which has a stable public address both peers can always reach. A full-cone NAT would not force this: a full-cone mapping accepts inbound packets from any source once it has been opened by any outbound packet, so even a symmetric peer could initiate directly to a full-cone peer's STUN-discovered address, and the full-cone side would simply reply to the per-destination mapping the symmetric side just created. That is why the case that genuinely requires a relay is symmetric-to-symmetric, not full-cone-to-symmetric.
TURN server security
Because a TURN server relays arbitrary traffic on behalf of clients, it can be abused as an open relay/proxy if left unauthenticated, similar in spirit to an open mail relay. Authenticate with short-lived, time-limited credentials rather than a static shared password: a common approach (used by the long-term-credential mechanism in the TURN specification) derives a username/password pair from a server-side secret and a timestamp, so credentials expire automatically and a leaked credential has a short useful window. Limit abuse further by capping per-user bandwidth and concurrent relay allocations, restricting which destination ports/protocols the relay will forward, and rate-limiting new allocation requests per source IP, combined with monitoring for any single client sustaining unusually high relay volume.
Describe the TCP three-way handshake in detail: which flags are set in each packet (SYN, SYN-ACK, ACK), how sequence and acknowledgment numbers are used, and what state each endpoint moves into after each step. Explain what problem the handshake actually solves.
Sample Answer
Direct answer
The TCP three-way handshake establishes a reliable connection before any data flows: the client sends a SYN, the server replies with a combined SYN-ACK, and the client finishes with an ACK. Its job is to let both sides agree on starting sequence numbers and confirm that both directions of the path actually work before committing application data to the wire.
Structured elaboration
- SYN: the client picks an initial sequence number (ISN, essentially a large pseudo-random 32-bit number) and sends a segment with the SYN flag set and that sequence number. The client moves to the
SYN-SENTstate. - SYN-ACK: the server, if it's listening on that port, picks its OWN initial sequence number, and replies with a segment that has both the SYN flag set (announcing the server's own sequence number) AND the ACK flag set (acknowledging the client's sequence number + 1). The server moves to the
SYN-RECEIVEDstate. - ACK: the client acknowledges the server's sequence number + 1 with a plain ACK segment. Both sides now move to
ESTABLISHED, and either side may now send data.
Why three steps rather than two: TCP needs BOTH sides' sequence numbers acknowledged, since TCP is full-duplex (both directions need independent sequence tracking). A two-way handshake could confirm only one direction; the third message is what confirms the client's original SYN actually arrived, closing the loop for the client's own sequence space.
Worked example
Suppose a client opens a TCP connection to a web server on port 443. The client sends SYN, seq=1000. The server responds SYN, ACK, seq=5000, ack=1001 (acknowledging the client's SYN by number+1). The client responds ACK, seq=1001, ack=5001. From this point, the client's next data byte will carry sequence number 1001, and the server's next data byte will carry sequence number 5001; each side tracks the OTHER side's sequence space independently via the ACK field of every following segment.
Trade-offs & pitfalls
A frequent mistake is describing the handshake as three round trips; it's actually one and a half round trips of latency, because the SYN-ACK piggybacks the server's SYN onto its ACK of the client's SYN. This is also exactly why TCP always incurs at least one round trip of setup latency before any data can flow, which is the whole motivation behind newer mechanisms like TCP Fast Open that try to send data alongside the very first SYN.
Design network isolation for a multi-tenant environment where tenants may have overlapping IP address space, each must be strictly isolated, and a few shared services (DNS, NTP) stay reachable. How do you handle controlled cross-tenant connectivity?
Sample Answer
Direct answer
Give every tenant its own VRF (virtual routing and forwarding instance: a private routing table, so two tenants can both use 10.0.0.0/24 without conflict) and its own VXLAN segments (Virtual Extensible LAN, defined in RFC 7348, which carries each tenant's Layer 2 segments over a shared IP fabric, identified by a 24-bit VNI, the VXLAN network identifier). No tenant VRF imports routes from another (a route leak is a route from one VRF being copied into another so traffic can cross; VRFs share routes by matching route targets, which are tags attached to routes saying which VRFs may import them, so importing another tenant's target would copy that tenant's routes in). Shared services (DNS and NTP) sit in their own VRF at provider-owned addresses that cannot collide with tenant space, and each tenant reaches them only through a per-tenant NAT (network address translation) on a firewall, so overlapping tenants appear to the services as unique addresses. Cross-tenant connectivity is an exception, built as a firewalled, NAT-published, directional, logged and time-limited rule, never as a route leak.
Isolation layers
| Layer | Control | What it stops |
|---|---|---|
| Layer 2 | One VNI per tenant segment; a server port belongs to exactly one VNI | Frame leakage; RFC 7348 notes that overlapping MAC addresses across segments never have traffic cross over because the VNI isolates it |
| Layer 3 | One VRF per tenant, no route import between tenants | Routing between tenants, and any ambiguity from overlapping prefixes |
| Policy | Per-tenant firewall zone, default deny | Misconfigured routes; every exception is explicit |
| Anti-spoof (spoofing is sending with a source address that is not yours) | Leaf port (the top-of-rack switch port a server plugs into) accepts only the tenant's own subnets as source | A tenant sending with another tenant's addresses |
| Fabric | Tenants cannot reach underlay (the physical IP network that carries the VXLAN tunnels) or management addresses; per-tenant rate limit at the leaf | Attacks on the fabric, noisy neighbours |
Tenant numbering for the 50-tenant colocation: tenant n gets a Layer 3 VNI of 50000 + n and up to three Layer 2 VNIs of 10000 + 10n + j (j = 1 to 3). That is 200 VNIs, all unique, out of 2^24 = 16,777,216 available, against 4,094 usable VLAN IDs. Fifty tenants fit in VLAN space, but a VLAN ID is a 12-bit, fabric-wide resource shared with infrastructure, while a VNI is scoped per tenant segment with room for growth.
If the platform has no VXLAN, use VRF-lite (a VRF per tenant, with 802.1Q subinterfaces on each link: a subinterface is a logical interface on a physical link that carries only frames tagged with one VLAN ID, so each tenant gets one per link): the same isolation, but each link must be configured per tenant, so it scales worse.
Shared services with overlapping tenants
The problem is not reachability; it is that fifty tenants can all send from 10.0.0.5. If they reached the DNS server directly, the server could not tell them apart and a reply could not find its way back. So:
graph LR
A["Tenant 1 VRF"] --> FW["Firewall, per-tenant NAT"]
B["Tenant 2 VRF"] --> FW
C["Tenant 50 VRF"] --> FW
FW --> S["Shared-services VRF"]
S --> D["DNS 203.0.113.2, .3"]
S --> N["NTP 203.0.113.4, .5"]
| Item | Plan | Check |
|---|---|---|
| Shared-services subnet | 203.0.113.0/28: firewall .1, DNS .2 and .3, NTP .4 and .5 | 14 usable addresses, 5 used |
| Per-tenant NAT address | Tenant n is NATed to 203.0.113.(63 + n) from 203.0.113.64/26 | Tenant 1 gets .64, tenant 50 gets .113; 50 of 64 addresses (78%) used |
| Policy | Permit only DNS (port 53) and NTP (UDP 123) to the services addresses | Anything else from a tenant drops |
The block 203.0.113.0/24 (RFC 5737) is reserved for documentation, so it is a stand-in here. In a real deployment, substitute a public range registered to the provider; the plan needs service addresses that cannot collide with tenants' private ranges, and the contract must forbid tenants from using those addresses internally. A tenant's host reaches 203.0.113.2 by a route in its VRF pointing at the firewall; the firewall translates its source to the tenant's unique address and forwards into the services VRF; the services VRF has a route for the NAT pool back to the firewall, which reverses the translation into the right tenant VRF.
A concrete trace, with illustrative addresses. Tenant 7's host 10.0.0.5 queries DNS at 203.0.113.2.
- The host sends source 10.0.0.5, destination 203.0.113.2. Tenant 7's VRF has a route for the services subnet pointing at the firewall.
- The firewall rewrites the source to tenant 7's address, 203.0.113.(63 + 7) = 203.0.113.70, and also rewrites the source port so one tenant address serves all of that tenant's hosts. It forwards source 203.0.113.70, destination 203.0.113.2 into the services VRF. Tenant 1's own 10.0.0.5 would leave as 203.0.113.64, so the DNS server sees two different clients.
- DNS replies with source 203.0.113.2, destination 203.0.113.70. The services VRF has a route for 203.0.113.64/26 pointing at the firewall.
- The firewall finds 203.0.113.70 and the port in its translation table, rewrites the destination back to 10.0.0.5 and delivers it into tenant 7's VRF.
This pool is 78% full at 50 tenants. The pool is exhausted at tenant 64, so order the second /26 when only 10 addresses are left (after tenant 54).
Controlled cross-tenant connectivity
Peering is an exception with a record: both tenants approve in writing, the rule lives in the source of truth (the inventory database that configurations are generated from), and an expiry date is set. It works through the same firewall:
- The destination tenant publishes only the specific service (for example a /32, a single-host route, and one port) at an address from a dedicated peering block, translated to its real address inside its own VRF. Illustrative example: tenant 3's database at 10.0.5.20 port 5432 is published as 198.51.100.17 port 5432 from the peering block 198.51.100.16/28 (RFC 5737 documentation space again). Tenant 1's host 10.0.0.5 connects to 198.51.100.17; the firewall translates the source to 203.0.113.64 and the destination to 10.0.5.20 and delivers it into tenant 3's VRF, so tenant 3's database logs a client at 203.0.113.64, never at 10.0.0.5.
- The source tenant is NATed to a unique address as above.
- The rule is directional (in the example above, tenant 1 to tenant 3 on TCP 5432 only), logged, and removed on expiry.
Never solve this by importing one tenant's route targets into another: with overlapping prefixes the leak is ambiguous and breaks isolation. If both tenants have non-overlapping space, routing their published prefixes through the firewall is acceptable, still not direct. Pairs grow fast (C(50, 2) = 1,225 possible), so keep peerings few and reviewed.
Operational details
- MTU: VXLAN wraps the tenant's whole Ethernet frame in an outer IP (20), UDP (8) and VXLAN (8) header, 36 bytes, and the tenant frame carries its own 14-byte Ethernet header, so a 1,500-byte tenant packet becomes 1,500 + 14 + 36 = 1,550 bytes of underlay IP packet. The usual quote of 50 bytes of overhead adds the outer Ethernet header (14 + 36) and matters for the frame size on the wire (1,514 + 50 = 1,564 bytes), while the underlay IP MTU needed is 1,550. Set the underlay to jumbo frames (Ethernet frames larger than the standard 1,500-byte payload, often 9,000 bytes) and test with a don't-fragment probe.
- Verification as code: place a canary host (a small test server) in each tenant, all with the same address (say 10.0.0.5), each answering with its tenant ID. From each tenant, probe 10.0.0.5 and confirm the answer carries your own ID: that proves overlapping space stays separate. Then test all 50 x 49 = 2,450 directed tenant pairs and expect failure except approved peerings. Run after every change.
The same design in a public cloud
Use one VPC (virtual private cloud, a tenant's private network inside the cloud) per tenant. AWS Transit Gateway (a regional router that connects VPCs) documentation says it does not support routing between VPCs with identical or overlapping CIDRs: if a newly attached VPC's CIDR is identical to or overlaps one already attached, its routes are not propagated to the transit gateway route table. Overlapping tenants therefore cannot be joined through a transit gateway, so shared services use AWS PrivateLink (which publishes a service to other VPCs through a private endpoint, without routing the networks together), which allows consumer and provider VPCs to have overlapping IP ranges; each tenant gets an endpoint with an address inside its own VPC.
You find an internal host beaconing to a suspicious internal IP in a different network zone, a sign of active lateral movement. Draft a containment plan using segmentation controls (access rule changes, microsegmentation, host-based firewall policy) that stops the spread while minimizing disruption to legitimate traffic, and describe how you would verify containment actually held.
Sample Answer
Contain fast without destroying evidence: isolate the host at the segmentation layer, not by powering it off, tighten its reachability to nothing except a monitored forensics path, and verify containment by confirming, from telemetry outside the host itself, that the beaconing traffic has actually stopped and the host can no longer reach anything it previously could.
Step 1: isolate without destroying evidence
Rather than shutting the host down, which can lose volatile evidence such as in-memory malware artifacts, or manually killing the suspicious process, which can tip off active command-and-control (some malware has dead-man-switch behavior), move the host into a quarantine segment or apply a host-based firewall policy that denies essentially all outbound and inbound traffic except a narrow, monitored path to incident-response tooling.
Step 2: contain with layered segmentation controls
Combine several levers rather than relying on one: revoke or change the host's existing access-rule membership (an access-rule change removing it from whatever security group previously granted it broad reach); apply microsegmentation-style explicit deny rules for the specific suspicious internal address and any other destinations flagged during triage; add a host-based firewall policy on the host itself as a second, independent layer; and, if the host holds a workload identity or certificate, revoke it so even a valid-looking authenticated request from it is rejected by other services' policy.
Step 3: narrow the blast radius further
Rotate any credentials or secrets the host had access to, on the assumption that isolation stops FUTURE misuse but doesn't undo anything already taken.
Step 4: verify containment actually held
Don't rely on "I applied the rule" as proof. Confirm from independent telemetry, network flow logs, the destination's own connection logs, or the segmentation control plane's enforcement confirmation, that the specific beaconing pattern has stopped appearing after the change, that the host can't reach any segment or service it could reach before, and that no OTHER host has started showing a similar beaconing pattern, which would indicate the compromise had already spread before containment.
Worked example
A host is observed beaconing every few minutes to a suspicious internal address in the payments segment. Containment: move the host's security-group membership from its normal tier to a quarantine group that denies all except a forensics jump host; add an explicit deny rule for the specific destination address at the payments segment boundary as a second layer, in case the quarantine change is delayed or incomplete; and revoke the host's workload certificate so any request it still manages to send is rejected by identity-aware policy on the receiving end, not just blocked at the network. Verification, roughly fifteen minutes later: flow logs show zero connections from the host to the previously targeted address, and payments-segment access logs show zero requests bearing the host's now-revoked identity, confirming both the network path and the identity path are closed, not just one of the two.
Trade-offs and pitfalls
Isolating too aggressively, killing the process or shutting the host down, can destroy forensic value and, in some cases, trigger a scripted destructive response from the malware faster than a quiet network-level isolation would; isolating too slowly to preserve evidence risks continued lateral movement while you wait. Most incident-response playbooks resolve this by favoring immediate network-level containment, fast and low-risk of tipping off the attacker, while deferring host-level forensic actions like memory capture or a process kill to a separate, deliberate step once network isolation is confirmed.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Network Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs