FAANG-Standard Network Engineer Interview Preparation Guide (Entry Level)
This guide is based on general FAANG interview practices and may not reflect specific company procedures.
Entry-level Network Engineer interviews at FAANG companies follow a rigorous but foundational assessment process designed to verify strong understanding of core networking principles, hands-on technical competency, systematic troubleshooting ability, and cultural fit. The interview pipeline progresses from initial recruiter screening through multiple technical assessments covering fundamental concepts and hands-on skills, culminating in behavioral evaluation and role alignment discussion. For entry-level positions, emphasis is placed on learning potential, attention to detail, problem-solving methodology, and ability to work independently with guidance rather than extensive production experience.
Interview Rounds
Recruiter Phone Screen
What to Expect
Initial conversation with a recruiter to assess background, motivation, and baseline fit for the Network Engineer role. This 15-30 minute call covers your path to networking, understanding of the role responsibilities, and initial technical readiness. Recruiter evaluates your communication skills, enthusiasm for the role, and availability. This is also an opportunity for you to ask clarifying questions about the position, team structure, and company expectations. Recruiter will ask about your educational background, any certifications, relevant projects or internships, and why you're interested in this specific company.
Tips & Advice
Prepare a clear, concise 2-3 minute summary of your background and why you're pursuing network engineering. Demonstrate genuine enthusiasm for the role and company. Have 2-3 thoughtful questions prepared about the team, technical stack, and growth opportunities. Be honest about your experience level—recruiters appreciate candidates who are realistic about entry-level status while showing eagerness to learn. Highlight any hands-on experience, certifications (CompTIA A+, Network+), or personal projects involving networking. Research the company's network infrastructure focus areas beforehand if possible.
Focus Topics
Motivation and Company Alignment
Articulate why you're interested in this specific company and role. Research the company's technology focus, culture, and any public information about their network infrastructure or engineering team.
Practice Interview
Study Questions
Certifications and Technical Foundation
Discuss any relevant certifications (CompTIA Network+, A+, CCNA studies) or hands-on experience with networking labs, home networks, or academic projects. Be honest about which areas you're strongest in and where you're still learning.
Practice Interview
Study Questions
Role Understanding and Expectations
Demonstrate understanding of Network Engineer responsibilities at an entry level, including network monitoring, troubleshooting, equipment configuration basics, and working with security technologies. Show awareness of what you'll be learning and growing into.
Practice Interview
Study Questions
Background and Career Path
Your educational background, relevant certifications, internships, or projects that led you to pursue network engineering. Be prepared to explain your understanding of what network engineers do and why this role appeals to you specifically.
Practice Interview
Study Questions
Technical Fundamentals Phone Screen
What to Expect
45-60 minute technical phone interview focused on core networking fundamentals. This round tests your understanding of foundational concepts essential to network engineering. Expect questions about the OSI model, TCP/IP protocol suite, IP addressing and subnetting, basic routing concepts, and network services like DNS and DHCP. The interviewer will ask both theoretical questions and practical scenarios. For entry-level candidates, strong fundamentals are more important than memorized facts. Interviewers evaluate your ability to think through problems, explain concepts clearly, and demonstrate solid foundational knowledge. You may be asked to work through subnetting problems or explain how network communications work at different layers.
Tips & Advice
Focus your preparation on truly understanding concepts rather than memorization. Be able to explain the OSI model clearly and describe what happens at each layer. Practice subnetting calculations until they become second nature. Understand TCP/IP flow and be able to trace how data moves from application to physical layer. When asked a question you're unsure about, think out loud and work through it systematically—this demonstrates problem-solving approach, which is valued for entry-level candidates. Use a whiteboard or paper to sketch out concepts if needed. If you don't know something, say so honestly and ask clarifying questions. Prepare 3-4 real-world networking scenarios (like 'what happens when you type a URL') and practice explaining them at different levels of detail.
Focus Topics
Network Topologies and Architecture
Understanding of basic network topologies (star, mesh, ring, bus) and when each is appropriate. Concepts of LAN (Local Area Network) vs WAN (Wide Area Network), network segmentation, and basic network architecture concepts.
Practice Interview
Study Questions
Routing Fundamentals
Basic understanding of routing concepts including routing tables, routing decisions based on longest prefix match, static vs dynamic routing introduction, and how routers forward packets. Know basic routing protocols at high level: RIP, OSPF, EIGRP concepts.
Practice Interview
Study Questions
DNS and DHCP
How DNS (Domain Name System) resolves domain names to IP addresses. Understand DNS query process, record types, and hierarchy. DHCP (Dynamic Host Configuration Protocol) basics: how it assigns IP addresses, lease process, and how it simplifies network management.
Practice Interview
Study Questions
OSI Model and Network Layers
Deep understanding of the seven-layer OSI model including Physical, Data Link, Network, Transport, Session, Presentation, and Application layers. Know what happens at each layer, which devices/protocols operate at each layer, and how data is processed as it moves through layers (encapsulation and decapsulation).
Practice Interview
Study Questions
TCP/IP Protocol Suite and Basics
Understanding of TCP/IP model vs OSI model, IPv4 and IPv6 addressing, protocol layering, and how the four layers of TCP/IP (Link, Internet, Transport, Application) correspond to networking concepts. Know basic protocols and their purposes: IP, ICMP, TCP, UDP, ARP.
Practice Interview
Study Questions
IP Addressing and Subnetting
Mastery of IPv4 address classes, subnet masks, CIDR notation, and subnetting calculations. Ability to determine network address, broadcast address, available hosts, and subnet ranges. Understanding of public vs private IP addresses and address conservation techniques.
Practice Interview
Study Questions
Network Technologies and Protocols Phone Screen
What to Expect
45-60 minute technical phone interview focused on networking technologies, protocols, and practical tools used in network operations. This round tests your knowledge of specific networking technologies, security fundamentals, network tools for diagnostics, and how different protocols work in practice. Expect questions about TCP vs UDP, network security basics, VPN concepts, firewalls, packet analysis, and common network diagnostic commands. The interviewer will present scenarios requiring you to recommend appropriate technologies or explain how specific protocols solve network problems. This round assesses both theoretical knowledge and practical understanding of how technologies are used in real networks.
Tips & Advice
Study the differences between connection-oriented and connectionless protocols deeply. Understand why you would choose TCP over UDP in specific scenarios and vice versa. Practice explaining network security concepts in simple terms—VPN, encryption, firewalls, and ACLs. Be familiar with common network diagnostic tools and commands: ping, traceroute, ipconfig/ifconfig, netstat, nslookup, arp. Know how to read packet traces and identify key information. When discussing security, think about the job description's emphasis on 'implementing network security measures'—be ready to discuss access control, encryption basics, and security protocols. If asked about a technology you're less familiar with, explain what you know and ask clarifying questions.
Focus Topics
Packet Analysis and Network Traffic Understanding
Basic understanding of packet structure, what information is in each layer's header, and how to analyze traffic captures. Know how to identify packet types, protocols, source/destination addresses, and ports. Understanding of network flows and communication patterns.
Practice Interview
Study Questions
Common Networking Protocols and Services
Understanding of commonly used protocols: HTTP/HTTPS, FTP, SSH, Telnet, SNMP, Syslog. Know the ports they use, their purposes, and when they're appropriate. Understand protocol layering and protocol suites.
Practice Interview
Study Questions
VPN Basics and Encryption
What VPNs (Virtual Private Networks) do and why they're used. Site-to-site vs remote access VPN concepts. Basic encryption principles: symmetric vs asymmetric encryption, SSL/TLS basics, and how encryption secures communications. Understanding of tunneling concepts.
Practice Interview
Study Questions
TCP vs UDP and Transport Layer Protocols
In-depth understanding of TCP (Transmission Control Protocol) as connection-oriented, reliable protocol and UDP (User Datagram Protocol) as connectionless, faster but unreliable protocol. Know when to use each, their respective advantages/disadvantages, header structures, and real-world applications. Understand ICMP for diagnostics.
Practice Interview
Study Questions
Network Diagnostic Tools and Commands
Practical knowledge of common network tools: ping (ICMP echo requests), traceroute (path tracing), ipconfig/ifconfig (IP configuration viewing), nslookup/dig (DNS queries), netstat (connection statistics), arp (address resolution protocol), and packet analysis tools like Wireshark. Know what each tool does, what output to expect, and how to interpret results.
Practice Interview
Study Questions
Network Security Fundamentals
Basic security concepts: confidentiality, integrity, and availability (CIA triad). Understanding of firewalls, access control lists (ACLs), encryption basics, and network security layers. Know what threats firewalls and ACLs protect against. Understand public key cryptography concepts at basic level.
Practice Interview
Study Questions
On-site Technical Assessment: Hands-on Troubleshooting and Lab
What to Expect
60-75 minute on-site interview combining hands-on lab work with troubleshooting scenarios. You'll be presented with network problems in a simulated or real lab environment and asked to diagnose and resolve issues. Common scenarios include connectivity failures, misconfigured devices, routing problems, or service unavailability. You'll have access to network equipment (routers, switches, or simulation tools) and be expected to use appropriate diagnostic tools to identify root causes and implement fixes. The interviewer observes your methodology, communication, technical skills, and ability to work through problems systematically. For entry-level candidates, the emphasis is on showing structured troubleshooting approach, using tools effectively, and thinking out loud about the problem.
Tips & Advice
Practice hands-on labs extensively before the interview using GNS3 or Cisco Packet Tracer. Develop a systematic troubleshooting methodology: gather information, identify symptoms, form hypotheses, test systematically, and implement solutions. When given a problem, start by understanding what's not working and what should be working. Use diagnostic commands confidently—ping, traceroute, show interfaces, show routing table, show IP route. Document your findings and explain your reasoning as you go. If you get stuck, ask clarifying questions rather than randomly trying things. Be methodical and logical. Practice explaining your thinking out loud—interviewers want to understand your problem-solving approach. Don't rush; accuracy and methodology matter more than speed for entry-level candidates.
Focus Topics
Performance Monitoring and Network Health Assessment
Understanding network performance metrics: bandwidth utilization, latency, packet loss, error rates. Using tools and commands to monitor performance: show interface statistics, network monitoring tools, understanding what normal performance looks like vs degraded performance.
Practice Interview
Study Questions
Lab Environment Navigation and Tool Proficiency
Comfort using GNS3, Cisco Packet Tracer, or real lab equipment. Ability to navigate interfaces, launch commands, interpret output, save configurations, and reset devices as needed. Understanding virtual networking in lab environments.
Practice Interview
Study Questions
Routing Troubleshooting and Verification
Examining routing tables, verifying static routes are configured correctly, understanding why packets aren't reaching destinations, diagnosing routing misconfigurations, and tracing routing paths. Using show ip route, traceroute, and related commands to troubleshoot routing problems.
Practice Interview
Study Questions
Cisco IOS Basics and Device Configuration
Basic Cisco Internetwork Operating System (IOS) commands for routers and switches: navigating command modes (user/enable/config), basic show commands (show interface, show ip route, show running-config), basic configuration commands (hostname, IP addressing, routing), and understanding command hierarchy.
Practice Interview
Study Questions
Systematic Network Troubleshooting Methodology
Structured approach to troubleshooting: define the problem clearly, gather symptoms, check physical connectivity, verify configurations, isolate the faulty component, identify root cause, implement fix, and verify the solution. Understanding of layers-based troubleshooting (start at physical layer, work up through OSI layers). Knowing when to escalate issues.
Practice Interview
Study Questions
Interface and Connectivity Troubleshooting
Diagnosing connectivity problems at interface level: checking interface status (up/down), verifying IP configuration correctness, checking for physical layer issues, understanding interface statistics (errors, collisions, drops). Troubleshooting ping failures, route reachability problems, and communication between network segments.
Practice Interview
Study Questions
On-site Technical Assessment: Network Configuration and Security
What to Expect
60-75 minute on-site interview focused on network configuration tasks, basic security implementation, and design thinking. You'll be given scenarios requiring you to configure network devices, implement security measures, or design small network changes. Scenarios might include: configure a new VLAN, set up basic ACLs for traffic filtering, implement security policies, design a small network segment for a new department, or configure interfaces for specific requirements. The interviewer assesses your ability to translate requirements into configurations, understand security implications of design choices, and think about scalability and best practices even at entry level. For entry-level candidates, the emphasis is on foundational security understanding and configuration capability.
Tips & Advice
Practice configuring VLANs, static routes, basic ACLs, and interface IP addresses in GNS3 or Packet Tracer extensively. Understand why you'd use each technology—for example, VLANs for segmentation and security. When given a requirement, ask clarifying questions before diving into configuration. Explain your design reasoning: why you're choosing specific technologies, how the design meets requirements, and what scalability considerations exist. Understand access control lists conceptually and practically—know what rules allow/deny and why you'd block specific traffic. For security questions, think about the CIA triad and how configurations support those principles. Start with fundamentals and work up in complexity. Document configurations and explain as you go.
Focus Topics
Small-Scale Network Design and Capacity Planning
Designing network segments for small departments or sites, considering current needs and modest growth, making technology choices (VLAN, subnetting, routing) appropriate to scale. Understanding when to upgrade, add redundancy, or change architecture. Basic capacity planning concepts.
Practice Interview
Study Questions
Network Device Firewalls and Security Appliances
Basic understanding of firewalls (stateful vs stateless), how firewalls protect networks, firewall policies and rule creation, common firewall technologies (packet-filtering, stateful inspection). Understanding where firewalls fit in network architecture.
Practice Interview
Study Questions
VLAN Configuration and Management
Creating VLANs for network segmentation, assigning ports to VLANs, configuring trunk ports, basic VLAN routing concepts, and understanding when to use VLANs. Know VLAN IDs, native VLAN concepts, and tagged/untagged traffic.
Practice Interview
Study Questions
Network Security Implementation Basics
Concepts of defense-in-depth, least privilege access, how to implement basic security controls, firewall rule concepts, and security policy implementation at network level. Understanding what threats ACLs and firewalls protect against.
Practice Interview
Study Questions
Router and Switch Configuration Fundamentals
Configuring IP addresses on interfaces, basic static routing, enabling interfaces, setting hostnames and administrative access, saving configurations. Understanding configuration modes and commands. Ability to verify configurations work as intended. Basic switch configuration including VLAN assignment and management.
Practice Interview
Study Questions
Access Control Lists (ACL) Basics
Understanding what ACLs do and why they're used for security. Basic ACL concepts: permit/deny rules, wildcards, port numbers, direction (inbound/outbound), sequence numbers. Ability to write simple ACLs that allow/deny specific traffic based on source, destination, protocol, or port.
Practice Interview
Study Questions
Behavioral and Cultural Fit Round
What to Expect
45-60 minute interview focused on behavioral competencies, teamwork, learning ability, and cultural alignment with FAANG values. Although specific company values weren't identified, FAANG companies typically evaluate how candidates demonstrate principles like customer focus, ownership, bias for action, learning and growing, earning trust, and delivering results in a team environment. For entry-level candidates, the focus is on learning potential, collaboration with team members, ability to handle challenges, and cultural fit. You'll be asked about past experiences (academic projects, internships, personal projects) that demonstrate these competencies. The interviewer evaluates communication skills, resilience, growth mindset, and ability to work effectively with others.
Tips & Advice
Prepare 4-5 specific stories from your background that demonstrate key competencies: learning ability (taking on new challenges, studying for certifications), teamwork (group projects, contributing to team success), handling failure (learning from mistakes, bouncing back), problem-solving approach, and adaptability. Use the STAR method: Situation, Task, Action, Result—be specific and quantifiable when possible. For entry-level, emphasize learning eagerness and growth mindset rather than achievements. Practice telling stories concisely in 2-3 minutes. Be genuine and authentic—FAANG companies value authenticity. Research company culture and values beforehand to align your responses. When asked about challenges, focus on what you learned rather than dwelling on the difficulty. Show enthusiasm for the role and company. Ask thoughtful questions that demonstrate understanding of what engineers do and company direction.
Focus Topics
Communication and Technical Clarity
Ability to explain technical concepts clearly to diverse audiences, document work clearly, ask good questions to understand requirements, and communicate status effectively. Demonstrating strong written and verbal communication skills.
Practice Interview
Study Questions
Handling Mistakes and Feedback
Stories about making mistakes, taking responsibility, learning from them, and improving. Demonstrating openness to feedback, ability to receive criticism constructively, and using feedback to improve performance.
Practice Interview
Study Questions
Adaptability and Handling Change
Stories about adapting to new situations, learning new technologies quickly, working in ambiguous environments, and managing change. Demonstrating flexibility and comfort with uncertainty.
Practice Interview
Study Questions
Learning Ability and Growth Mindset
Demonstrating curiosity, eagerness to learn new technologies, ability to study independently and gain new skills (e.g., pursuing certifications, building home labs, taking online courses). Stories about taking on unfamiliar challenges, making mistakes and learning from them, and adapting approach based on feedback.
Practice Interview
Study Questions
Problem-Solving Approach and Resilience
Demonstrating systematic thinking when facing challenges, persistence in working through difficult problems, ability to break down complex issues, and willingness to ask for help when appropriate. Stories about debugging network issues, overcoming obstacles, and pushing through challenges.
Practice Interview
Study Questions
Teamwork and Collaboration
Demonstrating ability to work effectively with others, support team members, communicate clearly, and contribute to shared goals. Stories about group projects, helping colleagues, receiving feedback well, and working toward team success despite challenges.
Practice Interview
Study Questions
Hiring Manager Round
What to Expect
30-45 minute final interview with the hiring manager or lead engineer for the team. This conversation focuses on role fit, your understanding of the position and team, alignment with team goals, and growth potential within the role and company. The hiring manager assesses whether you're ready for the position, whether you understand what you'll be doing day-to-day, and whether you'll grow into increasing responsibility. You'll discuss the team's challenges and projects, what success looks like in the first 90 days, how you'll be supported as entry-level, and career growth opportunities. This is also an opportunity to ask detailed questions about the role, team, and company. The manager evaluates both your technical capability and soft skills, ultimately determining if you're hire-worthy and will succeed in their team.
Tips & Advice
Research the team and its focus area if possible—look at company blog posts, LinkedIn profiles of team members, or project descriptions. Prepare specific, thoughtful questions about the role, team challenges, technology choices, and growth path. Listen carefully to the manager's description of the role and mirror that language back when discussing your interest. Demonstrate understanding of what day-to-day network engineering work involves based on the job description. Express enthusiasm for specific aspects of the role that appeal to you. Be honest about entry-level status while showing confidence in your ability to learn and contribute. Ask about support and mentoring for entry-level engineers. Discuss how you'll measure success in the first 90 days. Show that you want to understand and contribute to team goals. End by reiterating your interest and excitement about the opportunity.
Focus Topics
90-Day Success and Growth Path
Discussing what success looks like in your first 90 days, what skills you want to develop, and how you see yourself growing in the role and company. Demonstrating forward thinking and commitment to growth.
Practice Interview
Study Questions
Team and Company Alignment
Demonstrating interest in this specific team and company, not just any job. Showing that you've thought about how your goals align with team/company direction. Expressing genuine enthusiasm for the work the team does.
Practice Interview
Study Questions
Support and Learning Expectations
Asking thoughtful questions about how entry-level engineers are supported, mentoring structure, learning resources available, and realistic expectations for onboarding. Demonstrating that you understand you'll need support and guidance.
Practice Interview
Study Questions
Role Understanding and Day-to-Day Responsibilities
Demonstrating clear understanding of what you'll be doing as an entry-level Network Engineer on this team. Referencing job description responsibilities: network monitoring, troubleshooting, equipment configuration, security implementation, and documentation. Showing you understand the scope and have realistic expectations.
Practice Interview
Study Questions
Technical Readiness and Growth Potential
Demonstrating that you have foundational skills needed for the role and strong potential to grow. Discussing how you'll apply your preparation, what you're eager to learn, and how you'll contribute quickly. Showing realistic understanding of learning curve without underselling yourself.
Practice Interview
Study Questions
Frequently Asked Network Engineer Interview Questions
From the block 172.16.0.0/20, allocate the minimum-sized subnets required to support 12 separate networks each needing at least 150 usable hosts. For each allocated subnet provide the network address, CIDR mask, broadcast address, and usable host range. Explain your allocation algorithm and any trade-offs.
Sample Answer
Direct answer
Each network needs at least 150 usable hosts, and a subnet with h host bits has 2^h - 2 usable addresses (the network address and the broadcast address are not assignable to hosts). A /25 (h = 7) gives only 126, which is too few, so the minimum is a /24 (h = 8, 254 usable). Twelve /24s need 12 x 256 = 3,072 addresses, and the /20 has 4,096, so they fit with a /22 (1,024 addresses) left over: 172.16.12.0/22.
Allocation algorithm
- Find the smallest prefix that meets the requirement: smallest h with 2^h - 2 >= 150, so h = 8 and the prefix length is 32 - 8 = /24.
- Check capacity: the /20 holds 2^(24-20) = 16 subnets of /24, and 12 are needed.
- Allocate in ascending address order from the start of the block. Because every subnet is the same size, each one starts exactly where the previous one ended: a /24 is 256 addresses, so the third octet simply goes up by 1 (172.16.0.0, 172.16.1.0, 172.16.2.0, and so on). Every start address is therefore a multiple of 256, which is what a /24 boundary requires, so nothing is misaligned. (With mixed sizes you would sort largest first so each block stays aligned.)
- Summarise what is left: the remaining four /24s, 172.16.12.0 to 172.16.15.255, collapse into the single block 172.16.12.0/22.
The 12 subnets
Mask for every row: 255.255.255.0 (/24).
| # | Network | Broadcast | Usable host range |
|---|---|---|---|
| 1 | 172.16.0.0/24 | 172.16.0.255 | 172.16.0.1 to 172.16.0.254 |
| 2 | 172.16.1.0/24 | 172.16.1.255 | 172.16.1.1 to 172.16.1.254 |
| 3 | 172.16.2.0/24 | 172.16.2.255 | 172.16.2.1 to 172.16.2.254 |
| 4 | 172.16.3.0/24 | 172.16.3.255 | 172.16.3.1 to 172.16.3.254 |
| 5 | 172.16.4.0/24 | 172.16.4.255 | 172.16.4.1 to 172.16.4.254 |
| 6 | 172.16.5.0/24 | 172.16.5.255 | 172.16.5.1 to 172.16.5.254 |
| 7 | 172.16.6.0/24 | 172.16.6.255 | 172.16.6.1 to 172.16.6.254 |
| 8 | 172.16.7.0/24 | 172.16.7.255 | 172.16.7.1 to 172.16.7.254 |
| 9 | 172.16.8.0/24 | 172.16.8.255 | 172.16.8.1 to 172.16.8.254 |
| 10 | 172.16.9.0/24 | 172.16.9.255 | 172.16.9.1 to 172.16.9.254 |
| 11 | 172.16.10.0/24 | 172.16.10.255 | 172.16.10.1 to 172.16.10.254 |
| 12 | 172.16.11.0/24 | 172.16.11.255 | 172.16.11.1 to 172.16.11.254 |
Spare: 172.16.12.0/22 (172.16.12.0 to 172.16.15.255), room for four more /24 networks, or any mix of smaller blocks.
Checking the arithmetic
A short Python check regenerates the table from the rule, so no row depends on hand arithmetic:
import ipaddress
block = ipaddress.ip_network("172.16.0.0/20")
bits = 2
while 2**bits - 2 < 150:
bits += 1
subs = list(block.subnets(new_prefix=32 - bits))
for i, n in enumerate(subs[:12], 1):
h = list(n.hosts())
print(i, n, n.broadcast_address, h[0], h[-1])
print(list(ipaddress.summarize_address_range(subs[12].network_address, block.broadcast_address)))
Line by line: the while loop starts with 2 host bits and adds one until 2^bits - 2 reaches at least 150, which stops at 8. block.subnets(new_prefix=32 - bits) cuts the /20 into every /24 inside it, in address order (16 of them). The for loop prints the first 12; n.hosts() lists the usable addresses, so its first and last entries are the host range, and n.broadcast_address is the broadcast. The final line takes the range from the start of the thirteenth /24 to the end of the /20 and summarize_address_range returns the fewest prefixes that cover exactly that range. It prints [IPv4Network('172.16.12.0/22')], and the 12 rows match the table.
Trade-offs
- Waste is inherent, not a mistake. A 150-host requirement sits between /25 (126) and /24 (254). Each network leaves 254 - 150 = 104 usable addresses idle, 12 x 104 = 1,248 in total. That is the cost of the binary block sizes; no uniform prefix does better, and splitting a 150-host network across two /25s means two subnets (two gateways, two broadcast domains), which is worse operationally than the idle addresses.
- The idle addresses are the growth headroom. A network can grow from 150 to 254 hosts (about 69% more, (254 - 150) / 150) with no renumbering. Reserve the gateway and any redundancy-protocol addresses from that headroom (for example .1 gateway, .2 and .3 router interfaces) and it is still large.
- Keep the leftover block contiguous. Allocating from the bottom keeps 172.16.12.0/22 as one summarisable block; scattering the 12 subnets would fragment it.
- If "at least 150" later becomes 300, a network needs a /23 (510 usable), and a /23 must start on an even third octet (172.16.12.0/23 qualifies; 172.16.5.0/23 does not). Allocating adjacent /24s only helps if the pair is aligned, so decide before handing out the first block whether any network might double.
A stakeholder needs a number out of a part of the business you do not understand yet, and they need it this week. How do you get them something they can use without pretending to more certainty than you have?
Sample Answer
Direct answer
I give a bounded number fast rather than staying quiet while I chase precision I don't have time for: I state the number, the method behind it, and what I'm assuming, all in the same breath, and I commit openly to a tighter follow-up once there's more time. Silence until it's perfect helps nobody if the stakeholder has to decide by Friday either way.
Structured elaboration
- Find the fastest defensible path, not the most rigorous one. Given a week in an area I don't know well, I look first for existing data or dashboards that are already adjacent to the question, then a short conversation with whoever actually owns that part of the business to get the two or three facts that matter most, before I'd ever try to build something from scratch.
- Make a conservative first cut. Wherever I'm genuinely unsure, I lean toward the more cautious assumption, so if the number is wrong, it's wrong in the direction that's less likely to mislead the decision being made with it.
- Say explicitly what's left out. I tell the stakeholder plainly what the number does and doesn't cover, so they know its boundaries instead of assuming it accounts for everything.
- Put the assumptions right next to the number. Not buried in an appendix nobody reads: if the number depends on three specific assumptions, I say so in the same message the number appears in.
- Give a range, not false precision. A rounded range like "roughly 800 to 1,200" is more honest than a specific-looking figure like "947," because the second implies a level of rigor I don't actually have.
- Commit to and schedule the tightening pass. I say what additional data or time would sharpen the number, and when I'll have it, so the first answer is understood as a starting point rather than the final word.
Worked example
A stakeholder once needed an estimate of how much additional support-ticket volume a new customer segment would generate, before we finalized staffing for the following quarter, and I'd never analyzed that segment before. Rather than going quiet for a week to build a proper model, I spent half a day finding the closest available proxy: an existing segment with roughly similar product usage patterns, and its historical ticket rate per active user. I applied that rate to our projected user count for the new segment, deliberately rounding up the assumption about how "similar" the segments really were, since I wasn't confident and wanted the estimate to err toward not under-staffing. I sent the number as a range, with the two assumptions stated directly underneath it (the proxy segment's comparability, and the projected user count itself), and said I'd have a tighter number within two weeks once we had a few actual weeks of the new segment's real data. That let them staff conservatively now, and the follow-up estimate two weeks later came in close to the original range.
Trade-offs and pitfalls
The clearest pitfall is going quiet while trying to build something more rigorous than the deadline allows, since the stakeholder ends up deciding without you anyway, just with worse information. The opposite pitfall is handing over a specific-looking number without caveats, which invites the stakeholder to trust it further than it deserves and use it in ways it was never meant to support. The middle path, a clearly-labeled range with visible assumptions and a committed follow-up, is what actually respects both the deadline and the limits of what you know.
You observe periodic route oscillations for a set of prefixes approximately every 5 minutes. Explain how you would analyze for causes such as BGP MED oscillation, iBGP route-reflector feedback loops, next-hop unreachability flapping, route flap damping interactions, or upstream policy. Specify the logs, BGP update traces, timers, packet captures, and configuration changes you would examine and tests you would run to identify and mitigate the oscillation.
Sample Answer
Direct answer
A steady 5-minute period means something is driven by a timer or a repeating state machine, not by random failures. I would first prove what is oscillating (reachability, best-path choice, or only attributes), group the affected prefixes by what they share (next hop, neighbor AS, origin AS, community), and then match the period against each candidate timer. The five candidates leave different fingerprints, so a few targeted observations separate them before I change anything.
Vocabulary used below. A route reflector (RR) is an iBGP router (BGP between routers of one autonomous system) that re-advertises routes it learned from its clients, so those clients need not peer with every router; an RR and its clients form a cluster. Two attributes stop reflection loops: ORIGINATOR_ID (the router ID of the router that first injected the route into the AS) and CLUSTER_LIST (the CLUSTER_IDs of the clusters the route has been reflected through, newest first); the CLUSTER_ID is the cluster's name, shared by its redundant RRs. A confederation splits one AS into member ASes with the same effect of avoiding a full mesh. MED (Multi-Exit Discriminator) is a number a neighboring AS attaches to say which of its links it prefers you use, lower wins, and it is only comparable between paths from the same neighbor AS (RFC 4271 section 9.1.2.2). MRAI (MinRouteAdvertisementInterval) is the minimum time a router must leave between two UPDATEs about the same prefix to one peer (RFC 4271 section 9.2.1.1). BMP (BGP Monitoring Protocol, RFC 7854) lets a router stream a copy of the routes and updates it receives to a collector.
Step 1: Pin down the pattern (10 minutes of data)
- Pick two or three affected prefixes. Capture every UPDATE and WITHDRAW for them from a passive source: a BMP feed or a route collector if you have one, otherwise a debug of BGP updates filtered to those prefixes on one router, plus syslog timestamps. Record, per event: time, peer, whether it was a withdrawal or an announcement, and the attributes (AS path, next hop, MED, local preference, communities, ORIGINATOR_ID, CLUSTER_LIST).
- Run a packet capture on one affected session so the capture agrees with the router's own view:
tcpdump -nn -ttt -w bgp.pcap 'tcp port 179 and host 192.0.2.1'(the filter compiles as written; replace the address with the peer). Decode the UPDATEs offline. - Classify each event:
- Withdrawal then re-announcement: the prefix really becomes unreachable (next hop loss, upstream withdrawal, session event).
- Announcement only, attributes changed: the new UPDATE silently replaces the old path (an implicit withdrawal, no WITHDRAW message is sent). This is best-path churn. MED, route reflection and policy cause it.
- Group the prefixes. If they all share one next hop, suspect next-hop reachability. If they all come from two neighbor ASes, suspect MED. If they share an origin AS or a community, suspect upstream policy.
- Write down the period precisely (for example 300 s plus or minus 1 s) and the phase relative to other events (the IGP (interior gateway protocol) log, interface errors, scheduled jobs).
- Review configuration changes: diff the running configuration against the last known-good version and read the change log for the onset time of the first oscillation. Look for new or edited route-maps, MED or local-preference settings, damping, route-reflector cluster IDs, next-hop-self, timers, and IGP metric changes. A flap that began right after a policy change is the cheapest hypothesis to test.
Reading the router side: on a reflected path show ip bgp <prefix> prints a line such as Originator: 192.0.2.1, Cluster list: 2.2.2.2 (illustrative addresses). The Originator is the router that first learned the route inside the AS, and the Cluster list shows which clusters reflected it. If the Originator or Cluster list changes between consecutive updates for the same prefix, the route is arriving through different reflection paths each time, which is the reflection signature.
Step 2: Match the period to a timer
RFC 4271 section 10 suggests these defaults, which are the timers worth comparing against 300 seconds: the hold time 90 seconds with keepalive at one third of it, MinRouteAdvertisementInterval 30 seconds on eBGP and 5 seconds on iBGP, and the connect retry timer 120 seconds. Also compare against the IGP SPF (shortest path first) throttle (the delay a link-state IGP waits before re-running its path calculation after a change), the BGP next-hop tracking delay (BGP's watch on the IGP route to each next hop, so it reacts within seconds when one disappears; Cisco default 5 seconds), route flap damping half-life (default 15 minutes), and anything scheduled on the box (IP SLA, tracked objects, cron-style automation). None of the protocol defaults equals 300 seconds. A whole-session reset resets every prefix from that peer, so a flap limited to "a set of prefixes" points away from keepalive and hold-timer expiry and toward attribute or next-hop causes.
As a first suspicion, not proof: MED oscillation has no clock of its own, its rhythm comes from how quickly updates propagate, so a period that stays at exactly 300 s is more likely a configured timer, a scheduled job or an upstream's own policy. Rule those in or out with the capture first, then move to the route-reflection causes.
Step 3: Test each cause, in this order
| Cause | Fingerprint | Test | Mitigation |
|---|---|---|---|
| Next-hop unreachability | Withdrawals for every prefix sharing a next hop, at the moment the IGP route to that next hop disappears | show ip route <next-hop> repeatedly; correlate with IGP adjacency, LSA or LSP logs and interface counters; check whether the next hop is only reachable over a flapping link | Fix the link or IGP instability; BFD on the underlying link; keep next hops on stable loopbacks. Tuning bgp nexthop trigger delay (range 1 to 100 s, default 5 s) changes how fast BGP reacts, it does not remove the cause |
| MED-induced oscillation | The prefix stays reachable but its best path rotates among three or more paths, with the rotation visible as updates and withdrawals of reflected copies between the RRs; no external change | show ip bgp <prefix> on the route reflectors and a client: are there paths from two or more neighbor ASes with different MEDs, behind a route-reflector or confederation design? | RFC 3345 mitigations: make links between clusters cost more in the IGP than links inside a cluster, so each RR prefers its own cluster's exits; stop accepting MED (overwrite it with a constant on ingress); set local preference per advertising AS so MED is not compared across ASes; or, where the router count allows it, use a full iBGP mesh (RFC 3345 lists it and notes it does not scale) |
| Route-reflector feedback | Same prefix reflected between RRs with changing CLUSTER_LIST; or forwarding loops, not only control-plane churn | Read ORIGINATOR_ID and CLUSTER_LIST in show ip bgp <prefix>; check that redundant RRs in one cluster share a CLUSTER_ID (RFC 4456 section 7) and that RR topology follows the physical topology | Correct cluster IDs; follow the RFC 4456 loop-prevention design; align RR hierarchy with IGP topology |
| Route flap damping | Prefix disappears for tens of minutes after repeated flaps; show ip bgp dampening dampened-paths lists it | show ip bgp dampening flap-statistics, then show ip bgp dampening dampened-paths | Use RFC 7196 values (suppress threshold of at least 6000), or exempt the affected prefixes; damping hides a flap, it is not its source |
| Upstream policy | Announcements arrive with periodically changing attributes (AS-path prepend, MED, community) | Compare consecutive UPDATEs from the capture; ask the upstream for their change window; check whether the change is on their schedule | Escalate with the capture as evidence; filter or normalize the attribute on ingress; if only the MED changes, overwrite it |
How a route reflector plus MED makes the best path rotate
MED is only compared between paths learned from the same neighbor AS, and a route reflector chooses from only the paths it has been sent. Together these let two RRs knock each other's preferred path out in turn. The numbers below are the textbook case from RFC 3345 (Figure 1). Cluster 1 has RR Ra with clients Rb and Rc; cluster 2 has RR Rd with client Re. All three clients learn the same prefix from outside: Rb from AS10 with MED 10, Rc from AS6 with MED 1, Re from AS6 with MED 0. Ra's IGP costs are 5 to Rb, 4 to Rc and 13 to Re; Rd's are 12 to Re, 5 to Rc and 6 to Rb.
- Ra holds the AS6 path via Rc (MED 1, cost 4) and the AS10 path via Rb (cost 5). They come from different neighbor ASes, so MED is not compared and the lower IGP cost wins: Ra picks the AS6 path and advertises it to Rd.
- Rd now holds Re's AS6 path (MED 0, cost 12) and Ra's AS6 path (MED 1, cost 5). Same neighbor AS, so MED decides: MED 0 beats MED 1. Rd picks Re's path and advertises it to Ra.
- Ra now sees an AS6 path with MED 0 (via Rd, cost 13). That path eliminates the MED 1 path from Rc, leaving the AS6 MED 0 path (cost 13) against the AS10 path (cost 5). IGP cost now picks AS10, so Ra advertises the AS10 path and withdraws what it sent about AS6.
- Rd now holds Re's AS6 path (cost 12) and Ra's AS10 path (cost 6). IGP cost picks AS10 and Rd withdraws its AS6 advertisement.
- Ra no longer hears the MED 0 path, so the MED 1 path via Rc is back in play and beats AS10 on IGP cost (4 against 5). That is step 1 again, and the cycle repeats with no external change.
RFC 3345 states the conditions: a single-level reflection or confederation design, and MEDs from two or more neighbor ASes that are unique. Its preferred mitigation is the one in the table: had the inter-cluster IGP metrics been much larger than the intra-cluster ones, step 3 would not occur, because Ra would always prefer exits inside its own cluster.
Worked example: how damping would behave on a 5-minute flap cycle
Route flap damping gives each prefix a penalty score. Every flap adds points, the score decays over time and halves every half-life, and once it passes the suppress threshold the router stops using the route; it uses the route again when the score decays below the reuse threshold. Take the default Cisco damping parameters listed in RFC 7196 Table 1: a withdrawal adds 1000, an attribute change adds 500, suppress threshold 2000, reuse threshold 750, half-life 15 minutes, maximum suppress time 60 minutes. The ceiling is therefore 750 x 2^(60/15) = 12000. With one withdrawal every 5 minutes, the penalty just after each event is 1000, 1794, 2424, 2924, 3321, 3635, 3885, 4084 (computed, each step decays the previous value by 0.5^(5/15) and adds 1000). It crosses 2000 at the third flap (minute 10) and the route is suppressed from then on. The steady-state value approaches 1000 / (1 - 0.5^(5/15)) = about 4847, so once the upstream stops flapping the route stays suppressed for 15 x log2(4847/750) = about 40 minutes before reuse. That is the operational lesson: damping converts a 5-minute flap into a 40-plus-minute outage, so if the symptom stays a clean 5-minute cycle, damping is not suppressing it, and if reachability becomes much worse after the first 10 minutes, damping is amplifying it.
Step 4: Mitigate and prove it
Apply the one fix that matches the fingerprint, then verify for at least two full periods plus margin (a 12-minute window shows two cycles): zero UPDATEs for the affected prefixes on the capture, a stable best path in show ip bgp <prefix>, and no growth in flap statistics. If you cannot prove which cause it is, a safe interim step is to stop the loop from propagating (a stable static or aggregate at the edge), not to widen timers across the whole network.
Pitfalls
- Changing MRAI or damping first can hide the symptom while the root cause remains.
- MED is compared only between routes learned from the same neighboring AS (RFC 4271 section 9.1.2.2), so "always compare" settings change the whole decision process and need a network-wide change window.
- Treating route reflection and MED as two unrelated suspects: RFC 3345 shows the oscillation needs both (a route-reflection or confederation design plus unique MEDs from two or more ASes).
A DNS change went out and about 15% of users are getting errors. You roll it back, yet many clients keep resolving the bad or old answer for hours. Explain why, what you can do to shorten the damage, and what change process would prevent a repeat.
Sample Answer
Direct answer
A rollback changes the authoritative data, but nothing tells the thousands of resolver caches that already hold the bad answer. Each cache keeps it until its own timer runs out, and the timer is the TTL (time-to-live) that was published with the bad record, counted from when that particular cache fetched it, not from when you rolled back. In times: if the bad record carried a TTL of four hours and the last moment anyone could fetch it was the 09:40 rollback, a cache that fetched it at 09:40 keeps it until 13:40. That is why the damage after a rollback lasts up to one full TTL beyond the last moment anyone could still fetch the bad data. To shorten it, make the bad answer harmless (keep the old target serving traffic), flush the caches you or others can flush, and communicate; to prevent a repeat, lower TTLs ahead of risky changes, roll out to a small share first, and watch resolution from outside before and after.
Why stale answers survive a rollback (executed demonstration)
Setup in a Debian container: BIND is the authoritative server for example.lab with a record TTL of 600 s; Unbound is a caching resolver in front of it.
=== step 1: publish app A 192.0.2.10 with TTL 600; resolver asks and caches
app.example.lab. 600 IN A 192.0.2.10
=== step 2: bad change: app moved to 192.0.2.99 (authoritative now serves it)
authoritative says:
app.example.lab. 600 IN A 192.0.2.99
=== step 3: a few seconds later the resolver still answers from its cache, TTL counting down
app.example.lab. 596 IN A 192.0.2.10
=== step 4: the resolver operator flushes the name
ok
app.example.lab. 600 IN A 192.0.2.99
Step 3 is the whole story: the authoritative answer changed, the resolver kept serving its copy with a counting-down TTL, and only an explicit flush (unbound-control flush app.example.lab) removed it. Read the lines this way: the number after the name is the remaining TTL in seconds (600 when fresh, 596 a few seconds later because the resolver's copy is already four seconds old), and the last field is the address, which is still the old 192.0.2.10 while the authoritative server says 192.0.2.99.
For BIND the equivalent is rndc flushname app.example.lab. What success looks like differs between the two tools. unbound-control flush prints ok and exits 0. rndc flushname prints nothing and exits 0 (checked with echo $?). In a test with a forwarding BIND resolver, dig @resolver app.example.lab A returned the stale 192.0.2.10 with a TTL of 595 before the flush and 192.0.2.99 with TTL 600 after it. A flush clears that one name from that one resolver only, so the check that it worked is a fresh dig against the resolver, comparing the address and TTL with what the authoritative server returns.
The arithmetic of "hours"
Say the bad record had TTL 14400 (four hours), was published at 09:00, and rolled back at 09:40. A resolver that fetched it at 09:39 holds it until 13:39; the worst case, a fetch at 09:40, holds it until 13:40. With a TTL of 300 the same mistake would have been gone from every compliant cache by 09:45. (Computed with Python datetime: 09:40 plus 14400 s is 13:40; plus 300 s is 09:45.)
What stretches the tail beyond the TTL
- Resolvers that clamp TTLs. To clamp is to force a value into a limit: some resolvers enforce a minimum TTL and treat anything shorter as that minimum. A second Unbound instance with
cache-min-ttl: 300answeredshort.example.labwith TTL 300 while the authoritative server published 30. A published TTL of 60 behind a resolver with a 3600 s floor lasts until 10:40 in the example above. - Negative caching. If the bad change deleted a record, resolvers cache the NXDOMAIN (name does not exist) answer. RFC 2308 sets its lifetime to the smaller of the SOA record's own TTL and its MINIMUM field, and recommends keeping it to one to three hours, noting that more than a day is problematic. Unbound's default ceiling for negative answers is 3600 s.
- Other layers with their own timers: CNAME targets, NS (delegation) records, and caches in the operating system, browser and application runtime, plus connection pools (sets of already-open connections that an application reuses) that keep using an address they already hold.
- Serve-stale. Resolvers that implement RFC 8767 (serve-stale means answering from a record whose TTL has already run out) may answer from expired data, with a recommended TTL of 30 seconds on such answers, when the authoritative servers are unreachable. That prolongs outages, not bad answers from reachable servers.
How to shorten the damage
- Make the bad answer work. The fastest cure is to keep serving real traffic at the address the cached answer points to: proxy or redirect it to the good stack. This ends user impact without waiting for any cache.
- Flush what can be flushed: your own resolvers (commands above) and the large public ones. Google Public DNS provides a Flush Cache tool, and Cloudflare's 1.1.1.1 has a Cache Purge tool taking a domain name and record type. Confirm with
dig SOAagainst the resolver and an authoritative server that they agree. - Do not change the record again to "fix" the fix: each change creates a new cohort of caches (a group of caches that fetched the record in the same window, so they expire together) with different expiry times and makes the picture harder to read. Worked example with a TTL of 3600 s: the bad record goes live at 09:00, you roll back at 09:30, and then change it a third time at 09:50. Caches that fetched the bad record between 09:00 and 09:30 hold it until 10:00 to 10:30. Caches that fetched the rollback value between 09:30 and 09:50 hold it until 10:30 to 10:50. Caches that fetch after 09:50 hold the third value. So up to three different answers are live from 09:50 to 10:30, and the last cache with a stale value lets go at 10:50. Had you left the rollback alone, the last bad entry would have expired at 10:30 (computed in Python: 09:30 plus 3600 s is 10:30, 09:50 plus 3600 s is 10:50).
- Communicate the expected end time from the arithmetic above (last bad answer plus TTL), not a guess.
A change process that prevents the repeat
- Lower the TTL first for records you are about to change, and wait out the old TTL (the cutover method in a planned migration).
- Stage it: weighted or percentage-based DNS routing (the DNS service hands out the new address for only a set share of queries and the old one for the rest) for 1 to 5 percent of traffic first, so a bad answer reaches a small fraction and rolling back has a small tail.
- Lint and diff before publishing: zone-file checks (
named-checkzone) and a reviewed diff of what will change; treat warnings as failures where they matter. - Verify from outside: automated probes that query the authoritative servers and a few public resolvers right after publish and compare to expectations, with an automatic revert decision.
- Make rollback a first-class step: the previous record set is stored, the rollback is rehearsed, and the old target stays up for at least one old TTL after any change.
Pitfalls
- "We rolled back" is not the same as "users are fixed". The metric to watch is the share of successful resolutions seen from outside, over time.
- Resolver TTL clamps and application caches make an exact end time unknowable, so say "by about" and keep measuring.
Implement a simple reliable stop-and-wait protocol over UDP in Python: a send_reliable(sock, dest, payload, timeout) and a matching receive_reliable(sock). Use a single-bit sequence number, ACK packets, retransmit-on-timeout, in-order delivery, and duplicate handling. Explain what your implementation demonstrates about which parts of TCP's reliability UDP does not give you for free.
Sample Answer
Direct answer
Building reliability on top of UDP means implementing, by hand, the exact machinery TCP gives you for free: a sequence number to detect duplicates, explicit acknowledgments, and a retransmission timer. A single-bit (0/1) sequence number is enough for stop-and-wait specifically, because only one message is ever in flight at a time.
Structured elaboration (approach)
send_reliable sends the payload tagged with the current sequence bit, then blocks (with a timeout) waiting for a matching ACK; on a timeout it just resends the same packet, and on receiving an ACK for the WRONG sequence number (a stale ACK from a previous round) it keeps waiting rather than treating that as success. receive_reliable accepts a packet, immediately ACKs it (even if it's a duplicate, in case its own previous ACK was lost), and only hands NEW data (matching the expected next sequence bit) up to the caller, silently absorbing duplicates.
Worked example (code)
import socket, struct
HEADER = struct.Struct("!BB") # (sequence bit, type: 0=DATA, 1=ACK)
def send_reliable(sock, dest, payload, timeout, max_retries=5, state={"seq": 0}):
seq = state["seq"]
packet = HEADER.pack(seq, 0) + payload
sock.settimeout(timeout)
for _ in range(max_retries):
sock.sendto(packet, dest)
try:
data, addr = sock.recvfrom(4096)
except socket.timeout:
continue # retransmit on timeout
if len(data) < HEADER.size:
continue
ack_seq, ack_type = HEADER.unpack(data[:HEADER.size])
if ack_type == 1 and ack_seq == seq:
state["seq"] = 1 - seq
return True
return False
def receive_reliable(sock, state={"expected_seq": 0}):
while True:
data, addr = sock.recvfrom(4096)
if len(data) < HEADER.size:
continue
seq, pkt_type = HEADER.unpack(data[:HEADER.size])
if pkt_type != 0:
continue
payload = data[HEADER.size:]
sock.sendto(HEADER.pack(seq, 1), addr) # always ACK, even duplicates
if seq == state["expected_seq"]:
state["expected_seq"] = 1 - seq
return payload, addr
# else: duplicate, already ACKed above, loop for the real next message
This was executed against a deterministic loss-simulating wrapper (a UDP socket wrapper that drops a configurable fraction of outgoing packets using a seeded random generator, so the test is reproducible) sending 4 messages at loss rates of 0%, 30%, and 60%. At every loss rate tested, all 4 messages were delivered exactly once, in the original order, confirming both the retransmit-on-timeout path and the duplicate-suppression path work correctly under real, repeated loss.
Trade-offs & pitfalls (edge cases and complexity)
Complexity: with a single sequence bit and no pipelining, stop-and-wait can send only ONE unacknowledged message at a time, so throughput is bounded by one round trip per message (a real reliability layer would need a sliding window of sequence numbers, not just one bit, to use a high-bandwidth-delay-product link efficiently, exactly the same motivation as TCP's own window). Edge cases handled: a lost DATA packet (sender times out, retransmits), a lost ACK (receiver gets a duplicate DATA packet, re-ACKs it without re-delivering the payload to the application), and a delayed ACK arriving after the sender has already given up and retransmitted (the sender must ignore an ACK for the WRONG sequence number rather than treating it as confirmation, otherwise a stale ACK could be mistaken for acknowledging the NEXT message). What this exercise demonstrates: TCP is doing exactly this kind of bookkeeping (and much more, for a full sliding window, congestion control, and out-of-order buffering) on every connection, for free.
Tell me about a time you had to explain a complex incident to a non-technical team, for example legal, sales, or executives. What did you choose to include, what did you leave out, and what was the outcome with those stakeholders?
Sample Answer
Direct answer
The core move in an incident explanation to a non-technical audience is separating three layers up front: what happened (in plain terms, no root-cause mechanism), what it meant for them (impact, in terms they already track), and what's being done about it, then deliberately leaving out anything that doesn't serve one of those three. Below is an incident where I did that under time pressure, including delivering it live to a mixed engineering-and-business audience.
What to include, what to leave out, and how to decide
- Lead with impact, not sequence. Legal, sales, and executives care about what happened TO THEM first, which customers, how long, what's the exposure, the technical timeline is useful evidence, not the headline.
- Deliberately exclude logs, stack traces, and internal service names; they add authority for an engineering audience and add nothing but confusion for this one. A useful test: if a detail doesn't change what the listener should do next, leave it out.
- Give the cause in one plain sentence with no jargon, something like "a recent configuration change made one of our systems too slow to respond to a partner service in time," rather than either omitting cause entirely (which reads as evasive) or over-explaining the mechanism.
- When delivering this live rather than in a written report, whether it's a hallway update or presenting a postmortem verbally to a room that mixes engineers and business stakeholders, pause after the impact statement for questions before moving to cause. People worried about impact can't absorb a root-cause explanation until that worry is addressed first.
Worked example
Situation: during a high-traffic sales period, our payment service began intermittently failing checkout requests for roughly ninety minutes. Legal, sales leadership, and the executive team needed an explanation quickly.
Task: explain what happened clearly enough for them to act, communicate with affected customers, assess any obligations, decide on immediate next steps, without either alarming them with irrelevant detail or minimizing the impact.
Action: I opened with impact, in the terms they track: which customers were affected, for roughly how long, and that the issue was fully resolved and being watched closely. I gave the cause in one sentence: a recent configuration change made our payment service too slow to respond to our external payment gateway in time, causing some checkout attempts to fail. I described what we did in plain terms (reverted the change, increased how long we wait before giving up on a slow response, added an automatic circuit breaker so a slow dependency can't cascade into a wider outage) and what we were doing next (a deeper review, with a fuller technical writeup available to anyone who wanted it). I left out the specific error codes, service names, and configuration parameter, none of which changed what legal, sales, or the executives needed to do next. I paused for questions right after the impact statement, before moving on, and answered a legal question about customer notification obligations directly instead of routing it back to engineering jargon.
Result: legal and sales left with a clear, accurate picture of exposure and could communicate confidently with affected customers; the executive team approved the follow-up work (the circuit breaker and review) without needing to dig into implementation detail themselves, and a fuller technical postmortem was made available separately for the engineering team that wanted the mechanism-level explanation. I learned that pausing for questions right after the impact statement, before cause, kept people from tuning out a cause explanation they weren't ready to hear yet.
Trade-offs and pitfalls
Leaving out technical detail can read as evasive if you do it silently; I said "I'm not going to walk through the technical internals here, I'm glad to share those separately" so the omission was visible on purpose rather than hidden. The other pitfall is understating severity to keep the room calm, that erodes trust the moment the real scope becomes clear later. State the honest impact even when it's uncomfortable, and let the "what we're doing about it" section carry the reassurance instead of the impact statement itself.
You are building the access layer for a new data centre pod and the team's default is spanning tree with redundant uplinks. Where does spanning tree hurt you at that scale, what would you put in its place so every uplink carries traffic, and in what situation would you still keep it?
Sample Answer
Direct answer
Spanning tree is the protocol that stops loops in a switched network by electing one switch as the root bridge (the reference point, the one with the lowest bridge ID: the priority number first, with the MAC address as the tie-break) and blocking every redundant port so each switch keeps a single forwarding path toward the root. A pod is a group of racks built and managed as one unit, and a ToR (top-of-rack) switch is the switch at the top of each rack that the servers plug into. Spanning tree hurts at pod scale because it makes redundancy cost bandwidth: it blocks every redundant link so each switch has one forwarding path to the root bridge, which means half of dual-homed uplink capacity sits idle, and a topology change can disturb forwarding while it reconverges. For the access layer I would put a multi-chassis link aggregation (MLAG, called vPC, virtual PortChannel, on Cisco Nexus) pair at the top of each rack, with servers dual-homed using LACP (Link Aggregation Control Protocol, which bundles several cables into one logical link) and ToR uplinks routed (each uplink is a layer 3 point-to-point link with its own IP address, and the routing protocol spreads traffic over all equal-cost uplinks, so there is no layer 2 loop to block) or aggregated into a port-channel, so every link forwards and failure is handled by link aggregation, not by a spanning-tree recalculation. I would still keep spanning tree running as a safety net on edge ports to stop a mistaken cable from causing a loop, and keep it as the main mechanism where a legacy switch or a device without LACP support must attach with redundant links.
Where spanning tree hurts (computed example)
Spanning tree (Cisco defaults to Rapid PVST+ on Catalyst 9000, with bridge priority 32768 and port priority 128) forces redundant paths into a blocked state. Example: a pod where each of 20 access switches has two 40G uplinks to two aggregation switches. Under a single spanning tree shared by all VLANs (or per-VLAN trees that all have the same root), one uplink per switch forwards: usable = 20 x 40G = 800G of 1,600G installed, 50%. Per-VLAN trees (PVST+, one tree per VLAN) or multiple spanning tree (MST, one tree per group of VLANs) can alternate the blocked link per VLAN. Concretely, with 20 VLANs, make aggregation switch 1 the root for VLANs 1 to 10 and aggregation switch 2 the root for VLANs 11 to 20: on each access switch the uplink to switch 1 forwards VLANs 1 to 10 and blocks 11 to 20, and the uplink to switch 2 does the opposite. Each 40G uplink carries 10 VLANs, so both are used and, if the VLANs carry equal load, the pod uses its full 1,600G on average. It recovers capacity on average, but any one VLAN (and so any single large flow in it) still has only one 40G path, and the load balance is a manual VLAN-to-root assignment that drifts as workloads move.
Other costs at scale:
- Blocked links are not tested by traffic until failure, so a bad standby link surfaces during an outage.
- Root election is fragile: a switch added with a lower bridge priority value than the intended root wins the election and reshapes the whole tree.
- Large single broadcast domains and many VLANs on shared trees amplify any mistake into a pod-wide event.
What to use instead, at layer 2
| Need | Mechanism | How it removes the block |
|---|---|---|
| Server dual-homed to two switches | MLAG pair + LACP port-channel | The two ToRs look like one switch, so both server links forward |
| ToR to aggregation, every uplink used | MLAG on the uplinks (port-channel across the aggregation pair) | Spanning tree sees one logical link |
| Servers attached to several leaf switches with no peer link between them | EVPN multihoming (Ethernet VPN, an overlay control plane) | The leaves advertise a shared server attachment through BGP (the routing protocol that carries reachability information between devices), so no switch pair is coupled by a peer link |
The MLAG design needs a peer link, consistent VLAN and spanning-tree configuration on both peers, a heartbeat path to avoid split-brain (the two peers lose contact, each assumes the other is dead and both act as the active unit), and a staged upgrade procedure; the Arista reference documents show mlag config-sanity and states that the global spanning-tree configuration is taken from the primary peer. |
When I would still keep spanning tree
- Always on edge ports as a protection: PortFast (moves the port straight to forwarding) plus BPDU guard (error-disables the port if it receives spanning-tree messages), configured with
spanning-tree portfastandspanning-tree bpduguard enableon Cisco, or globally withspanning-tree portfast default(PortFast on all non-trunking access ports) plusspanning-tree portfast bpduguard default(BPDU guard on PortFast ports; the second command alone does not turn PortFast on). A server NIC bridge or a stray switch then cannot create a loop. - As the main protocol where it is the only safe option: attaching a legacy switch with redundant links to the pod, a small branch or lab pod where two cables are the entire redundancy need, or a device that cannot run LACP.
- Root guard (
spanning-tree guard root) and a deliberately set root where access switches connect to a shared layer.
Pitfalls
- Disabling spanning tree entirely "because MLAG is loop-free": a miscabled port-channel or a bridged server then floods the pod.
- Treating MLAG as free: it couples two switches' software versions and makes the peer link critical.
- Mixing MST and PVST+ domains without checking the root placement.
- Believing the hashing balances evenly: the hash (a fixed recipe that turns address bits into a link number) sends every frame of one flow down the same link so frames stay in order, which is why a few large flows can still fill one member link.
Design a DMZ architecture for hosting public web applications across a three-tier model (load balancers, web/application tier, database). Describe network segmentation/subnets, firewall rules between tiers, NAT/load-balancer placement, health-check requirements, and ways to minimize attack surface while supporting HTTPS and scaling to 10k RPS across two availability zones with high availability.
Sample Answer
Direct answer
Front the design with load balancers spread across two availability zones (AZs) for both scale and failover, put the web and application tier in a subnet per AZ reachable only through the load balancer, keep the database tier in a private subnet per AZ reachable only from the application tier, and size each tier so that losing one entire AZ still leaves enough capacity in the surviving AZ to serve the 10,000-requests-per-second (RPS) target.
Structured elaboration
Tiers and subnets, duplicated per AZ.
- Public subnet: load balancers only, terminating HTTPS/TLS (Transport Layer Security), health-checking the tier behind them.
- Demilitarized zone (DMZ) subnet: web and application servers, reachable from the load balancer on the application port only, never directly reachable from the internet.
- Private subnet: database tier, reachable only from the DMZ subnet's application servers on the database's own port.
flowchart TB
subgraph AZ1[Availability Zone 1]
LB1[Load balancer] --> WEB1[Web app tier]
WEB1 --> DB1[Database tier]
end
subgraph AZ2[Availability Zone 2]
LB2[Load balancer] --> WEB2[Web app tier]
WEB2 --> DB2[Database tier]
end
INET[Internet] --> LB1
INET --> LB2
Firewall rules between tiers. Load balancer to web/application tier: allow only the application's listening port. Web/application tier to database tier: allow only the database's own port. An explicit deny applies to everything else at each boundary, identically in both AZs, so failover into the surviving AZ does not quietly relax the security posture there.
NAT/load-balancer placement. The load balancer is the only internet-facing component; application servers have no public address at all and are reachable only through the load balancer's private-subnet path. Outbound internet access the application or database tiers might still need, for patching or package repositories, goes through a NAT gateway per AZ rather than a public address on those tiers directly, keeping inbound exposure (nothing but the load balancer) and outbound exposure (routed and still filterable) deliberately asymmetric.
Health checks. The load balancer's health check should exercise a real application endpoint, not just confirm the port is open, and health checks should run per AZ so the load balancer can route traffic away from an unhealthy AZ's instances without waiting for an explicit failover event.
Minimizing attack surface while meeting the scale target. Transport Layer Security terminates at the load balancer, optionally re-encrypting to the backend for defense in depth, so the DMZ subnet and the internet only ever exchange encrypted traffic and certificate management lives in one place rather than scattered across every instance. Only the load balancer ever has a public listener; the web/application and database tiers have zero public addresses; and administrative access to any tier goes through a separate bastion path, not through the same subnets serving customer traffic.
Worked example
Illustrative capacity math, explicitly an assumed input rather than a measured benchmark: suppose each web/application instance can sustain roughly 500 requests per second before its own response-time service-level objective (SLO) degrades. Serving 10,000 RPS from a single AZ then needs about 10,000 divided by 500, which is 20 instances. For the design to survive losing one AZ entirely and still serve the full 10,000 RPS from the surviving AZ alone, that AZ's own capacity needs to already be sized (or able to scale quickly) to roughly 20 instances on its own, meaning close to 40 instances total across both AZs if the organization wants that headroom running in steady state rather than scaling reactively during the failure. That capacity-versus-cost decision has to be made explicit rather than assumed away by simply saying "two AZs."
Trade-offs and pitfalls
Sizing each AZ for only half of the 10,000 RPS target, with no failover headroom, means an AZ failure causes a real capacity shortfall exactly when it can least be afforded, so having two AZs alone does not guarantee availability unless capacity is planned for the failure case specifically. Giving application servers public addresses "just for easier administrative access" defeats the entire point of the DMZ design and is a common shortcut under deadline pressure. Forgetting that the load balancer itself needs to be deployed redundantly across both AZs leaves a single point of failure even when everything behind it is properly duplicated. A health check that only confirms a port is open, rather than real application health, will happily keep routing traffic to a broken instance.
You need to process a set of packet captures (or a live traffic feed) to answer a concrete question, e.g. which flows are retransmitting the most, or what the handshake latency looks like per flow, without opening each one by hand. Describe how you'd script this (naming the library or tool you'd reach for), what fields you'd extract, and what you'd have to be careful of, like out-of-order packets or multiple capture points, for the numbers to be trustworthy.
Sample Answer
Direct answer
For a concrete question like "which flows are retransmitting the most" or "what's the per-flow handshake latency," script it with a pcap-parsing library (Python's scapy, or shelling out to tshark's field-extraction mode) rather than opening captures by hand; the key engineering concerns are correctness (handling out-of-order and duplicate packets) and scale (not loading an entire large capture into memory at once).
Structured elaboration
- Choose the extraction path:
tshark -r file.pcap -T fields -e ip.src -e ip.dst -e frame.lenstreams field values without building Python objects for every packet, and scales best for very large captures; scapy'sPcapReader(as opposed tordpcap) reads packets one at a time, which is the right choice when you need custom logic per packet in Python. - Aggregate incrementally: keep running totals in a dictionary keyed by the flow tuple (source IP, destination IP, and for TCP also ports), rather than storing every packet, so memory use stays flat regardless of capture size.
- Handle multiple capture points and out-of-order packets explicitly: if flows are captured at more than one point in the path, do not assume packet order in the file matches wall-clock order; sort by timestamp within each flow before computing anything that depends on ordering (like handshake latency, which needs to match a SYN to its SYN-ACK, not just the Nth and N+1th packet).
- Validate against a known-answer capture before trusting the script's output on real production data.
Worked example
Here is a minimal top-talkers-by-bytes script, executed against a synthetic 60-packet capture (3 flows) built with scapy for this validation:
from collections import defaultdict
from scapy.all import PcapReader, IP
def top_talkers(path, n=10):
counts = defaultdict(int)
byte_totals = defaultdict(int)
with PcapReader(path) as reader:
for pkt in reader:
if IP in pkt:
key = (pkt[IP].src, pkt[IP].dst)
counts[key] += 1
byte_totals[key] += len(pkt)
return sorted(byte_totals.items(), key=lambda kv: kv[1], reverse=True)[:n]
Run against the test capture, this correctly reported two flows: 10.1.1.10 -> 10.2.2.20 with 40 packets totaling 28,700 bytes, and 10.1.1.11 -> 10.2.2.21 with 20 packets totaling 11,700 bytes, matching the known composition of the synthetic file exactly. For handshake-latency specifically, the same pattern applies but keyed by the full 4-tuple, matching each SYN to the SYN-ACK sharing its source/destination ports and computing the timestamp delta.
Trade-offs & pitfalls
Loading an entire multi-gigabyte capture with rdpcap (which reads the whole file into memory as a list) is a common and expensive mistake; use a streaming reader instead. When captures come from multiple points, do not assume a packet appearing in one capture but not another means it was dropped; capture loss at the tap/collection point itself is a real, separate failure mode from network packet loss, and conflating the two produces wrong conclusions.
Before interviewing for a role like this, walk me through the research plan you'd run. What sources would you consult (for example, the company's engineering or product blog, LinkedIn, GitHub, public filings, or recent news), what facts or signals you'd try to extract from each, and what one or two red flags versus positive signals would most change how you'd approach the role?
Sample Answer
Direct answer
A strong pre-interview research plan works from the most authoritative sources outward: company-controlled material first (site, blog, filings), then people-and-code sources (LinkedIn, GitHub), then outside voices (news, reviews), synthesized into a short brief with the one or two findings that would most change your approach flagged as open questions to raise or confirm.
Structured elaboration
What to pull from each source:
- Engineering or product blog: what they're building and what they've recently launched or pivoted away from. Tells you current priorities, not just stated ones.
- LinkedIn: team size and growth trend, who the hiring manager's other reports are, recent hires or departures (a churn signal), and the seniority mix of recent hires.
- GitHub (if they have public repos): commit cadence, open-issue volume and age, contributor count. This is a proxy for engineering health and how much attention a given system gets.
- Public filings (a 10-K or 10-Q, the annual and quarterly financial reports public companies must file with regulators) or funding announcements for a startup: which business segment is growing, headcount trends, and risks the company names about itself.
- Recent news: layoffs, funding rounds, leadership changes, or product launches, which tell you whether the team is likely riding a tailwind or working through a headwind.
Most of what you find is context. Two findings are worth naming explicitly as the ones that would change how you approach the role, because they change what you do, not just what you know:
- The red flag that matters most: departures concentrated in the team you'd join. Not company-wide churn, which usually reflects the market, but several people at your level leaving that one team within a couple of quarters, especially alongside a posting that has been re-listed. That pattern points at the team rather than the industry. It changes your approach from selling yourself to diligence: you'd want the hiring manager's own account of what changed and who left, and you'd weight what individual engineers say about it far above what the recruiter does.
- The positive signal that matters most: recent, specific, public evidence the team ships. A dated engineering post describing a system they actually built and what went wrong, a changelog with real entries rather than "bug fixes," a conference talk by someone still there. Claiming a culture of ownership is free; a public record of shipping is expensive to fake. This one raises how much scope you'd be willing to take on at the stated level, rather than negotiating for a safer, narrower remit.
Two things also matter more than any single source:
- When sources conflict, don't average them. If LinkedIn shows fast headcount growth but reviews mention understaffing complaints, hold both as live hypotheses and pick the one interview question that would resolve the conflict, rather than quietly picking whichever story you like better.
- Synthesize into something short you'd actually use. A one-page brief (mission in one line, likely challenges ranked, two or three open questions) beats a folder of tabs you never revisit.
Worked example
Say you're interviewing for a Backend Developer role at a mid-size fintech company. The blog's last three posts are about migrating a monolith to services. LinkedIn shows engineering headcount growing noticeably over the past year, and most recent hires are titled "Senior," none "Staff." Their one public GitHub repo, a customer-facing SDK (software development kit, a packaged set of tools other companies use to integrate with a product), has a large number of open issues, the oldest well over a year old. News shows a funding round closed several months ago.
Reading this together: the funding round is the likely reason the headcount line moved, which makes this growth a funded plan rather than a churn backfill, and senior-heavy hiring during a service migration suggests they need people who can operate with less hand-holding. Those two together are the positive signal, and they're worth saying out loud because they tell you the role is probably real scope rather than a replacement seat. The stale open-source issues on a customer-facing SDK is the actual red flag worth raising, not as a criticism, but as a genuine question: is that repo neglected, or intentionally deprioritized while the team focuses elsewhere? Note it isn't a strong enough flag to change whether you'd take the role, only what you'd ask.
Trade-offs and pitfalls
The most common mistake is treating any single source as ground truth, LinkedIn headcount counts are noisy (they include inactive profiles and contractors), so weight repeated, specific signals over one-off numbers. It also helps to say your assumptions out loud early in the process, whether to the recruiter or the hiring manager, rather than presenting inference as settled fact, since that's what catches a wrong read before it steers your whole set of questions in the wrong direction. A related failure is collecting a fact and then never using it, if the funding round or the layoff never shows up in your reading of the team, it wasn't research, it was browsing. Finally, don't let the research become the performance itself, its job is to sharpen two or three real questions, not to be recited back verbatim.
Recommended Additional Resources
- Cisco CCNA Study Guide (100-101) by Todd Lammle - comprehensive entry-level certification preparation
- CompTIA Network+ Study Guide by Mike Meyers - foundational networking knowledge and entry-level certification
- The Illustrated Network: How TCP/IP Works in a Modern Network by Walter Goralski - deep conceptual understanding of networking
- GNS3 Emulator (https://www.gns3.com/) - free network simulation tool for hands-on lab practice
- Cisco Packet Tracer (https://www.netacad.com/) - Cisco's network simulation tool for practice
- Wireshark (https://www.wireshark.org/) - free packet analysis tool for understanding network traffic
- LinkedIn Learning: Network Administration and Cisco certification courses
- Udemy: CompTIA Network+ and CCNA courses by instructor Jason Dion or Chris Bryant
- YouTube: Professor Messer's CompTIA Network+ video series - free comprehensive preparation
- FAANG Technical Interview Preparation: 'Cracking the Coding Interview' by Gayle Laakmann McDowell - for understanding problem-solving methodology (applicable to troubleshooting mindset)
- LeetCode - while primarily for coding, useful for understanding systematic problem-solving approaches
- System Design Primer (GitHub) - foundational system thinking even for entry-level network roles
- Official Cisco IOS Documentation and Command References - reference material for hands-on practice
Search Results
48 Networking Engineering Interview Questions (With Answers)
Tell me about yourself. · Why did you decide to become a network engineer? · How did you hear about the organisation? · What specifically about this role appeals ...
CCNA Certification: Top 60 Interview Questions and Answers - Jetking
21. What is the purpose of the ping command? Ping checks the connectivity between two network devices. 22. What is a MAC address?
▷ Cybersecurity Interview Questions and Answers (2025 Guide)
21. What is encryption, encoding and hashing? 22. What is Perfect Forward Secrecy? 23. What is WEP crack? 24. What is meant by network sniffing? 25. What do you ...
Top 73 CCNA Interview Questions and Answers
1) What is a MAC Address? Answer: A Media Access Control (MAC) address is a unique identifier assigned to a Network Interface Card (NIC) by the manufacturer. It ...
This interview preparation guide was generated using AI-powered research from the sources listed above. While we strive for accuracy, we recommend verifying critical information from official company sources.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Network Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs