Google Network Engineer (Entry Level) Interview Preparation Guide
Google's network engineer interview process for entry-level candidates typically consists of a recruiter screening phase, followed by 1-2 technical phone screens focusing on networking fundamentals and troubleshooting, and 4-5 onsite rounds covering technical networking depth, practical troubleshooting scenarios, network design basics, and behavioral/culture fit assessment. The process evaluates foundational networking knowledge, problem-solving ability, communication skills, and alignment with Google's collaborative culture.
Interview Rounds
Recruiter Screening
What to Expect
Initial conversation with a Google recruiter to assess your background, verify your interest in the network engineer role, and determine baseline fit. This includes discussion of your networking experience, education, motivation for joining Google, and clarification of role expectations. The recruiter will also explain the interview process and timeline.
Tips & Advice
Be enthusiastic and clear about your interest in network engineering. Highlight any internships, labs, certifications (CCNA, Network+), or personal projects involving network configuration. Prepare a concise summary of your networking experience and why you want to work at Google on infrastructure. Ask thoughtful questions about the role and team to show genuine interest.
Focus Topics
Questions to Ask About the Role
Prepare thoughtful questions about the network engineer position, the team structure, onboarding process, and typical projects you'd work on.
Practice Interview
Study Questions
Communication and Professionalism
Communicate clearly, answer questions directly, and maintain professional demeanor. Demonstrate listening skills and ask relevant follow-up questions.
Practice Interview
Study Questions
Background and Motivation
Articulate your networking experience, educational background, and genuine interest in Google's infrastructure and the specific network engineer role.
Practice Interview
Study Questions
Technical Phone Screen - Networking Fundamentals
What to Expect
First technical interview conducted over phone or video with a Google network engineer. This round focuses on foundational networking concepts, TCP/IP stack, addressing, routing basics, and your understanding of core protocols. You may be asked to explain networking concepts, work through simple network scenarios, or discuss how you've applied these concepts in past work or learning. This assesses your baseline technical knowledge and communication ability.
Tips & Advice
Review OSI model layers, TCP/IP stack, IPv4/IPv6 addressing, subnetting, routing concepts (static vs dynamic), and common protocols (TCP, UDP, DNS, DHCP, ARP). Be ready to explain network concepts clearly and ask clarifying questions if a scenario is unclear. Use a whiteboard or notepad to sketch network diagrams while explaining. For entry-level, focus on explaining fundamental concepts correctly rather than advanced optimization. Reference the job description: be ready to discuss how you understand network infrastructure design, protocols, and basic troubleshooting.
Focus Topics
Network Design Concepts
Understand basic network architecture concepts like LAN, WAN, VLANs, network segmentation, redundancy, and how these relate to reliability and scalability.
Practice Interview
Study Questions
Routing Fundamentals
Understand static routing, dynamic routing basics (distance-vector vs link-state), routing tables, default gateways, and how routers forward packets. Know concepts like convergence and metric.
Practice Interview
Study Questions
Basic Network Troubleshooting Tools
Know common Linux/Unix networking tools like ping, traceroute, netstat, ss, ifconfig/ip, dig, nslookup, arp, and tcpdump. Understand what each tool shows and when to use it.
Practice Interview
Study Questions
IPv4 Addressing and Subnetting
Master IPv4 address classes, subnet masks, CIDR notation, calculating network/broadcast addresses, and determining host ranges. Understand private vs public IP ranges and special-use addresses.[1]
Practice Interview
Study Questions
OSI Model and TCP/IP Stack
Understand the 7 layers of the OSI model and the TCP/IP model, including the protocols at each layer and how data flows through layers. Know common protocols like TCP, UDP, IP, ICMP, DNS, DHCP, ARP, and HTTP/HTTPS.
Practice Interview
Study Questions
Technical Phone Screen - Network Troubleshooting and Problem-Solving
What to Expect
Second technical phone interview focusing on practical troubleshooting scenarios and problem-solving methodology. You'll be presented with network connectivity or performance issues and asked to diagnose the problem methodically. This assesses your troubleshooting approach, logical thinking, ability to ask clarifying questions, and understanding of common network failure modes. The goal is to see how you apply foundational knowledge to real-world scenarios, not necessarily to arrive at the 'correct' answer quickly.
Tips & Advice
When presented with a troubleshooting scenario, start by gathering information: ask what symptoms are being observed, what was working before, what changed. Use a systematic approach—work from Layer 1 upward or follow the network path. Ask clarifying questions about the environment (VLANs, firewalls, NAT, cloud vs on-prem). Explain your reasoning aloud so the interviewer understands your methodology. Don't jump to conclusions; rule out possibilities methodically. Reference the search results[1] which show common troubleshooting patterns: IP connectivity issues, service port accessibility, DNS resolution failures, routing to specific subnets, inter-VLAN communication, and application-level issues.
Focus Topics
VLAN and Layer-2 Issues
Understand VLAN basics, inter-VLAN routing configuration, trunk vs access ports, VLAN tagging, and how to troubleshoot communication failures between VLANs. Know when to check port assignments and router interfaces.[1]
Practice Interview
Study Questions
Service and Application Connectivity Issues
Differentiate between network-level problems and application-level issues. Know how to verify service listening (ss -lntp), check port binding, validate load balancer configuration, and understand when to involve application teams.[1]
Practice Interview
Study Questions
DNS Troubleshooting
Understand DNS resolution flow, know how to test DNS with dig and getent commands, differentiate between DNS failures and other connectivity issues, identify resolver configuration issues via /etc/nsswitch.conf, and diagnose when DNS works but applications fail.[1]
Practice Interview
Study Questions
Routing and Path Issues
Diagnose missing routes, overlapping CIDRs, policy-based routing issues, and understand how to use 'ip route' command to view and troubleshoot routing. Identify when selective routing failures occur.[1]
Practice Interview
Study Questions
Connectivity Troubleshooting Methodology
Systematic approach to diagnosing connectivity issues: verify local configuration (IP, gateway), check layer-2 connectivity (ARP, MAC), test layer-3 routing, verify service listening on correct port, and rule out firewall/NAT issues. Know when to use ping, traceroute, arp, netstat, and ss.[1]
Practice Interview
Study Questions
Onsite Interview 1 - Network Infrastructure Design and Protocols
What to Expect
In-person or video interview with a senior network engineer or infrastructure architect. This round assesses deeper understanding of network infrastructure design, protocol choices, and architectural thinking at a basic level. You'll discuss designing simple network topologies, selecting appropriate protocols and technologies for scenarios, understanding performance and reliability trade-offs, and explaining your design reasoning. This evaluates whether you understand how individual network components work together and can think about infrastructure holistically, not just individual device configuration.
Tips & Advice
Prepare by understanding network design trade-offs: redundancy vs cost, performance vs complexity, security vs usability. Be ready to discuss when to use different technologies: static vs dynamic routing, VLAN vs subnet segmentation, L3 switching vs separate routers. Draw network diagrams and explain your reasoning. For entry-level, focus on explaining fundamental design decisions (e.g., why use OSPF over static routing) rather than complex optimization. Discuss the job responsibilities: how would you design network infrastructure for reliability, security, and scalability? Reference Google's emphasis on infrastructure reliability. Ask clarifying questions about requirements before proposing designs.
Focus Topics
IP Addressing Strategy and IPAM
Design addressing schemes using subnetting and hierarchical addressing. Understand IPAM (IP Address Management) concepts, address space planning, and how addressing scales with network growth.
Practice Interview
Study Questions
Routing Protocol Selection
Understand characteristics of different routing protocols: static routing, RIP, OSPF, BGP. Know when to use each based on network size, complexity, and convergence requirements. Discuss metrics, cost, and scalability.
Practice Interview
Study Questions
Switching and VLAN Architecture
Understand switched network design, VLAN architecture for segmentation, trunk port design, spanning tree basics for preventing loops, and Layer 3 switching. Know how switches provide scalability and segmentation.
Practice Interview
Study Questions
Network Security Fundamentals in Design
Understand how firewalls fit into network design, DMZ concepts, network segmentation for security, ACL basics, and how security considerations influence architecture. Know the job requirement to implement network security measures.
Practice Interview
Study Questions
Network Topology Design Principles
Understand basic network topologies (star, mesh, hybrid), their advantages and disadvantages. Know when to use redundancy, failover mechanisms, and how topology affects reliability and scalability.
Practice Interview
Study Questions
Onsite Interview 2 - Advanced Troubleshooting and Performance Monitoring
What to Expect
Interview with a network operations engineer or performance specialist. This round combines challenging troubleshooting scenarios with discussion of network monitoring, performance analysis, and operational metrics. You'll work through complex multi-layer issues that require synthesizing knowledge across protocols, tools, and infrastructure. The interview also assesses your understanding of network performance monitoring requirements and how to measure network health. This evaluates problem-solving depth, ability to handle ambiguity, and understanding of operational requirements.
Tips & Advice
Prepare for complex scenarios combining multiple failure modes (e.g., intermittent connectivity that turns out to be MTU issues, or DNS working for some clients but not others). Approach systematically but be willing to pivot when new information changes your hypothesis. Discuss monitoring: what metrics matter (latency, jitter, packet loss, throughput), what tools to use (NetFlow, sFlow, SNMP, packet analysis), and how to detect problems proactively. Reference the job responsibility for monitoring network performance and traffic. Be comfortable discussing tools like Wireshark for packet analysis, NetFlow for traffic analytics, and SNMP for device monitoring. For entry-level, show that you understand the 'why' behind monitoring rather than just 'what to measure.'
Focus Topics
Capacity Planning and Network Growth
Understand how to plan for network growth: monitoring utilization trends, predicting when upgrades are needed, understanding bottlenecks, and planning expansion. Link to the job responsibility of capacity planning.
Practice Interview
Study Questions
Packet Analysis and Protocol Inspection
Understand how to capture and analyze packets using Wireshark or tcpdump. Know how to interpret packet headers, identify protocol anomalies, and use packet analysis to diagnose issues like DNS failures, TCP retransmissions, or protocol-specific problems.
Practice Interview
Study Questions
Network Performance Analysis and Metrics
Understand key performance metrics: latency, jitter, packet loss, throughput, and how they relate to user experience. Know tools for measuring performance: ping, iperf, traceroute with latency analysis, and understanding when performance is inadequate.
Practice Interview
Study Questions
Network Monitoring and Observability
Understand monitoring approaches: SNMP for device metrics, NetFlow/sFlow for traffic analysis, packet capture and analysis with tcpdump/Wireshark, and syslog for event logging. Know what each tool provides and limitations.
Practice Interview
Study Questions
Complex Multi-Layer Troubleshooting Scenarios
Diagnose issues that span multiple network layers or combine multiple failure modes. Practice scenarios like MTU issues in tunnels, packet fragmentation problems, performance issues with specific traffic types, and intermittent connectivity. Know when to involve application teams.
Practice Interview
Study Questions
Onsite Interview 3 - Network Security, Configuration, and Behavioral
What to Expect
This round combines technical assessment of network security and device configuration practices with behavioral evaluation. A network security specialist or senior engineer will assess your understanding of firewall concepts, access control, security best practices, and how you configure network equipment securely. Additionally, a Google team member will conduct a behavioral interview using Google's standard framework to assess collaboration, teamwork, communication, learning ability, and cultural fit. The behavioral portion uses the STAR method and asks about your past experiences demonstrating these qualities.
Tips & Advice
For the technical portion: Study firewall concepts (stateful vs stateless filtering, ACLs, NAT/PAT), understand basic cryptography (why encryption matters), know security hardening practices (least privilege, default deny, logging), and be ready to discuss how to configure network equipment securely. Reference the job requirement to implement network security measures. For behavioral: Prepare 5-6 stories using the STAR method about times you collaborated with teammates, learned something new, solved a problem creatively, handled feedback, or dealt with ambiguity. Google values learning ability and collaboration highly for entry-level candidates. Emphasize growth mindset and willingness to learn from team members.
Focus Topics
Google Behavioral: Problem-Solving and Initiative
Discuss times you identified problems proactively, took initiative to solve them, thought creatively about solutions, and followed through. Show both technical and interpersonal problem-solving examples.[2]
Practice Interview
Study Questions
Secure Network Device Configuration
Know best practices for configuring network equipment securely: strong authentication, logging and monitoring configuration changes, disabling unnecessary services, keeping firmware updated, and using secure management protocols.
Practice Interview
Study Questions
Google Behavioral: Learning Ability and Growth Mindset
Show eagerness to learn, examples of quickly picking up new technologies or domains, adaptability to change, and ability to learn from feedback. Demonstrate intellectual curiosity about networking.[2]
Practice Interview
Study Questions
Network Access Control and Least Privilege
Understand the principle of least privilege, how to design access control policies, segmentation using firewalls/ACLs/VLANs, and verification that users/systems have only necessary access. Know why this matters for security.
Practice Interview
Study Questions
Firewall Fundamentals and ACLs
Understand stateful vs stateless firewalls, how firewalls inspect traffic, basic ACL syntax and logic, implicit deny rules, and how to write firewall rules for common scenarios. Know limitations and when firewalls alone aren't sufficient.
Practice Interview
Study Questions
Google Behavioral: Collaboration and Teamwork
Demonstrate ability to work effectively with team members, ask for help when needed, share knowledge, and contribute to team goals. Use STAR method to discuss times you collaborated cross-functionally or supported teammates.[2]
Practice Interview
Study Questions
Frequently Asked Network Engineer Interview Questions
What is the difference between 'culture fit' and 'culture add', and which do you think better describes you as a candidate? Give one concrete example of a perspective, skill, or way of working you would bring to a team that is not already well represented there.
Sample Answer
Direct answer
Culture fit asks whether you already share a team's existing norms and behaviors; culture add asks what you would bring that the team does not already have. I would describe myself mostly as a culture add: I share the fundamentals a team needs to trust me (reliability, candor, respect for other people's time), but the useful thing I offer beyond that is a genuinely different working background rather than a mirror of the team that is already there.
Structured elaboration
- Define both terms precisely before answering for yourself. Culture fit is about alignment on shared behaviors and values: does this person operate the way we already operate. Culture add is about complementary difference: does this person's background, working style, or perspective fill a gap the team doesn't currently have.
- Explain why the distinction matters, not just define it. A team optimized purely for fit tends toward groupthink: everyone reasons the same way, so blind spots go unchallenged and the same kinds of mistakes recur. A team that only adds without any shared fit becomes uncoordinated: people can't predict each other's reasoning enough to move fast together. The healthy target is fit on a small number of load-bearing behaviors (honesty, follow-through, respect) plus deliberate add on everything else.
- Give a genuine, specific example of your own add, not a generic trait. Vague claims ("I bring diverse perspectives") are the single most common failure mode here; a strong answer names the concrete gap and the concrete evidence.
- Anticipate the natural follow-up: how do you know your difference is actually useful, versus just different for its own sake. The answer is to point at a specific decision, disagreement, or piece of feedback that changed because of the difference you brought, not just a credential or background fact.
Worked example
Suppose your last two teams were both product engineering teams building consumer-facing features, and the team you're interviewing for is mostly staffed by engineers with that same background. Your own prior role was on a data-platform team, closer to the systems that feed those consumer features than to the features themselves. A concrete add-story: in a past project, a product team wanted to ship a new recommendation feature quickly; because of your platform background, you asked a question the rest of the team hadn't raised (whether the upstream data pipeline's freshness guarantees actually matched what the feature's UI implied to users), which surfaced a real gap between a 24-hour batch refresh and a UI copy that said "updated just for you." The team fixed the copy and adjusted the refresh cadence before launch rather than after a user complaint. That is a genuine add: a different background produced a question the existing team composition was less likely to ask on its own, and it changed a real outcome.
Trade-offs & pitfalls
The common failure is answering only the definitional half (correctly explaining fit versus add) and then, when asked for a personal example, retreating to generic self-description ("I'm a good communicator", "I care about quality") that any candidate could say and that does not actually demonstrate difference. A second pitfall is overcorrecting into implying you don't fit at all; the strongest answers are explicit that you also share the small set of behaviors every functioning team needs, and that add is about everything on top of that baseline, not a replacement for it.
A non-technical customer asks whether to spend on more bandwidth or on reducing latency. How do you explain the trade-off and guide the decision?
Sample Answer
Direct answer
Bandwidth is how much you can move per second. Latency is how long one message takes to get there and back. Which one to buy depends on what the customer does: big transfers are limited by bandwidth, interactive or chatty work is limited by latency. Start by measuring, not buying: if the link is already close to full at peak, buy bandwidth. If it is mostly idle and things still feel slow, bandwidth will not help and the fix is shorter distance or fewer round trips.
Explaining it simply
Use a road. Bandwidth is the number of lanes: more lanes carry more cars at once. Latency is the length of the road: a longer road means every car takes longer, no matter how many lanes. A truck delivering 10,000 parcels cares about lanes. A courier who must make twelve trips to the same office to finish a job cares about the length of the road.
A worked comparison with numbers
Assume a web page of 2 MB that needs 12 round trips to load (an assumed figure for the example: connection setup, then several rounds of requests). Time is about (round trips x round-trip time) plus (size / bandwidth).
| Plan | Round-trip time | Bandwidth | Time to load |
|---|---|---|---|
| Today | 80 ms | 100 Mbit/s | 12 x 0.08 + 16 / 100 = 0.96 + 0.16 = 1.12 s |
| 5 times the bandwidth | 80 ms | 500 Mbit/s | 0.96 + 0.032 = 0.992 s |
| Half the latency (move closer, for example a nearby region or a CDN, content delivery network) | 40 ms | 100 Mbit/s | 0.48 + 0.16 = 0.64 s |
Five times the bandwidth saves 0.128 s here. Halving the latency saves 0.48 s. For a large file transfer the answer reverses: 2 GB (16,000 Mbit) at 100 Mbit/s takes about 160 s, at 500 Mbit/s about 32 s, while the round-trip time barely matters.
A single connection can also be limited by latency itself: A TCP window is how much data the sender may have in flight, sent but not yet acknowledged; it must then wait a full round trip for the acknowledgement before sending more. With a 64 KB (65,535 byte) window, throughput is at most window / RTT: 65,535 x 8 = 524,280 bits per round trip, so at 100 ms that is 524,280 / 0.1 = 5.24 Mbit/s and at 10 ms it is 524,280 / 0.01 = 52.4 Mbit/s, however big the link. (Modern systems scale the window, but the principle stands: a long path needs more data in flight to fill it. For a 1 Gbit/s link at 100 ms, that is 1,000,000,000 bits/s x 0.1 s / 8 = 12,500,000 bytes, 12.5 MB.)
How I would guide the decision
- Ask what they do: large backups and video downloads, or interactive apps, calls and many small API requests.
- Measure: peak and 95th percentile utilization on the link, the round-trip time and loss to the places that matter.
- Apply a rule: peak utilization above about 70 to 80 percent for sustained periods means buy bandwidth (the exact threshold depends on how bursty the traffic is). Low utilization with slow interactive work means attack latency.
- Say what lowers latency, and that you cannot buy it from the ISP past a limit: distance sets a floor, so the fix is placing services nearer, using a CDN, reducing the number of round trips in the application, or choosing a better path.
- Compare cost per outcome, such as seconds saved per transaction, not price per Mbit/s.
What I would recommend
For a customer that does not know: measure for a week, then spend on whichever the data points to. If forced to choose without data and the workload is interactive or real-time, put the money toward latency. If it is bulk transfer, put it toward bandwidth.
Pitfalls
- Buying bandwidth because a speed test number looks low, when the real problem is round trips or packet loss.
- Promising a latency improvement the physical distance cannot deliver.
- Using the vendor's "up to" speed instead of measured throughput.
Why do IPv6 LAN subnets almost always use a /64? Are there legitimate cases for any other prefix length on a link?
Sample Answer
Direct answer
An IPv6 address is 128 bits, and the standard architecture splits a unicast address (one that names a single interface) into a network prefix and a 64-bit interface identifier (the host part). Automatic address configuration on a LAN and several security and privacy mechanisms assume that 64-bit host part, so the LAN prefix must be exactly 64 bits long: a /64. Yes, other lengths are legitimate on links that do not use those mechanisms: /127 on router-to-router point-to-point links (RFC 6164), and /128 for a single loopback address.
Why /64 on a LAN
SLAAC is the direct dependency; the other points are further mechanisms written around the same 64-bit boundary.
- SLAAC (stateless address autoconfiguration): a host builds its own address by taking the /64 prefix from a router advertisement (the router's periodic or on-request message announcing prefixes, sent with ICMPv6, the IPv6 control-message protocol) and appending an interface identifier it generates itself. For Ethernet-style links the interface identifier is 64 bits, so a prefix longer than /64 leaves no room for it and SLAAC cannot work. RFC 4862 (the SLAAC specification, section 5.5.3) says a Prefix Information option MUST be ignored when the prefix length plus interface-identifier length does not equal 128 bits, and RFC 7421 analyses how the 64-bit identifier is assumed across many specifications. So on an Ethernet-style link a prefix longer than /64 does not autoconfigure, and you need manually configured addresses or DHCPv6.
Worked build: the router advertises 2001:db8:301:10::/64 and the host picks the 64-bit identifier 0211:22ff:fe33:4455 (illustrative). Prefix, then identifier: 2001:db8:301:10: + 0211:22ff:fe33:4455 = 2001:db8:301:10:211:22ff:fe33:4455. The prefix supplies the first 64 bits and the identifier the last 64, so a /72 prefix would leave only 56 bits and the host could not fill the gap. - Privacy and security mechanisms: temporary privacy addresses and cryptographically generated addresses are built on a 64-bit identifier. CGA (cryptographically generated addresses, where the identifier is derived from the host's public key so a host can prove an address is its own; Secure Neighbor Discovery uses it) is defined around a 64-bit identifier, so RFC 7421 notes that moving the /64 boundary would invalidate that definition. The same RFC also points out that a short host part lets a scanner find every host by trying neighbouring addresses: entropy (the number of equally likely values, here 2^64 identifiers) is what makes guessing impractical, and long prefixes such as /120 leave only a few hundred values to try.
- Other protocols assume it: the multicast method of RFC 3306 (building a group address that embeds the network's unicast prefix) assumes a longest prefix of 64 bits, and some transition mechanisms (for example 6to4) assume a 64-bit identifier too, so a non-/64 LAN works only until it meets one of them. NAT64 (which translates between IPv6 and IPv4) is more tolerant: RFC 7421 says it works with a subnet boundary out to /96, which is why it is not a reason for /64.
- Space is not the constraint: a /64 holds 2^64 = 18,446,744,073,709,551,616 addresses, and a /48 site allocation contains 2^16 = 65,536 /64s, so not saving addresses is the intended trade. Subnetting a LAN smaller is therefore not an optimisation.
When another length is right
| Length | Where | Why |
|---|---|---|
| /127 | Router-to-router point-to-point links | RFC 6164 requires routers to support it. On a prefix shorter than /127, such as /64, a packet to an unassigned address on the link used to be able to bounce between the two routers until its hop limit (the counter every router decrements, dropping the packet at 0) expired: for example router A sends a packet for the unused ::99 across the link, router B sees the address as on that same link and sends it back, and so on (the ping-pong problem). Current ICMPv6 (RFC 4443, section 3.1) largely closes that loop, because a router that receives a packet on a point-to-point link for an address in that link's subnet, other than its own, MUST NOT forward it back onto the arrival link, and RFC 6164 notes this mitigation; the loop is therefore a historical reason and a risk only on routers that do not implement it. The stronger remaining reason is neighbor cache exhaustion: an attacker can send traffic to many of the unassigned addresses (about 2^64 of them on a /64), and each one makes a router create an unresolved entry in its neighbor cache (the table mapping IPv6 addresses to link-layer addresses) and send a lookup that is never answered, which eventually exhausts memory and processing. A /127 leaves exactly two addresses, so neither the loop nor the exhaustion is possible, and RFC 6164 says assigning /127 eliminates the exhaustion problem completely. |
| /128 | Loopback or a single service address | Identifies one host or router, with no link attached. |
| /64 anyway on router links | Some operators still do this | Consistency and simpler automation, mitigated with ACLs (access control lists) that drop traffic to unused addresses and neighbor-cache limits, but /127 removes the problem rather than mitigating it. |
Cloud networks follow the same convention: the AWS documentation example gives a VPC a /56 IPv6 block and shows a subnet carved from it as a /64.
Worked example
A site is delegated 2001:db8:301::/48 (a documentation prefix). Its user VLAN is 2001:db8:301:10::/64 and hosts self-configure from it. The WAN link between the site router and the regional router is 2001:db8:301:ffff::2/127, which holds exactly two addresses: the site router takes ::2 and the regional router ::3, and nothing else on that link exists to attract traffic. The /127 deliberately does not start at ::0: RFC 6164 advises against assigning an address whose last 64 bits are all zero, because that is the Subnet-Router anycast address (and routers must disable that anycast for a /127). The router loopback is a /128 taken from a separate loopback block.
Pitfalls
- A /112 or /120 LAN to "limit the hosts" breaks SLAAC: hosts do not autoconfigure, and the fix is the prefix, not a router flag. Advertising a prefix shorter than /64 for SLAAC does not work either; the length SLAAC expects is 64.
- Treating /64 as waste and trying to conserve; the scarce thing in IPv6 is routing-table size and plan discipline, not addresses.
Design a pair of POPs that deliver private layer 3 VPN service to an enterprise customer with control-plane and data-plane redundancy. How does the design behave when one POP fails, and how do you avoid loops and black holes during a partial failure?
Sample Answer
Direct answer
Build two POPs (points of presence), each with provider edge routers (PEs, which hold per-customer VRFs, i.e. separate routing tables), core routers (P) and one route reflector (RR, a router that relays BGP routes so PEs need not peer with each other). The customer is dual-homed: one circuit to a PE in each POP. Give every PE its own route distinguisher (RD) per VRF, so the RR passes both paths instead of hiding the backup. When a POP fails, BFD (a fast liveness check) drops the attachment in 0.15 s, the surviving POP already holds the customer's remote routes, and the loop and black-hole controls described below keep the partial failures from becoming outages.
Plain-language terms used below. A CE (customer edge) is the customer's own router. A route distinguisher (RD) is an 8-byte tag the provider prepends to a customer prefix so that two customers using the same address range, or two paths to one prefix, stay distinct in BGP (RFC 4364). A route target (RT) is a tag on a route that says which VRFs may import it; the export list is the tags a VRF attaches and the import list is the tags it accepts. An IGP (interior gateway protocol, for example OSPF or IS-IS) carries the provider's internal routes; LDP (label distribution protocol) hands out the MPLS labels that forward packets across the core. iBGP is BGP between routers in the same AS (autonomous system, identified by an AS number). A next hop is the router a route says to send traffic to. A black hole is traffic that is silently dropped.
Layout (names are examples)
- POP-1: pe-1a, pe-1b, p-1, rr-1. POP-2: pe-2a, pe-2b, p-2, rr-2.
- Customer ce-x connects to pe-1a and pe-2a with one 1 Gbps circuit each. This one customer's peak is 0.6 Gbps; the 30 Gbps used in the capacity paragraph is the combined peak of all customers on the POP pair.
- Provider AS 65000 (example). Every PE has iBGP sessions to rr-1 and rr-2. rr-1 and rr-2 share one cluster ID (RFC 4456 says redundant RRs in one cluster share it, so they can discard each other's reflected copies), and each uses its own router ID.
- Core links run an IGP with BFD, MPLS labels via LDP or segment routing, and 2 x 100G between POPs.
Control-plane redundancy
Two RRs, every PE peering with both, means losing one RR changes nothing in forwarding. Unique RDs matter: RFC 4364 allows different routes to the same CE to carry different RDs. Concretely (illustrative values), ce-x's prefix 10.1.0.0/24 is advertised by both PEs:
| RD choice | What pe-1a advertises | What pe-2a advertises | What the RR sends to a remote PE |
|---|---|---|---|
| Unique RD per PE | 65000:101 + 10.1.0.0/24, next hop pe-1a | 65000:201 + 10.1.0.0/24, next hop pe-2a | both, because they are different routes |
| Shared RD | 65000:100 + 10.1.0.0/24, next hop pe-1a | 65000:100 + 10.1.0.0/24, next hop pe-2a | only the best one, say pe-1a's |
With a shared RD the RR would pick one best path and send only that to clients, so the backup path through the other POP would be invisible and convergence would wait for a BGP update. The alternative is BGP Add-Path (RFC 7911), which lets a router advertise several paths per prefix; I prefer unique RDs because they need no capability negotiation.
RR scale: with 400 customer VPNs, 150 routes each and 2 PEs per customer, each RR receives 400 x 150 x 2 = 120,000 VPN paths in either design. With unique RDs they are 120,000 distinct VPN routes and the RR reflects all of them, so every PE that imports them holds 120,000; with a shared RD they collapse to 60,000 routes, the RR selects one best path for each and reflects only those, so clients hold 60,000. Unique RDs therefore cost twice the VPN routes in every client and in each RR's outbound updates, which is the price of keeping the backup path visible. Partition customers over more RR pairs when a pair reaches a set fraction of your RR platform's documented limit; 70% is an illustrative headroom choice, not a vendor figure.
Data-plane redundancy and behaviour when a POP fails
| Failure | Detection | What happens |
|---|---|---|
| CE to PE circuit fails | BFD 50 ms x 3 = 0.15 s | PE withdraws the route, traffic uses the other POP |
| pe-1a node fails | IGP BFD on core links, 0.15 s (assuming the same 50 ms x 3 timers as the customer circuits) | Remote PEs lose the next hop for pe-1a, drop its paths and use pe-2a's |
| rr-1 fails | iBGP session down | Paths from rr-2 already present, no data-plane change |
| POP-1 entirely down | All of the above | Customer traffic uses POP-2 only |
Capacity: assume the combined peak of all customers on the POP pair is 30 Gbps (an assumption; ce-x is 0.6 / 30 = 2% of it), spread evenly over the four PEs. Each PE has 2 x 100G = 200G to the core, so normal load is 30 / 4 = 7.5 Gbps per PE, 3.75% of 200G. After a POP failure the two PEs of the surviving POP carry 30 / 2 = 15 Gbps each, 7.5% of 200G. Customer circuits: active-active puts 0.3 Gbps on each 1 Gbps circuit and 0.6 Gbps (60%) on the one left. Nothing runs at 100%.
Avoiding loops
- Site of Origin (SoO): RFC 4364 defines an extended community that identifies the originating site and says a route must never be redistributed to a CE at that site. Set it on each customer site so a route learned from ce-x via pe-1a is not sent back to it via pe-2a. This matters especially when sites reuse one AS number, because ordinary BGP loop detection then fails.
- RR loops: ORIGINATOR_ID and CLUSTER_LIST (RFC 4456) let routers drop a route that comes back to the router or cluster that originated it.
- Customer backdoor links between sites need their own metrics so they do not become a transit path between POPs.
Avoiding black holes in a partial failure
- Control plane up, labels down: IGP reaches a PE loopback but LDP has no label for it, so BGP says "reachable" and packets drop. Turn on LDP-IGP synchronisation (RFC 5443), which gives a link maximum cost until LDP is up on it, or use segment routing, so a link is not used until labels exist.
- A PE cut off from the core but with live customer sessions: it keeps attracting customer traffic it cannot forward. Track the core-facing interfaces and, if all are down, shut the CE-facing sessions so the CE fails over.
- Stale remote next hops: remote PEs use next-hop tracking, so loss of a PE loopback in the IGP invalidates its VPN paths in IGP time, not BGP hold-timer time.
Route leaks, traffic engineering and RTBH
Prefix filtering with a max-prefix limit and RTBH apply to every customer; the traffic-engineering tools are used only when a customer buys a specific service, and the BGP roles attribute applies on the provider's own external peers.
- Leak prevention: accept from each CE only prefixes the customer owns and set a max-prefix limit that shuts the session (RFC 7454). Keep RT import and export lists narrow, with no default route leaking between VRFs. On the provider's own external peers use BGP roles and the Only to Customer attribute (RFC 9234).
- Traffic engineering: run IGP metrics with equal-cost multipath first. Add segment routing policies, or IGP Flexible Algorithm minimum-delay paths (RFC 9350), only for a customer who buys a latency tier. RSVP-TE adds per-tunnel state that this scale does not need.
- RTBH (remotely triggered black hole): a customer sends a /32 inside its own space with the BLACKHOLE community 65535:666 (RFC 7999). The PE accepts it only if it is a /32 within the customer's registered prefixes, sets the next hop to a discard route, and keeps it from being exported (RFC 7999 says to add NO_EXPORT or NO_ADVERTISE and that honouring it must be agreed in advance).
Pitfalls
- Verify failover with a test, not a diagram: pull each circuit and each POP link in a lab and record packet loss.
- A shared RD and a single RR pair looks redundant and hides the backup path.
- Do not claim sub-second end-to-end convergence from BFD times alone. Only the detection step is 0.15 s. Route withdrawal and forwarding-table updates add to it, and need measuring on your platform.
Construct a BPF filter for tcpdump or a display filter for Wireshark/tshark that isolates TCP retransmissions on a busy interface, and explain the difference between what a capture filter and a display filter can each see and why that distinction matters when you're trying to keep a production capture small.
Sample Answer
Direct answer
A capture filter (tcp[tcpflags] & (tcp-syn|tcp-ack) != 0 alone won't isolate retransmissions specifically, since BPF has no built-in concept of "this segment I've seen before") is fundamentally limited here; retransmission detection genuinely requires STATE across multiple packets (has this exact sequence number been seen already), which a stateless capture filter can't express, so isolating retransmissions is really a DISPLAY-filter (or post-processing) job, not a capture-filter one.
Structured elaboration
- Why capture filters (BPF) can't do this well: BPF filters evaluate each packet independently, with no memory of prior packets; "is this a retransmission" requires comparing THIS packet's sequence number against what's already been seen for this same TCP stream, which is inherently stateful and outside what a capture filter can express.
- What Wireshark's display filter CAN do, because it operates AFTER full protocol dissection with state:
tcp.analysis.retransmissionis a Wireshark-computed field, built by Wireshark's own stateful TCP stream tracking during dissection, not something derivable from a stateless per-packet filter; this display filter, applied in Wireshark or viatshark -Y, correctly isolates retransmissions because Wireshark has already done the stateful bookkeeping. - The practical implication for capturing on a busy interface: since you can't cheaply filter for retransmissions AT CAPTURE TIME, the practical approach is to capture the full (or reasonably filtered by host/port) traffic to a file, then apply the display filter afterward during analysis, accepting a larger capture file in exchange for being able to ask this specific, stateful question after the fact.
- Capture filters versus display filters, the general principle: capture filters (BPF) are for REDUCING VOLUME at the point of capture using only per-packet, stateless criteria (host, port, protocol, flags); display filters (Wireshark's own syntax) can express much richer, STATEFUL, cross-packet logic, because they run against already-captured, already-parsed data where the tool has built up state across the whole stream.
- Why you'd prefer one over the other in production: a capture filter is cheaper (reduces what's written to disk or memory at all) and appropriate when you know in advance a simple, stateless criterion (a specific host/port) will scope the capture usefully; a display filter is necessary whenever the question itself requires state (retransmissions, duplicate ACKs, a specific stream's full analysis), and in that case you must capture broadly enough first, then filter afterward.
Worked example
On a busy interface, capture with a modest capture filter scoping to the relevant host/port (tcp and host 10.1.1.10) to keep the file a manageable size, then open it in Wireshark or run tshark -r file.pcap -Y 'tcp.analysis.retransmission' to isolate the retransmissions specifically; attempting to write an equivalent BPF capture filter for "retransmissions only" isn't achievable, because BPF has no mechanism to remember which sequence numbers it's already seen for a given stream.
Trade-offs & pitfalls
A common misunderstanding is assuming any packet-matching criterion can be expressed as a capture filter if you just find the right syntax; retransmission detection specifically cannot, because it requires state BPF fundamentally doesn't carry across packets. Recognizing this distinction (stateless capture-time filtering versus stateful post-capture analysis) up front saves time that would otherwise be spent trying to force a capture filter to do something it structurally cannot.
Compare terminating TLS at the load balancer (offloading) versus end-to-end TLS between clients and backend services in a microservices architecture. Discuss impacts on security (encryption, inspection), observability (logging, tracing), performance, certificate management, and intrusion detection. Provide a recommended configuration for a moderate-security environment and justify trade-offs.
Sample Answer
Terminating Transport Layer Security (TLS, the protocol that encrypts and authenticates a connection) at the load balancer decrypts traffic once at the network edge and forwards it onward; end-to-end TLS keeps every hop between client and backend encrypted independently. The right choice for a moderate-security environment is usually neither pure extreme but a hybrid: terminate and re-encrypt.
Comparison
| Dimension | Terminate at load balancer | End-to-end TLS |
|---|---|---|
| Security | Public hop encrypted; internal hop between load balancer and backend is plaintext unless separately protected, widening blast radius if the internal network is compromised | Every hop encrypted, no plaintext segment even inside the data center |
| Observability | Load balancer, web application firewall (WAF), and network-based intrusion detection can inspect decrypted content directly | Network-based tools behind the load balancer see only ciphertext; visibility requires app-layer or service-mesh-level logging instead |
| Performance | TLS handshake and cryptographic cost absorbed once, often on hardware built for it | Every service-to-service hop pays its own handshake cost, mitigated by session reuse/connection pooling but still added overhead |
| Certificate management | One certificate (or a small set) at the edge to issue and rotate | Every service needs its own identity certificate; effectively requires an automated internal certificate authority to be operable at scale |
| Intrusion detection | Full inline visibility for anything placed behind the load balancer | A network-based sensor behind the load balancer is blind to payload; detection must move to the service layer |
Recommended configuration for a moderate-security environment
Terminate public TLS at the load balancer or WAF so it can inspect and enforce policy on the incoming traffic, then re-encrypt (mutual TLS, mTLS, where both sides present certificates) from the load balancer into the internal network rather than passing plaintext to the backend. This "terminate and re-encrypt" pattern keeps the inspection and simple centralized certificate management benefits of edge termination while removing the plaintext internal hop that pure offloading leaves exposed. It is a meaningfully lower operational cost than full service-to-service mTLS across every microservice, while closing the biggest single gap (an unencrypted internal segment) that pure offloading has.
When full end-to-end mTLS is worth it
Justify moving to full mesh-wide mTLS between every service pair once compliance requirements demand it (handling regulated data where no internal plaintext hop is acceptable under any circumstance) or once the organization is deliberately building toward a zero-trust internal network posture and can absorb the operational cost of automated short-lived certificate issuance and rotation across every service, since doing that manually at microservice scale is not realistic.
How does continuous authentication and authorization differ from a one-time login? What signals (behavioral, location, device posture) should trigger re-authentication or an adaptive change in access, and how do you avoid re-prompting the user so often that they get fatigued?
Sample Answer
Direct answer
A one-time login checks identity once, at the start of a session, then trusts that session until it expires or is explicitly logged out. Continuous authentication and authorization keep re-evaluating trust throughout the session as new signals arrive, so access can be tightened, challenged, or revoked mid-session if something changes, not only at the door.
Structured elaboration
Signals that should trigger re-authentication or an adaptive change in access:
- Behavioral: a user performing actions well outside their normal pattern, such as bulk-querying data they never normally touch, or an unusually rapid sequence of requests that looks automated.
- Location: a login or request originating from a network location or geography inconsistent with the user's recent activity, sometimes described as impossible travel, an active session in one place and a new request appearing to originate from somewhere else shortly after.
- Device posture: the device's security state changing mid-session, disk encryption disabled, an outdated patch level, malware detection triggering, or the device no longer matching the enrolled, managed device that started the session.
Not every signal should trigger the same response. A well-designed system grades severity: a low-confidence anomaly might just be logged; a moderate anomaly might trigger a step-up challenge, such as an additional multi-factor authentication (MFA) prompt, scoped to the specific sensitive action being attempted rather than the whole session; a high-confidence anomaly, a clearly failed device-posture check or a hijacking signal, should terminate the session and force full re-authentication.
Avoiding re-prompt fatigue:
- Scope step-up challenges to the specific risky action, not the whole session, so a user browsing normal, low-sensitivity resources is never interrupted.
- Use risk-based thresholds rather than fixed intervals. Prompting every user every fifteen minutes regardless of behavior trains people to reflexively approve prompts, a habit attackers exploit through prompt bombing; prompting only when a meaningful signal changes keeps prompts rare enough to be taken seriously.
- Prefer passive signals, device posture, network reputation, behavioral baselining, over active prompts wherever possible, since passive checks add friction to the system rather than to the user, escalating to an active prompt only when those passive signals actually indicate elevated risk.
Worked example
A user logs in from a managed laptop on the corporate network at 9am, low risk, no prompts needed for normal work. At 2pm, the same session starts downloading a far larger volume of customer records than that user has ever accessed in one sitting, a behavioral anomaly, while the device's posture check reports its disk encryption was disabled ten minutes earlier, a device-posture anomaly. Both signals firing together push the computed risk well past a step-up threshold into a high-risk band, so the system does not just ask for MFA, it suspends the download and forces full re-authentication plus a device compliance check before any further access to customer data, while the user's earlier, low-risk browsing that morning was never interrupted at all.
Trade-offs and pitfalls
The common wrong turn is treating every anomaly as equally important and re-prompting for anything unusual, which produces fatigue and trains users to click through prompts without reading them. The fix is grading signals by confidence and severity and scoping the response to the specific action at risk, not the entire session.
Tell me about something technical you taught yourself recently that nobody asked you to learn. What made you decide it was worth your time, how did you go about it, and what changed at work because you did?
Sample Answer
Direct answer
In the last year I taught myself how to read query execution plans and reason about indexing, not because anyone assigned it, but because a recurring internal report kept getting slower and nobody had the bandwidth to look into why. I spent a handful of evenings learning to read plan output and understand how the database chooses an access path, then applied it directly to that report's query rather than treating it as a side hobby, and the fix noticeably shortened a report that had become one of the slowest in the weekly batch.
Structured elaboration
- Justify the "why this and not something else": pick something tied to a real, recurring cost you already feel, a slow report, a repeated manual step, a bug class that keeps recurring, rather than a trending technology with no attachment to your actual work.
- Keep the learning self-structured: with no assigned curriculum, the plan is whatever sequence of official docs and small experiments gets to "I can predict what this will do" fastest.
- Validate the new understanding against people who already know the area, even when nobody assigned this; a quick review confirms the understanding is actually right, not just plausible.
- Land it back in the work rather than a personal notebook; the skill only counts, for real impact and for describing it later, once it is applied to something that mattered.
- Check whether it stuck: months later, are you still reaching for it, or did it fade once the original problem was solved?
Worked example
A weekly finance reconciliation report kept taking noticeably longer to run as data grew, and it kept getting flagged as "just slow" without anyone owning a fix. Outside assigned work, I spent a handful of evenings over two weeks working through documentation on how a query planner chooses between an index and a full scan, reproducing small example queries locally rather than only reading passively. I then applied the plan-inspection tooling directly to the report's slowest query and found it was doing a full table scan on a column with no index, caused by an implicit type mismatch in a join condition. I added the right index and fixed the mismatch, and had a senior engineer sanity-check the change before it shipped, since this was genuinely new territory for me. The report went from being flagged in every week's slow-query review to not appearing at all. I kept using the same read-the-plan-first habit on later slow queries, so the skill stuck well past the original problem.
Trade-offs and pitfalls
- Self-taught understanding validated only against your own intuition, with no outside check, risks confidently shipping a fix that happens to work on the case you tested but does not generalize.
- Picking a skill purely because it is trendy, with no real problem behind it, produces knowledge that is hard to defend as impact and often does not stick.
- There is a real risk of scope creep: fixing one query can turn into re-architecting a system nobody asked you to touch; the discipline is applying the new skill to the specific problem, not treating it as license for a bigger, unrequested project.
Explain how deploying services across multiple availability zones and regions changes network capacity planning. Cover the effect on latency, inter-region bandwidth, and replication costs, and how you'd factor cross-AZ and cross-region transfer fees into your sizing decisions.
Sample Answer
How the topology changes network capacity planning
Within a single AZ, inter-service network hops are effectively free to plan for: bandwidth is high, latency is low and stable, and cloud providers typically do not charge for traffic that stays within the same AZ. The moment a service spreads across AZs or regions, network capacity planning has to account for three things that were not there before: added latency, the bandwidth the replication/traffic itself needs, and the transfer fees the provider charges for that traffic.
Latency
Cross-AZ traffic within a region adds a small, fairly predictable latency (commonly single-digit milliseconds), usually fine to absorb into a synchronous call path. Cross-region traffic adds latency driven by physical distance (commonly tens to well over a hundred milliseconds), generally too slow for a synchronous request path, so it has to be planned for as asynchronous replication instead. That changes the bandwidth math: you are now sizing for a steady replication stream, not a single request/response pair.
Inter-region bandwidth and replication costs
Size cross-region bandwidth for the actual replication volume, not the request volume: for a database, that is write throughput times average write size times the number of regions receiving a copy; for event/log shipping, it is the raw event volume. Build in headroom above the steady-state rate, because if a target region falls behind (a network blip, a slow follower), it needs to catch up faster than steady-state without permanently falling further behind, which needs spare bandwidth on top of the baseline.
Factoring transfer fees into sizing
Cloud providers typically charge nothing, or very little, for traffic within an AZ, a small per-GB fee for cross-AZ traffic within a region, and a meaningfully larger per-GB fee for cross-region traffic. When sizing, that means: minimize what actually needs to leave the region (only replicate what genuinely needs a remote copy, not everything by default), track the recurring transfer cost as an ongoing line item that scales with traffic growth, not a one-time setup cost, and weigh it in the same trade-off as compute cost when deciding how many regions to run in and how much to replicate to each.
Explain the seven layers of the OSI model. For each layer, state its primary responsibility, the name of its protocol data unit (PDU), and one or two protocols or technologies that commonly operate there. Then explain why identifying which layer a failure sits at is useful before you jump to a fix.
Sample Answer
Direct answer
The OSI model splits network communication into seven layers, each handing off a well-defined unit of work to the layer above and below it: Physical, Data Link, Network, Transport, Session, Presentation, and Application. Knowing which layer a protocol or symptom belongs to lets you reason about failures systematically instead of guessing.
Structured elaboration
| Layer | Responsibility | PDU (protocol data unit) | Common protocols/tech |
|---|---|---|---|
| 7. Application | Provides the interface applications use to talk over the network | Data | HTTP, DNS |
| 6. Presentation | Translates, encrypts, and compresses data into a form the application layer can use | Data | TLS, character encoding |
| 5. Session | Establishes, manages, and tears down a logical session between two hosts | Data | RPC session handling |
| 4. Transport | End-to-end delivery between processes: reliability, ordering, flow control | Segment (TCP) / Datagram (UDP) | TCP, UDP |
| 3. Network | Logical addressing and routing across networks | Packet | IP, ICMP |
| 2. Data Link | Framing and addressing on a single link (same broadcast domain) | Frame | Ethernet, ARP |
| 1. Physical | Raw bit transmission over a physical medium | Bit | Ethernet PHY, fiber, radio (Wi-Fi) |
The mnemonic that matters more than memorizing names is the DIRECTION of responsibility: each layer only needs to trust the layer directly below it to deliver its unit of data, and it only exposes a clean interface to the layer above. That's what lets a Transport-layer protocol like TCP work identically over Ethernet, Wi-Fi, or a VPN tunnel: it never needs to know which Layer 1/2 technology is underneath.
Worked example
Say a user reports "the site is down." Layer-by-layer reasoning turns that vague complaint into a specific hypothesis:
- Physical/Data Link symptom: the NIC shows no link light, or
ip linkreports the interface asDOWNa cable or switch-port problem. - Network layer symptom:
pingto the server's IP times out but the local gateway responds a routing problem somewhere between here and there. - Transport layer symptom:
pingsucceeds but a TCP connection to the port hangs or resets a firewall, a service that isn't listening, or a transport-level issue. - Application layer symptom: the TCP connection completes but the HTTP response is an error or garbage the service is up but misbehaving.
Each of these is a different team, a different fix, and a different urgency. That's the actual payoff of the model: it turns "the site is down" into "which layer, which evidence."
Trade-offs & pitfalls
The OSI model is a teaching and troubleshooting framework, not how real stacks are literally implemented. Session and Presentation are rarely separate pieces of code in modern systems; TLS, for instance, is commonly described as "sits between Transport and Application" rather than cleanly as Layer 6. Don't over-fit a real symptom to exactly one layer: a firewall dropping SYN packets looks like a Transport-layer symptom (connection never establishes) but the actual cause and fix live at a security-policy layer that OSI doesn't model at all.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Network Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs