Senior Network Engineer Interview Preparation Guide - Google
Google's interview process for Senior Network Engineer typically includes an initial recruiter screening, technical phone screens focused on networking fundamentals and troubleshooting, followed by 4-5 onsite rounds covering network architecture design, infrastructure troubleshooting, Google-specific infrastructure knowledge, and behavioral/culture fit assessment. The process emphasizes both deep technical expertise and alignment with Google's values of collaboration, bias to action, and comfort with ambiguity.
Interview Rounds
Recruiter Screening
What to Expect
Initial phone conversation with a Google recruiter to assess your background, experience, motivation for joining Google, and alignment with the role. The recruiter will verify your eligibility, discuss compensation expectations, and determine if you should proceed to technical interviews. This round is non-technical and focuses on your career trajectory and interest in the specific role.
Tips & Advice
Be prepared to discuss your career progression from early networking roles to senior-level expertise. Have specific examples of your achievements and impact ready. Show genuine interest in Google's infrastructure and mission. Ask thoughtful questions about the team, projects, and growth opportunities. Be clear about your salary expectations and flexibility on start date.
Focus Topics
Role-Specific Experience
Your hands-on experience with network architecture design, infrastructure operations, troubleshooting at scale, and work with routing, switching, and security technologies
Practice Interview
Study Questions
Career Trajectory and Progression
Your evolution from junior networking roles to senior-level expertise, key projects that shaped your skills, and why you're ready for a senior role at Google
Practice Interview
Study Questions
Motivation for Google
Your specific reasons for wanting to join Google, familiarity with Google's infrastructure approach, and how this role aligns with your career goals
Practice Interview
Study Questions
Technical Phone Screen 1: Networking Fundamentals and Protocol Deep Dive
What to Expect
First technical phone interview focused on core networking protocols, OSI model layers, and fundamental networking concepts. The interviewer will ask questions about IPv4/IPv6, routing protocols (BGP, OSPF), switching, VLAN concepts, DNS, TCP/IP behavior, and network troubleshooting methodology. You should be able to explain protocols clearly and discuss trade-offs between different approaches. This round assesses your foundational knowledge and communication ability.
Tips & Advice
Review networking protocols in depth, focusing on BGP (critical for senior roles), OSPF, and EIGRP. Be comfortable explaining packet flow at different layers. Know the differences between connection-oriented and connectionless protocols. Practice explaining why certain protocols are used in specific scenarios. Use a whiteboard or shared screen to diagram network flows. Be prepared for follow-up questions that test your understanding, not just memorization. Explain your troubleshooting approach clearly—many questions will be framed as "what would you check first" scenarios.
Focus Topics
Switching, VLANs, and Layer 2
VLAN tagging and trunking, STP/RSTP for loop prevention, MAC learning, spanning tree topology selection, access vs trunk ports, and inter-VLAN routing design
Practice Interview
Study Questions
DNS and Name Resolution
DNS query process, recursive vs iterative resolution, DNS record types (A, AAAA, CNAME, MX, TXT, NS), DNS caching behavior, and common DNS troubleshooting scenarios
Practice Interview
Study Questions
Routing Protocols and BGP
Deep understanding of Border Gateway Protocol (BGP), interior gateway protocols (OSPF, EIGRP), path selection, convergence behavior, and when to use each protocol in different network designs
Practice Interview
Study Questions
TCP/IP Protocol Suite
TCP vs UDP behavior, connection establishment, flow control, congestion control, segment sizing, retransmission behavior, and how these affect application performance
Practice Interview
Study Questions
Network Troubleshooting Methodology
Systematic approach to isolating network issues: verifying connectivity at each layer, using tools (ping, traceroute, netstat, ss, tcpdump), distinguishing DNS vs routing vs firewall issues, and identifying the root cause logically
Practice Interview
Study Questions
Technical Phone Screen 2: Network Design and Troubleshooting at Scale
What to Expect
Second technical phone interview focused on network architecture design, infrastructure troubleshooting at scale, and system-level thinking. You may be given a scenario like designing a network for a distributed data center environment, handling traffic growth, or troubleshooting a complex connectivity issue. The interviewer assesses your ability to think about trade-offs, scalability, redundancy, and security. You should discuss design decisions and explain why you chose specific approaches. This round often includes a semi-open-ended design problem or complex troubleshooting scenario.
Tips & Advice
Practice designing networks end-to-end: start with requirements gathering, discuss constraints (bandwidth, latency, cost, availability), propose architecture with redundancy, consider failure modes, and evaluate trade-offs. For troubleshooting scenarios, explain your diagnostic approach out loud and ask clarifying questions about symptoms. Think about monitoring, observability, and how you'd detect issues. Be ready to explain why you avoided certain technologies or approaches. Discuss capacity planning, growth scenarios, and how your design scales. Mention security considerations and how your design addresses them. Show familiarity with modern approaches like Software-Defined Networking (SDN), automation, and cloud-native networking if relevant to your experience.
Focus Topics
Network Security in Design
Incorporating security into network architecture: DMZs, firewalling, access control lists (ACLs), network segmentation, VPN design, DDoS mitigation, and balancing security with performance
Practice Interview
Study Questions
Network Monitoring and Observability
Designing systems to be observable: flow monitoring, NetFlow/sFlow, SNMP, metrics collection, logging, alerting strategy, and using data to optimize network performance
Practice Interview
Study Questions
Network Architecture Design for Scale
Designing resilient, scalable network topologies supporting distributed systems. Understanding multi-layer architecture, redundancy strategies, load balancing, and handling growth while maintaining performance and reliability
Practice Interview
Study Questions
Complex Troubleshooting Scenarios
Diagnosing multi-layer issues involving routing, firewalls, DNS, and application behavior. Isolating which layer is causing problems, collecting relevant data, and forming hypotheses based on evidence
Practice Interview
Study Questions
High-Availability and Redundancy Design
Designing networks with no single points of failure, failover mechanisms, multi-path routing, redundant hardware, and assessing trade-offs between complexity and availability
Practice Interview
Study Questions
Onsite Round 1: Network Architecture Design Interview
What to Expect
In-person or video interview focused on designing a medium-to-large scale network architecture. You'll be given a business scenario (e.g., connecting multiple data centers, supporting a growing user base, or migrating infrastructure) and asked to design a network that meets specified requirements. The interviewer will probe your design decisions, ask about trade-offs, and explore edge cases. You should demonstrate understanding of scalability, redundancy, cost optimization, and security. This round assesses your architectural thinking and communication ability.
Tips & Advice
Start by clarifying requirements and constraints before designing. Ask about scale (number of users/locations), latency requirements, bandwidth needs, availability targets, and budget. Propose a clear architecture with diagrams. Explain component choices and why you selected specific technologies. Walk through traffic flows in normal and failure scenarios. Discuss scalability as traffic grows. Address security, redundancy, and monitoring from the start. Be prepared to modify your design based on feedback. Show that you're thinking about operational aspects: How is this managed? How are updates deployed? What breaks and how do you detect it? Practice designing on a whiteboard or digital drawing tool. Use standard network notation. Ask clarifying questions when scenarios are unclear.
Focus Topics
Firewall and Security Architecture
Designing stateful firewalls, establishing security zones, implementing access policies, understanding DDoS mitigation, and integrating security without excessive latency
Practice Interview
Study Questions
IP Addressing and CIDR Planning
Designing IPv4 and IPv6 addressing schemes that scale, understanding supernetting and aggregation, planning for NAT-less architectures, and managing address space efficiently
Practice Interview
Study Questions
Capacity Planning and Growth Scenarios
Estimating bandwidth requirements, planning for traffic growth, understanding headroom requirements, and designing infrastructure that can scale incrementally
Practice Interview
Study Questions
Data Center Networking and Inter-DC Design
Designing networks connecting multiple data centers, handling traffic between data centers efficiently, understanding latency constraints, and implementing disaster recovery and business continuity
Practice Interview
Study Questions
Network Topology Trade-offs
Understanding mesh vs hierarchical topologies, spine-leaf architecture, and when to use each approach based on requirements, costs, and operational complexity
Practice Interview
Study Questions
Onsite Round 2: Network Operations and Troubleshooting
What to Expect
Interview focused on operational excellence and troubleshooting complex network issues. You'll be presented with real-world scenarios (e.g., intermittent packet loss, latency spikes, asymmetric routing) and asked to diagnose root causes. The interviewer assesses your systematic troubleshooting methodology, use of tools, and ability to prioritize when multiple issues could explain observed behavior. This round tests practical expertise and your ability to debug production systems under pressure.
Tips & Advice
For troubleshooting scenarios, establish a clear mental model: gather facts first, form hypotheses, test hypotheses, and iterate. Describe the tools you'd use (tcpdump, netstat/ss, ping, traceroute, BGP route inspection, SNMP, NetFlow, etc.) and what you'd learn from each. Practice describing packet flows through your network topology. Be comfortable with Linux networking tools. When faced with ambiguous symptoms, discuss what additional information would be helpful. Show that you understand the difference between correlation and causation. Discuss how you'd approach this in production (e.g., using non-invasive monitoring before running tests that might impact service). Mention metrics you'd review: latency percentiles, packet loss, jitter, retransmission rates. Be ready to pivot if your initial hypothesis is wrong and choose the next most likely cause logically.
Focus Topics
Performance Analysis and Optimization
Identifying performance bottlenecks, analyzing latency sources (propagation, processing, queuing), understanding bandwidth constraints, and optimizing network configurations for application requirements
Practice Interview
Study Questions
Handling Asymmetric Routing and MTU Issues
Understanding and troubleshooting path asymmetry that causes return traffic to take different routes, fragmentation vs MTU path discovery (PMTUD), and implications for TCP performance
Practice Interview
Study Questions
Firewall and NAT Troubleshooting
Diagnosing firewall rule issues, understanding stateful firewall behavior, troubleshooting NAT and port forwarding problems, and interpreting connection tracking information
Practice Interview
Study Questions
Network Tool Proficiency and Instrumentation
Deep knowledge of Linux networking tools (netstat, ss, ip route, iptables, tcpdump), SNMP monitoring, NetFlow analysis, syslog interpretation, and knowing which tool answers which question
Practice Interview
Study Questions
Advanced Troubleshooting and Root Cause Analysis
Systematic methodology for diagnosing complex multi-layer issues, using network tools effectively (tcpdump, netstat, ss, mtr, ping, traceroute, dig), interpreting packet captures, and distinguishing symptoms from root causes
Practice Interview
Study Questions
Onsite Round 3: Google Infrastructure and Cloud Networking
What to Expect
Interview focused on your understanding of modern infrastructure approaches, Google Cloud Platform (GCP) networking services, and how networking evolves in cloud environments. You may be asked about Software-Defined Networking (SDN), network automation, cloud-native networking architectures, and GCP services like Cloud VPN, Cloud Interconnect, Cloud Load Balancer, and VPC networking. The interviewer assesses your awareness of modern approaches and willingness to learn new technologies. This round is particularly important for Google roles since they heavily use cloud infrastructure internally.
Tips & Advice
Research GCP networking services thoroughly: understand Cloud VPC, subnets, firewall rules, Cloud Load Balancer, Cloud CDN, Cloud Interconnect, Cloud VPN, and Shared VPC. Understand the differences between GCP networking and traditional networking. Learn about SDN principles and how they differ from traditional hardware-based switching. Discuss automation: Infrastructure as Code (IaC), configuration management, and how you've used these in previous roles. Talk about containers and Kubernetes networking implications. If you have GCP experience, discuss real projects. If not, discuss how your networking knowledge applies to cloud platforms. Show understanding that cloud networking is more software-driven, more ephemeral, and requires different operational approaches. Discuss observability in cloud environments and differences from traditional network monitoring.
Focus Topics
Observability in Cloud and Modern Infrastructure
Understanding metrics, logs, and traces for cloud infrastructure. Using cloud-native observability platforms, understanding network flow data in cloud contexts, and monitoring microservices architectures
Practice Interview
Study Questions
Container and Kubernetes Networking
Understanding container networking models, overlay networks, Kubernetes networking concepts (Services, Ingress, NetworkPolicy), and how networking works differently in containerized environments
Practice Interview
Study Questions
Network Automation and Infrastructure as Code
Using Terraform, Ansible, or similar tools to define and manage network infrastructure, benefits of IaC (repeatability, version control, auditability), and integrating automation into deployment pipelines
Practice Interview
Study Questions
Software-Defined Networking (SDN) Concepts
Understanding the shift from hardware-based to software-defined networking, control plane vs data plane separation, network virtualization, and implications for network design and operations
Practice Interview
Study Questions
GCP Networking Services and Architecture
Understanding Google Cloud Platform networking fundamentals: VPC design, subnets, firewall rules, Cloud Load Balancer, Cloud CDN, Cloud Interconnect, Cloud VPN, and how networking is abstracted in cloud environments[4]
Practice Interview
Study Questions
Onsite Round 4: Google Behavioral and Culture Fit Interview
What to Expect
Interview focused on assessing your alignment with Google's culture and values. The interviewer will ask behavioral questions about your past experiences to evaluate: your ability to demonstrate leadership even without formal authority, how you handle ambiguity and complexity, your collaboration skills, your bias toward action, your communication ability, and your growth mindset. Unlike technical interviews, these questions assess soft skills, decision-making under uncertainty, and team dynamics. You should prepare specific examples from your past that illustrate key competencies. This round is critical for senior-level candidates as leadership and influence are core expectations[2].
Tips & Advice
Prepare 5-7 concrete examples from your career that demonstrate different competencies. Use the STAR method (Situation, Task, Action, Result) but focus on your specific contributions and how you influenced outcomes. For senior roles, emphasize: leading without formal authority, influencing team direction, mentoring others, handling ambiguity, making difficult trade-off decisions, and driving results in complex environments. Practice explaining failures and what you learned. Have examples showing collaboration across teams, driving process improvements, and contributing to team culture. Be authentic and specific—avoid generic corporate language. Show genuine enthusiasm for Google's mission and products. Ask thoughtful questions about the team and role. Research Google's leadership principles and weave them into your answers[2]. Discuss how you've operated at a higher level as you progressed in your career.
Focus Topics
Learning, Growth, and Adaptation
Examples of learning new technologies, adapting to changing business needs, seeking feedback, and demonstrating continuous growth throughout your career[2]
Practice Interview
Study Questions
Collaboration and Teamwork
Examples of working effectively with difficult stakeholders, collaborating across functional teams, resolving conflicts, and contributing to team culture[2]
Practice Interview
Study Questions
Technical Decision-Making and Trade-offs
Examples of making architectural or technical decisions that involved trade-offs, explaining your decision-making process, considering multiple perspectives, and defending your choices
Practice Interview
Study Questions
Bias to Action and Execution
Examples of moving quickly to deliver results, taking ownership of problems, driving projects to completion, and balancing speed with quality[2]
Practice Interview
Study Questions
Managing Ambiguity and Complexity
Examples of making decisions with incomplete information, operating in ambiguous situations, adapting to changing requirements, and helping teams navigate uncertainty[2]
Practice Interview
Study Questions
Emergent Leadership and Influence Without Authority
Demonstrating leadership when you don't have formal authority over others. Examples of stepping up to lead technical initiatives, mentoring team members, influencing decisions, and driving change through collaboration and technical credibility[2]
Practice Interview
Study Questions
Frequently Asked Network Engineer Interview Questions
Design and configure DHCP relay so that clients in VLAN 20 (subnet 10.20.20.0/24) receive DHCP from a centralized server at 10.0.0.10. Provide the switch SVI configuration, the router subinterface configuration for router-on-a-stick including ip helper-address, and explain the UDP/port and broadcast behaviors to check when clients receive no address.
Sample Answer
Approach (brief)
Create an SVI on the L2 switch for VLAN 20 so clients have gateway, trunk to router carrying VLAN 20, and configure router-on-a-stick subinterface with ip helper-address pointing to the DHCP server.
Switch SVI / VLAN config
interface Vlan20
description Clients-VLAN-20
ip address 10.20.20.1 255.255.255.0
no shutdown
interface GigabitEthernet1/0/1
switchport trunk encapsulation dot1q
switchport mode trunk
Router (router-on-a-stick) subinterface
interface GigabitEthernet0/0.20
encapsulation dot1Q 20
ip address 10.20.20.254 255.255.255.0
ip helper-address 10.0.0.10 ! forwards DHCP & other UDP services
no shutdown
Why ip helper-address / behavior to check when no lease
- DHCP uses UDP: client -> server uses source UDP 68 dest UDP 67; server replies src 67 dst 68.
- Router relay receives client broadcast (DHCPDISCOVER) and unicasts to 10.0.0.10 with giaddr = 10.20.20.254; server uses giaddr to allocate correct subnet.
- If clients get no address check:
- Is trunk up and VLAN 20 allowed between switch and router?
- Is SVI up and correct IP/mask?
- ACLs or firewall blocking UDP ports 67/68 on path to 10.0.0.10.
- Confirm helper configured on correct subinterface and that server responds to giaddr subnet.
- Use packet captures / debug:
- On router: debug ip dhcp server packets / debug ip packet detail to see forwarded UDP 67.
- On server: check received packet source and giaddr.
- Check ARP: router must ARP for client when replying (if relay, server reply returns to router which forwards to client).
- Optional: consider DHCP Relay Agent Information (option 82) if server expects it; can disable or configure appropriately.
Design an event-driven remediation pipeline that consumes streaming telemetry (gNMI/telemetry), publishes events to Kafka, runs anomaly detection, and triggers automated remediation playbooks (e.g., Ansible or Nornir) for incidents like interface flaps or BGP session drops. Describe components, message formats, orchestration logic, safety checks, and how to avoid repeated actions or race conditions.
Sample Answer
Overview
I’d build an event-driven remediation pipeline: telemetry → ingest → Kafka topics → stream processors (anomaly detection) → orchestration service → remediation runners (Ansible/Nornir) with safety and audit.
Components
- Telemetry collectors (gNMI collectors like OpenConfig/gNMI clients) pushing JSON/GPB to an ingest gateway.
- Kafka cluster with topics: telemetry.raw, telemetry.parsed, anomalies.detected, remediation.requests, remediation.results.
- Stream processing (Kafka Streams / Flink) for normalization, enrichment (device metadata), and anomaly detection rules/models.
- Orchestrator service (stateless microservice) that decides remediation, applies mutex/locks, and calls runners.
- Remediation runners (Ansible AWX or Nornir workers) invoked via REST or message queue.
- DB (Postgres/Redis) for state, locks, and dedup keys; and Prometheus/Grafana + audit logs.
Message formats
- Telemetry.parsed (JSON)
{ "device": "leaf1", "path": "interfaces/eth1/state/counters", "ts": 167..., "value": {...} } - anomalies.detected
{ "anomaly_id": "uuid", "type":"BGP_DOWN","device":"rtr1","ts":...,"evidence":[...], "confidence":0.92 } - remediation.request
{ "request_id":"uuid","anomaly_id":"uuid","action":"reload-bgp","device":"rtr1","params":{...},"ttl":300 }
Orchestration logic
- Stream processor emits anomaly event to anomalies.detected.
- Orchestrator consumes event, enrichment from CMDB, evaluates playbook mapping and risk policy.
- Acquire distributed lock in Redis: lock key = device + remediation_type. If locked, check lock TTL and skip/queue.
- Dedup: check Postgres for recent remediation for same anomaly fingerprint (hash of device+type+signature). If within suppression window, ack and skip.
- Apply safety checks: maintenance window, recent config change, manual override flag, device reachability test, impact blast-radius lookup.
- If safe, create remediation.request, persist state, publish to remediation.requests, call runner, and monitor result. Use exponential backoff for retries.
Safety & Race Conditions
- Distributed locking ensures single active remediation per device/type.
- Idempotent playbooks: design actions to be safe if re-run (check current state before change).
- Suppression windows and fingerprinting avoid repeated actions for flapping events.
- TTLs and heartbeats for locks to avoid deadlocks.
- Manual abort and approval gates for high-impact actions (notify via Slack/ITSM).
- Post-remediation verification step: validate the issue is resolved; if not, escalate to human.
Monitoring & Audit
- Emit remediation.results to Kafka; persist logs, success/failure metrics, and RCA artifacts.
- Alert on repeated failures or high-frequency remediations for operator intervention.
I’d implement incrementally: start with passive detection + alerting, then safe read-only remediation (e.g., clear counters), then escalate to config changes with approvals.
How do you handle receiving critical comments in a public setting, a Slack thread, a PR, or an all-hands, that question your competence or undermine your work in front of others? Walk through a specific time this happened: how you de-escalated in the moment, how you followed up privately to resolve the substance, and how you preserved the relationship afterward.
Sample Answer
Direct answer
In the moment, I de-escalate with a short, calm acknowledgment rather than defending point by point in public, then move the substance to a private conversation quickly. I resolve it there before returning anything final to the public thread, and I preserve the relationship by treating the private follow-up as genuine curiosity about what they saw, not damage control.
Structured elaboration
In the moment. A short, non-defensive acknowledgment that doesn't concede a point you haven't verified. Don't litigate the technical substance in the public thread; a back-and-forth there tends to escalate, since both people end up performing for an audience instead of solving the actual problem. This can be as simple as one line said in the room, or one line posted back in the thread: "good flag, let me look into that and follow up."
Follow up privately, quickly. Same day if possible, to actually understand the concern in full, away from the public framing, which is often more compressed or more pointed than intended.
Resolve the substance there, then close the loop publicly. A short, factual update, what was found and what changed, so the thread doesn't stay hanging as an unresolved public criticism. In a pull-request context, that update is often best as a concrete follow-up comment on the same thread, showing the fix rather than re-litigating the tone.
Preserve the relationship. Treat the private conversation as genuine curiosity about what they saw, not a negotiation to get them to back down.
Worked example
A senior engineer left a pull-request comment, visible to the whole team, saying my approach "showed a lack of understanding of how the service handles concurrency" and suggesting my change would cause a race condition (two operations happening at the same time in a way that produces the wrong result depending on which one finishes first). In the thread, I replied briefly: "good flag, let me dig into the concurrency behavior and get back to you," rather than apologizing profusely or arguing my design was fine. I then messaged them directly and asked them to walk me through the specific race they were worried about.
They were right about a real edge case under high load, though the public phrasing read as a broader judgment about my competence than the actual, narrower issue. We worked through the fix together on a quick call. I then posted a follow-up comment on the same pull request: "Confirmed the race condition under concurrent writes, thanks for catching it. Fixed by adding the lock (a mechanism that only lets one operation touch the shared data at a time) around the update in commit X, added a test for the concurrent case." That comment did double duty: it closed the technical thread with a concrete fix and quietly reset the public record from "you don't understand this" to "a specific bug was found and fixed," without relitigating the original tone.
Trade-offs and pitfalls
Defending yourself point by point in the same public thread usually escalates rather than resolves, since both people end up performing for the audience instead of solving the problem. Going silent in public after being publicly challenged can read as agreement or avoidance, so a brief acknowledgment before moving it private still matters. And posting only a public fix without ever actually talking to the person leaves the relationship unrepaired even if the code got fixed.
Tell me about a time you led troubleshooting of a major outage caused by a routing issue. Use the STAR method: describe the Situation, the Task you owned, the Actions you took (technical commands, configuration changes, and coordination), and the Result. Explain what you changed afterward (process, config, monitoring) to prevent recurrence.
Sample Answer
Situation
During a peak business day our primary region experienced 40% packet loss and service degradation. Customers reported timeouts; internal dashboards showed large AS-path changes and reachability loss to several subnets.
Task
I led the incident as the on-call Network Engineer: identify root cause, restore full connectivity, coordinate cross-team actions, and own the post-incident remediation.
Actions
- Immediate diagnosis: logged into core routers and ran:
show ip route 10.0.0.0/16
show bgp summary
show bgp neighbors <peer-ip> received-routes
traceroute <affected-prefix>
- Found an incorrect route leak from a newly onboarded transit peer that injected more-specific routes; BGP local-pref for leaked routes caused traffic black-holing.
- Tactical mitigation:
- Applied a route-map to deny the leaked prefixes from that peer:
route-map BLOCK-LEAK deny 10
match ip address prefix-list LEAKED
!
ip prefix-list LEAKED seq 5 deny 10.0.0.0/8 ge 16 le 32
- Reset BGP session to relearn correct routes:
clear ip bgp <peer-ip> softthenclear ip bgp <peer-ip>. - Coordination: informed NOC, cloud peering contact, security and app owners via conference bridge; synchronized change window for immediate route-map push; kept stakeholders updated every 5–10 minutes.
Result
Within 6 minutes of applying the route filter and soft-resetting BGP, full reachability and normal latency returned. SLA impact minimized; no data loss. Delivered a 2-hour postmortem within 24 hours.
Post-incident changes
- Config: enforced strict inbound prefix-lists on all external peers, set max-prefix limits and BGP dampening for noisy peers, and default local-pref for customer routes.
- Process: added mandatory validation checklist for onboarding peers (proof of route filters, test sessions), and dual-approval for peer config pushes.
- Monitoring: implemented BGP route-advertisement alerts, automated detection for sudden AS-path changes and more-specific prefix spikes; added synthetic probes for critical prefixes and an incident runbook.
- Learned to require staged rollouts and automated config checks to prevent future leaks.
Explain how NAT gateway or firewall connection-tracking table exhaustion can cause intermittent connection failures at scale. Describe how you would detect conntrack exhaustion, what metrics to monitor, and how you'd decide between raising the table's capacity versus addressing what's actually driving the churn. Include example commands to view the conntrack table on Linux-based appliances.
Sample Answer
Direct answer
A NAT gateway or firewall's connection-tracking (conntrack) table has a finite size; once it fills up, new connections can't be tracked and are dropped or fail intermittently, and this failure mode is distinctive because it appears only under LOAD (a specific connection count or churn rate), not as a constant, deterministic failure.
Structured elaboration
- Understand what conntrack tracks and why it can exhaust: every distinct connection passing through a stateful NAT or firewall device consumes an entry in its connection-tracking table until that connection closes (or its entry times out); a very large number of concurrent connections, or a high RATE of short-lived connections that don't clean up entries fast enough, can exhaust the table's capacity even if the device's raw throughput is nowhere near its limit.
- Detect conntrack exhaustion specifically: on Linux-based appliances,
conntrack -C(or reading/proc/sys/net/netfilter/nf_conntrack_countagainstnf_conntrack_max) shows current usage against the configured maximum; a count consistently near or at the maximum, correlated with the timing of reported failures, is the direct signature. - Monitor the right metric proactively: track conntrack table utilization as a percentage of its maximum, not just raw connection counts, since the maximum itself may have been sized for a different traffic pattern than what's actually occurring now.
- Remediation options, each with real trade-offs: increasing
nf_conntrack_maxraises the ceiling directly but consumes more memory and doesn't address WHY the table is filling (a genuine traffic-pattern change, or a bug causing connections to accumulate without properly closing); tuning conntrack timeouts more aggressively (reclaiming entries for idle or half-closed connections sooner) helps if many entries are lingering longer than they need to; deploying additional NAT gateways to spread the load reduces per-device table pressure but adds complexity and cost; and stateless NAT (where architecturally feasible) avoids the conntrack limit entirely for the traffic it handles, at the cost of losing some of the visibility and flexibility a stateful table provides. - Consider port-preservation strategies: some NAT implementations can be tuned to reuse or preserve source ports more efficiently, reducing table churn for high-connection-rate workloads without simply raising the ceiling.
Worked example
conntrack -C shows a count sitting at 98% of nf_conntrack_max during the exact windows when clients report intermittent connection failures, and comfortably lower outside those windows. Correlating this against traffic patterns shows a recent change introduced a much higher rate of short-lived connections (a client library switched from connection pooling to opening a fresh connection per request), each consuming a conntrack entry briefly but at a volume the existing table sizing didn't anticipate. Raising nf_conntrack_max provides immediate relief, while the client library's connection-pooling regression is flagged as the actual root cause worth fixing at the source.
Trade-offs & pitfalls
Simply raising the conntrack table's maximum size treats the symptom, not necessarily the cause; if a client-side regression (like the pooling example above) is generating far more connections than intended, raising the ceiling delays the next exhaustion event rather than preventing it. Always correlate exhaustion against what changed in the actual traffic pattern, not just the device's own capacity, before deciding whether to fix the ceiling or the traffic generating the load.
Starting from 192.168.100.0/24, allocate VLSM subnets to support networks requiring 30, 14, 6, and 2 hosts respectively. For each subnet provide the CIDR, network and broadcast addresses, and usable host ranges. Explain your allocation order and justify minimal waste.
Sample Answer
Approach & justification
Allocate largest subnets first (largest-to-smallest) to minimize fragmentation and waste. Required host counts map exactly to standard CIDR sizes: 30 → /27 (30 usable), 14 → /28 (14 usable), 6 → /29 (6 usable), 2 → /30 (2 usable).
Allocations (from 192.168.100.0/24)
-
30 hosts — /27
- Network: 192.168.100.0/27
- Usable: 192.168.100.1 – 192.168.100.30
- Broadcast: 192.168.100.31
-
14 hosts — /28
- Network: 192.168.100.32/28
- Usable: 192.168.100.33 – 192.168.100.46
- Broadcast: 192.168.100.47
-
6 hosts — /29
- Network: 192.168.100.48/29
- Usable: 192.168.100.49 – 192.168.100.54
- Broadcast: 192.168.100.55
-
2 hosts — /30
- Network: 192.168.100.56/30
- Usable: 192.168.100.57 – 192.168.100.58
- Broadcast: 192.168.100.59
Why minimal waste
Each subnet chosen exactly matches the nearest power-of-two size that accommodates required hosts (including network/broadcast). Allocating largest-first ensures contiguous free space for smaller blocks and avoids gaps that would prevent fitting larger subnets later. Remaining space (192.168.100.60–192.168.100.255) is available for future VLSM allocations.
What is a runbook, and what does a good one actually need to contain to be useful when someone's paged at 3am? Sketch what you'd want in one for a failed database migration.
Sample Answer
Direct answer
A runbook is a step-by-step operational document for handling a specific, known failure mode: what to check first, what commands to run, when to roll back versus push forward, and who to call if it gets worse. A good one is written so that someone half-awake at 3am who has never touched this exact system before can follow it without reconstructing context from scratch. The test of a good runbook is whether a different engineer than its author can execute it correctly under pressure.
What a runbook needs to contain
- Scope and severity: what specific failure this covers, and the severity/priority it corresponds to.
- Preconditions: what access, credentials, or tools you need before starting.
- Immediate triage steps: the first few things to check, in order, to confirm the diagnosis.
- Remediation steps: copy-pasteable commands with the expected output at each step, not prose descriptions of what to do.
- Verification: how to confirm the fix actually worked, not just that the command ran.
- Rollback path: a safe way back if remediation makes things worse, including its own preconditions (e.g. "requires a backup from the last 24 hours").
- Escalation: who to page next and when, by name/role/contact, not just "escalate if needed."
- Post-incident: where to file the incident ticket, and a note to update the runbook itself if a step was wrong or missing.
Worked example: failed database migration runbook
- Scope: prod schema migration failed mid-deploy. Severity: P1 if writes are blocked, P2 if only the migration job failed cleanly.
- Triage (first 5 minutes): tail the migration tool's log for the exact error; check
SELECT count(1) FROM pg_stat_activity WHERE state <> 'idle';to see if the migration left long-running locks; check the app's error dashboard for whether requests are actually failing yet. - Remediation, case A (migration failed cleanly, nothing partially applied): re-run the migration tool in dry-run mode first, then apply.
- Remediation, case B (partially applied, schema now inconsistent): put the app in read-only/maintenance mode to stop new writes, then decide between manually completing the migration or rolling back.
- Rollback (case B, if completing isn't safe): confirm the most recent backup timestamp, pause replication, restore with
pg_restore --clean --no-owner <backup>, then run a smoke test against a few critical read/write paths before removing maintenance mode. - Verification: re-run the app's smoke tests, spot-check row counts on affected tables against the pre-migration baseline.
- Escalation: if triage doesn't identify the cause within 10 minutes, or rollback is being considered, page the on-call DBA by name/rotation, not just "the DBA team."
- Post-incident: file the incident ticket with the timeline and root cause, and if any step above was missing or wrong, fix the runbook in the same pass as the incident writeup, not "later."
Trade-offs and pitfalls
- A runbook that's too generic ("check the logs, investigate, fix it") isn't actually a runbook, it's a checklist item pretending to be one; specificity is what makes it useful at 3am when judgment is impaired by fatigue.
- Runbooks rot: a step that references a tool or dashboard that got replaced months ago is worse than no runbook, because it wastes time and erodes trust in the whole document. Tie runbook review to any change in the system it covers, not a fixed calendar cadence alone.
- Over-indexing on "never improvise" can be as dangerous as no runbook at all: a good runbook documents when to deviate (e.g. "if replication lag exceeds a set threshold, stop and escalate instead of continuing") rather than pretending every failure mode was anticipated.
- Untested runbooks are a liability. The rollback path above should actually be exercised in a drill, not just written down, since commands like
pg_restore --cleanbehave differently depending on schema ownership and extensions that may not match what was true when the runbook was written.
Describe Kubernetes Service types (ClusterIP, NodePort, LoadBalancer, ExternalName). Explain how each type routes traffic to pods, typical cloud provider integrations, operational limitations, and how kube-proxy implementation affects packet forwarding behavior.
Sample Answer
A Kubernetes Service is a stable virtual address in front of a changing set of pods. The four types differ in who can reach that address and how: ClusterIP is internal-only, NodePort adds a port opened on every node, LoadBalancer asks the cloud provider to provision an external load balancer in front of that, and ExternalName is different in kind: a DNS alias to something outside the cluster with no pod routing involved at all.
Service types
| Type | Reachable from | How traffic reaches pods | Typical cloud integration |
|---|---|---|---|
| ClusterIP | inside the cluster only | kube-proxy rewrites the Service's virtual IP to a pod IP | none needed |
| NodePort | any node's IP, on a fixed port in the 30000-32767 range | same in-cluster rewrite as ClusterIP, reached via the node's IP first | often sits behind a hand-rolled external load balancer |
| LoadBalancer | the internet, or a private network, via a provisioned load balancer's address | the provider's load balancer forwards to node ports, which kube-proxy then routes to pods | automatic provisioning on AWS, GCP, or Azure through the cloud-controller-manager |
| ExternalName | resolves to an external DNS name; no cluster-internal routing at all | none; CoreDNS returns a CNAME and the client connects directly | giving a managed external service (e.g., a managed database) a cluster-local name |
A fifth variant sits alongside these four rather than replacing any of them: a headless Service (clusterIP: None). It skips kube-proxy's virtual-IP rewriting entirely, so instead of one stable IP load-balancing across pods, CoreDNS returns the individual pod IPs behind the Service directly to the client (sourced from the same EndpointSlices kube-proxy would otherwise consume). This is why StatefulSets pair with a headless Service: each replica needs its own stable, individually addressable DNS name rather than one shared address in front of all of them.
kube-proxy: the mechanism underneath
kube-proxy is what actually implements the ClusterIP, NodePort, and LoadBalancer rewriting, and it has more than one implementation:
- iptables mode (the long-standing default): programs Linux netfilter rules to destination-NAT (DNAT, rewriting a packet's destination address) Service traffic to a pod IP. Simple, but rule evaluation is roughly linear in Service count, which shows up as latency at high Service counts.
- ipvs mode (IP Virtual Server): uses the kernel's IPVS load-balancing tables instead of a rule chain, giving better performance and more scheduling algorithm choices at scale.
- nftables mode: the newer replacement backend, built on the modern Linux nftables framework, which reached general availability in Kubernetes 1.33. It targets the same correctness as iptables mode with materially better performance at high Service counts, though iptables remains the cluster-wide default for now.
Older material sometimes still mentions a userspace mode; it was deprecated years ago and fully removed in Kubernetes 1.26, so it should not appear in any current design.
Worked example: tracing one request through a LoadBalancer Service
flowchart LR
Client --> ELB[Cloud load balancer]
ELB --> NP[Node: NodePort]
NP --> KP[kube-proxy DNAT rule]
KP --> PodA[Pod A]
KP --> PodB[Pod B]
The client hits the cloud load balancer's public address, which forwards to a NodePort on whichever node it picked, where kube-proxy's DNAT rule sends the packet to one of the pods listed in the Service's EndpointSlices. With externalTrafficPolicy left at its default Cluster, the node that first receives the packet can forward it on to a pod on a different node, an extra hop, but with even load spread across all pods. Setting it to Local skips that extra hop and preserves the client's real source IP, at the cost of only routing to pods that happen to already be on the node the load balancer chose, so an uneven pod distribution across nodes turns directly into uneven traffic.
Trade-offs and pitfalls
- ExternalName is often reached for as 'the DNS one' when someone actually wants a placeholder Service for pods that do not exist yet. Because it has no selector and no health checking, nothing about it verifies the external target is even up.
- NodePort's fixed 30000-32767 range and one-port-per-Service-per-node model does not scale past a modest number of exposed Services and puts every node's IP in the attack surface; it is a building block for an external load balancer, rarely the final production answer on its own.
- LoadBalancer provisions one cloud load balancer per Service by default, which gets expensive and slow to provision as the number of Services grows; that pressure is usually solved by putting an Ingress controller behind a single LoadBalancer Service, a platform decision distinct from Service type selection itself.
- kube-proxy's mode is a cluster-wide setting, not a per-Service choice; moving from iptables to ipvs or nftables needs a validated cluster-wide rollout plan, not a Service-by-service change.
You are asked to integrate eBPF-based telemetry (e.g., Cilium) into an existing observability stack that uses Prometheus (metrics) and ELK (logs). Describe the types of telemetry eBPF can produce, how to map those outputs to metrics/logs/traces, export mechanisms and performance considerations, and strategies to reduce high-cardinality noise from eBPF events.
Sample Answer
High-level summary
eBPF (via Cilium) produces packet/flow, socket, syscall, L7, and kprobe/uprobe events. For a network engineer, map those outputs to metrics (Prometheus), logs (ELK), and traces (OpenTelemetry) while controlling performance and cardinality.
Types of telemetry & mapping
- Flow/Conntrack events (connect/accept/close, 5-tuple, bytes, RTT) → Prometheus metrics: connection rates, error rates, P50/P95 latency histograms; aggregated by service/namespace.
- L7/App-level observability (HTTP method, path, status) → traces/spans (OpenTelemetry) for request flow; or ELK logs when full request context required.
- DNS and TLS events (queries, SNI) → logs in ELK for security forensics; metrics for query volume and error rates.
- Socket/syscall traces and drops → Prometheus counters + ELK logs for detailed forensic records.
- Kernel/kprobe events (drops, retransmits) → metrics (alerts) + sampled logs.
Export mechanisms
- Prometheus: use Cilium’s metrics endpoint or an exporter that converts eBPF map counters → /metrics. Aggregate in user-space before exposing (avoid raw cardinality).
- Logs: send detailed event streams from Cilium/Hubble Relay to Fluentd/Fluent Bit → ELK. Use structured JSON with fields for indexing.
- Traces: emit OpenTelemetry spans from Cilium/Hubble or sidecar that correlates flows to traces; sample high-volume flows.
- Alternatives: push via Kafka for downstream processing, or use Hubble UI/Relay for interactive exploration.
Performance considerations
- eBPF overhead scales with event rate and map ops. Reduce per-packet work in eBPF program; do aggregation in user-space.
- Use ring buffers/perf buffers with batch reads. Tune batch sizes and reader frequency.
- Keep eBPF programs simple and bounded (bounded loops, limited maps).
- Kernel & BPF verifier: prefer BPF CO-RE and recent kernels to reduce complexity.
- Rate-limit exports and apply backpressure (drop non-essential events when CPU/queue pressure evident).
Cardinality reduction strategies
- Whitelist important labels (service, namespace, status), drop high-cardinality ones (client IPs) from metrics — keep them only in sampled logs.
- Pre-aggregate in kernel or user-space: counts per service rather than per source IP.
- Use sampling (systematic or adaptive): full logs for 1% of connections, metrics aggregated for all.
- Bucketing/histograms: use fixed buckets for latency/size to avoid unique value explosion.
- Dynamic rollups: retain detailed telemetry for short TTL then roll up to coarser aggregates.
- Enforce cardinality quotas in exporters and tag cardinality monitoring.
Example approach
- Emit aggregated Prometheus metrics from Cilium Hubble (conn/sec, drop/sec, P95 latency).
- Stream detailed JSON events to Fluent Bit with sampling rules (e.g., 1:1000 for client IPs) and forward to ELK only when matched by anomaly rules.
- Send sampled spans via OpenTelemetry collector to tracing backend for top-N slow endpoints.
This balances visibility, performance, and manageable cardinality while enabling security and network troubleshooting.
Describe the TCP three-way handshake in detail: which flags are set in each packet (SYN, SYN-ACK, ACK), how sequence and acknowledgment numbers are used, and what state each endpoint moves into after each step. Explain what problem the handshake actually solves.
Sample Answer
Direct answer
The TCP three-way handshake establishes a reliable connection before any data flows: the client sends a SYN, the server replies with a combined SYN-ACK, and the client finishes with an ACK. Its job is to let both sides agree on starting sequence numbers and confirm that both directions of the path actually work before committing application data to the wire.
Structured elaboration
- SYN: the client picks an initial sequence number (ISN, essentially a large pseudo-random 32-bit number) and sends a segment with the SYN flag set and that sequence number. The client moves to the
SYN-SENTstate. - SYN-ACK: the server, if it's listening on that port, picks its OWN initial sequence number, and replies with a segment that has both the SYN flag set (announcing the server's own sequence number) AND the ACK flag set (acknowledging the client's sequence number + 1). The server moves to the
SYN-RECEIVEDstate. - ACK: the client acknowledges the server's sequence number + 1 with a plain ACK segment. Both sides now move to
ESTABLISHED, and either side may now send data.
Why three steps rather than two: TCP needs BOTH sides' sequence numbers acknowledged, since TCP is full-duplex (both directions need independent sequence tracking). A two-way handshake could confirm only one direction; the third message is what confirms the client's original SYN actually arrived, closing the loop for the client's own sequence space.
Worked example
Suppose a client opens a TCP connection to a web server on port 443. The client sends SYN, seq=1000. The server responds SYN, ACK, seq=5000, ack=1001 (acknowledging the client's SYN by number+1). The client responds ACK, seq=1001, ack=5001. From this point, the client's next data byte will carry sequence number 1001, and the server's next data byte will carry sequence number 5001; each side tracks the OTHER side's sequence space independently via the ACK field of every following segment.
Trade-offs & pitfalls
A frequent mistake is describing the handshake as three round trips; it's actually one and a half round trips of latency, because the SYN-ACK piggybacks the server's SYN onto its ACK of the client's SYN. This is also exactly why TCP always incurs at least one round trip of setup latency before any data can flow, which is the whole motivation behind newer mechanisms like TCP Fast Open that try to send data alongside the very first SYN.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Network Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs