Google Staff Network Engineer Interview Preparation Guide
Google's Staff Network Engineer interview process typically consists of an initial recruiter screening, 2 technical phone rounds focusing on networking fundamentals and system design, followed by 5-6 onsite rounds covering system architecture, troubleshooting, security, behavioral assessment, and cross-functional collaboration. The process emphasizes deep technical expertise, strategic thinking, leadership capability, and alignment with Google's engineering culture. For Staff level, expect evaluation of your ability to influence beyond your immediate team and drive complex infrastructure initiatives.
Interview Rounds
Recruiter Screening
What to Expect
Initial 30-minute call with Google recruiter to assess background fit, career motivations, compensation expectations, and availability. Recruiter will verify your experience aligns with Staff-level expectations (12+ years in networking/infrastructure). May include a brief follow-up call after initial technical screens.
Tips & Advice
Have a concise 2-minute pitch about your background, emphasizing leadership experience and complex projects you've led. Be clear about why you're interested in Google specifically and the Staff role. Ask about team structure, immediate challenges, and career growth. Highlight cross-functional collaboration and mentorship experience.
Focus Topics
Relevant Project Experience
Summarize 2-3 significant network architecture or infrastructure projects demonstrating complexity, scope, and your technical depth.
Practice Interview
Study Questions
Leadership and Scope of Impact
Describe teams led (directly or indirectly), infrastructure initiatives you've driven, and how you've influenced beyond your immediate scope.
Practice Interview
Study Questions
Career Narrative and Motivation
Articulate your 12+ year journey in networking, key career inflection points, and why you're pursuing this Staff role at Google now.
Practice Interview
Study Questions
Technical Phone Screen 1: Networking Fundamentals and Troubleshooting
What to Expect
60-minute technical phone interview with a Google engineer focusing on core networking knowledge, hands-on troubleshooting skills, and diagnostics. Expect scenario-based questions where you explain how to diagnose and resolve network issues. You may be asked to sketch network topologies or discuss configuration approaches verbally. This round validates that you have strong fundamentals despite your seniority.
Tips & Advice
Review networking fundamentals: TCP/IP stack, DNS resolution, routing protocols (BGP, OSPF), switching concepts, VLAN management, and packet flow. Study common troubleshooting scenarios (connectivity issues, packet fragmentation, port reachability, DNS failures, selective routing problems, VLAN routing). Use the search results' troubleshooting methodology: trace through layers systematically (physical → link → network → application). Explain your thinking process clearly. Reference diagnostic tools like ping, traceroute, dig, getent, ip route, ss, netstat, tcpdump. At Staff level, you should explain not just 'how to fix' but 'why' issues occur and architectural trade-offs.
Focus Topics
Routing Protocols and Path Selection
Static vs. dynamic routing, BGP fundamentals, IGP protocols (OSPF, ISIS), metric calculation, convergence, and policy-based routing.
Practice Interview
Study Questions
Switching and VLAN Architecture
VLAN design, trunk vs. access ports, inter-VLAN routing configuration, spanning tree protocol, MAC learning, and layer 2 switching behavior.
Practice Interview
Study Questions
Diagnostic Tools and Network Analysis
Practical use of ping, traceroute, dig, getent, ip route, ss, netstat, tcpdump, netcat, and curl for network diagnostics and packet analysis.
Practice Interview
Study Questions
TCP/IP Stack and Protocol Fundamentals
Deep understanding of OSI model layers, TCP/IP protocols, IPv4/IPv6 addressing, subnetting, routing, ARP, ICMP, DNS, and protocol interactions.
Practice Interview
Study Questions
Network Troubleshooting Methodology
Systematic approach to diagnosing connectivity issues: layered diagnostics (physical → link → network → application), MTU problems, packet fragmentation, firewall/NAT issues, DNS resolution failures, and routing anomalies.
Practice Interview
Study Questions
Technical Phone Screen 2: Advanced Architecture and Design Thinking
What to Expect
60-minute technical phone interview focusing on system design, architectural thinking, and handling complex infrastructure problems. Expect open-ended design questions where you propose network architectures for scenarios (e.g., 'Design a global load balancing strategy for multi-region traffic,' 'Design a network security posture for microservices'). Interviewer will probe your trade-offs, scalability considerations, and how you'd handle failures. This round assesses your ability to think strategically, not just execute tactically.
Tips & Advice
Practice design thinking: clarify requirements, discuss trade-offs (cost vs. performance vs. security), propose solutions, and iterate based on feedback. Structure your answers: problem statement → constraints → proposed architecture → trade-offs → monitoring/observability. Be comfortable discussing multiple solutions (active-active vs. active-passive, overlay vs. underlay networks, centralized vs. distributed control). Reference Google's published infrastructure (Spanner, Colossus, Maglev if applicable to roles you're familiar with). At Staff level, interviewers expect you to bring up considerations that junior engineers miss: automation, monitoring, rollback strategies, blast radius limitations. Discuss how you'd mentor team members through these decisions.
Focus Topics
Cloud and Hybrid Infrastructure Considerations
Networking in cloud environments (VPCs, subnets, NAT, VPN tunnels), hybrid setups (on-prem to cloud), inter-cloud connectivity, and vendor lock-in considerations.
Practice Interview
Study Questions
Reliability, Observability, and Operational Excellence
Designing for reliability (SLO/SLA definition, failure scenarios, graceful degradation), observability (monitoring, alerting, logging), and automation for operational simplicity.
Practice Interview
Study Questions
Network Security Architecture
Designing secure network architectures: microsegmentation, firewall strategies, DDoS mitigation, encryption in transit, VPNs, identity/access control, and security boundaries.
Practice Interview
Study Questions
Large-Scale Network Architecture Design
Designing networks for scale (multi-region, multi-cloud, high availability): load balancing strategies, traffic engineering, redundancy patterns, and handling millions of requests.
Practice Interview
Study Questions
System Design Trade-offs and Decision Making
Evaluating trade-offs: latency vs. throughput, consistency vs. availability, cost vs. complexity, centralized vs. distributed control, manual vs. automated operations.
Practice Interview
Study Questions
Onsite Round 1: Network System Design Deep Dive
What to Expect
90-minute onsite whiteboarding session focused on complex network architecture design. You'll receive a scenario (e.g., 'Design a global CDN network for video delivery,' 'Design network infrastructure for a company scaling from 1M to 100M users') and asked to architect a solution. Interviewer will ask probing questions about scalability, failure modes, monitoring, and trade-offs. You'll be expected to draw diagrams, discuss components, and think through operational aspects. Expect questions on why you made certain choices and how you'd validate your design decisions.
Tips & Advice
Start by clarifying requirements and constraints (geography, scale, latency requirements, failure tolerance, cost). Sketch architecture: draw network topologies, identify critical paths, and mark failure points. Use clear notation and explain each component's role. Discuss redundancy and failover mechanisms. Address monitoring and observability early. Be prepared to pivot based on interviewer feedback. For Staff level, explicitly discuss: automation and self-healing, blast radius limitations, how you'd rollout changes, and team scaling (how many engineers to operate this?). Walk through a failure scenario and explain how your design handles it.
Focus Topics
Network Automation and Orchestration
Infrastructure-as-code principles, automated provisioning, configuration management, self-healing networks, and runbook automation for common operations.
Practice Interview
Study Questions
Capacity Planning and Growth Scaling
Forecasting infrastructure needs, non-disruptive upgrades, scaling strategies (horizontal vs. vertical), and cost optimization for growing deployments.
Practice Interview
Study Questions
Failure Analysis and Resilience Design
Identifying failure modes (link failure, switch failure, regional outage, software bug), designing graceful degradation, fast failover mechanisms, and testing strategies.
Practice Interview
Study Questions
Global Network Architecture at Scale
Multi-region network design, global load balancing, traffic engineering for optimal routing, geo-redundancy, and handling asymmetric network conditions.
Practice Interview
Study Questions
Onsite Round 2: Network Troubleshooting and Incident Response
What to Expect
75-minute technical interview simulating real-world troubleshooting scenarios. You'll be given a scenario (e.g., 'Users report 500ms latency spike affecting east region traffic; walk me through your diagnostics') and asked to systematically diagnose and resolve the issue. Interviewer plays devil's advocate, adding complexity (cascading failures, conflicting signals). You'll whiteboard your diagnostic approach, explain tools you'd use, and discuss how you'd communicate with stakeholders during an incident.
Tips & Advice
Follow a structured troubleshooting methodology: define scope (affects what? when did it start?), narrow down layers (physical → link → network → application), gather data (logs, metrics, packet captures), form hypotheses, and test. Use search results methodology: if IP ping works, routing and interface are fine; if domain lookup fails, DNS is suspect. At Staff level, show leadership in incidents: how you'd delegate, communicate impact, and prevent recurrence. Discuss coordination with other teams (application teams, security, DBAs). Mention post-incident review and documentation.
Focus Topics
Packet-Level Diagnostics and Protocol Analysis
Using tcpdump, Wireshark, or similar tools to capture and analyze packets; understanding protocol sequences, identifying anomalies, and correlating network behavior with application issues.
Practice Interview
Study Questions
Network Performance Analysis and Optimization
Identifying performance bottlenecks (latency, throughput, jitter), understanding packet loss, analyzing traffic patterns, and implementing optimizations (QoS, traffic engineering, compression).
Practice Interview
Study Questions
Incident Response and Crisis Management
Escalation procedures, stakeholder communication, impact assessment, mitigation steps, coordination with multiple teams, and post-incident reviews.
Practice Interview
Study Questions
Systematic Troubleshooting Methodology
Layered diagnostic approach (physical → link → network → application), hypothesis formation, data gathering, and iterative narrowing of root cause.
Practice Interview
Study Questions
Onsite Round 3: Network Security and Risk Management
What to Expect
75-minute interview focused on network security architecture, threat analysis, and risk mitigation. You'll be asked questions like 'How would you secure communication between microservices?' or 'Your public API is under DDoS attack; how do you respond?' Interviewer expects you to discuss security principles, threat modeling, defense-in-depth strategies, compliance considerations, and how you'd balance security with operational efficiency. At Staff level, expect questions about security culture and mentoring junior engineers on security practices.
Tips & Advice
Ground your answers in threat models: identify assets, threat actors, attack vectors, and impacts. Discuss defense-in-depth (multiple layers). For DDoS scenarios, explain upstream mitigation (BGP flowspec, anycast), rate limiting, and traffic scrubbing. For microservices, mention mTLS, zero-trust networking, and service meshes. Reference compliance frameworks (PCI-DSS, HIPAA, SOC2) if relevant. At Staff level, discuss how you'd lead security initiatives, build threat modeling into architecture reviews, and foster security awareness in team. Mention balancing security with developer experience and operational complexity.
Focus Topics
Compliance and Governance in Network Design
Understanding compliance frameworks (PCI-DSS, HIPAA, SOC2), data residency requirements, audit trails, and how to design compliant infrastructure.
Practice Interview
Study Questions
DDoS Mitigation Strategies
Understanding DDoS attacks (volumetric, protocol, application-layer), upstream mitigation (BGP flowspec, anycast), rate limiting, traffic scrubbing, and capacity planning for attack resilience.
Practice Interview
Study Questions
Threat Modeling and Risk Assessment
Identifying threat actors, analyzing attack vectors, assessing impact, and prioritizing security controls based on risk.
Practice Interview
Study Questions
Network Security Architecture and Defense-in-Depth
Designing layered security: perimeter defense (firewalls, WAF), internal segmentation (microsegmentation, VLANs), encryption (mTLS, VPN), and identity/access controls.
Practice Interview
Study Questions
Onsite Round 4: Behavioral and Leadership
What to Expect
60-minute behavioral interview assessing leadership, collaboration, conflict resolution, and alignment with Google values. You'll be asked about challenging situations you've navigated (difficult team members, conflicting priorities, failed projects, change management). Interviewer will probe how you led without formal authority, mentored others, and handled ambiguity. Based on Google's known behavioral interview patterns, expect questions about teamwork, leadership, project ownership, and decision-making under pressure.
Tips & Advice
Prepare 6-8 concrete stories using the STAR method (Situation, Task, Action, Result), emphasizing your impact. For Staff level, focus on: leading cross-functional initiatives, mentoring engineers, driving architectural decisions that influenced others, handling dissent constructively, and driving culture. Reference Google's known values if applicable (focus on users, bias for action, collaboration). Discuss team growth: how you've developed junior engineers and created opportunities for them. Mention conflict resolution: disagreements with security, application teams, or other infrastructure engineers. Show humility: discuss what you learned from failures and how you'd do things differently.
Focus Topics
Handling Failure and Driving Improvement
Examples of significant failures, what you learned, how you led recovery, and how you prevented recurrence through systemic improvements.
Practice Interview
Study Questions
Cross-Functional Collaboration and Stakeholder Management
Working effectively with application teams, security, SREs, product managers; managing competing priorities and building alignment.
Practice Interview
Study Questions
Leadership and Influence Without Authority
Demonstrating leadership when not formally managing: driving architectural decisions, building consensus across teams, influencing others through expertise and credibility.
Practice Interview
Study Questions
Mentorship and Team Development
Examples of developing junior engineers, creating learning opportunities, providing feedback, and building a culture of growth and excellence.
Practice Interview
Study Questions
Onsite Round 5: Strategic Thinking and Cross-Team Impact
What to Expect
60-minute interview with a senior leader (potentially manager's manager level) assessing strategic thinking, business acumen, and organization impact. Expect questions like 'How would you prioritize between three competing infrastructure initiatives?' or 'How would you build a business case for a major network infrastructure project?' You'll discuss how you balance technical excellence with business realities, how you'd scale your impact beyond hands-on work, and how you think about multi-year infrastructure strategy. This round evaluates fit for Staff-level scope: influence, judgment, and business alignment.
Tips & Advice
Think like a business partner, not just an engineer. Prepare to discuss projects through a business lens: costs, benefits, risk, ROI, timeline. Have examples of trade-off decisions you've made: why you chose solution A over B, what you'd do differently. Discuss how you've communicated complex technical topics to non-technical stakeholders. Share examples of influencing strategy: maybe you advocated for a migration, drove a cost optimization initiative, or built a roadmap. Be prepared to discuss your vision for the team's role and how you'd contribute to it at Google. Show comfort with ambiguity and ability to drive clarity. Ask thoughtful questions about business priorities and team challenges.
Focus Topics
Organizational Impact and Scaling Influence
Examples of impact beyond direct projects: shaping team culture, driving organizational changes, establishing practices or standards, and multiplying impact through others.
Practice Interview
Study Questions
Communication with Executive Stakeholders
Distilling complex technical topics for non-technical audiences, presenting business cases, managing expectations, and driving visibility for critical initiatives.
Practice Interview
Study Questions
Business Alignment and Technical-Business Trade-offs
Understanding business impact of technical decisions, quantifying trade-offs (cost vs. latency, complexity vs. simplicity), and making decisions aligned with business goals.
Practice Interview
Study Questions
Strategic Roadmap Planning and Prioritization
Multi-year infrastructure strategy, balancing innovation with operational needs, resource allocation, and driving organizational alignment on priorities.
Practice Interview
Study Questions
Frequently Asked Network Engineer Interview Questions
A replication service running over a high-latency WAN link achieves only 20% of the link's theoretical throughput. Walk through the end-to-end set of transport-layer explanations you would check, in a sensible order: congestion-control algorithm choice, socket buffer sizing, TCP window scaling and MSS, and NIC offload settings. For each, explain what evidence would tell you it is (or isn't) the cause.
Sample Answer
Direct answer
For a WAN replication job stuck at 20% of theoretical throughput, walk the transport-layer stack in this order: congestion-control behavior first (is loss even happening, and if so is the algorithm reacting sensibly), then window sizing relative to the path's bandwidth-delay product, then socket buffers, then NIC-level settings, since each of these can independently cap throughput and the cheapest checks come first.
Structured elaboration
- Congestion-control algorithm and loss: check whether the connection is experiencing any packet loss at all (via retransmit counters,
ss -i'sretransfield). If there's meaningful loss and the algorithm is a loss-based one like Cubic, and the path has ANY non-congestive baseline loss (common on long WAN paths), the algorithm may be needlessly throttling itself; switching to BBR is a plausible fix here specifically because it doesn't over-react to non-congestive loss the way Cubic does. - Window size versus bandwidth-delay product: compute the path's bandwidth-delay product (bandwidth times round-trip time) and compare it to the connection's actual window (
ss -i'scwndand the negotiated window scale). If the window can never grow large enough to cover the bandwidth-delay product, no amount of congestion-control tuning will help, the connection is fundamentally window-limited, not congestion-limited. - Socket buffer sizing: even with window scaling negotiated, if the OS's actual send/receive socket buffers (
net.ipv4.tcp_wmem/tcp_rmemon Linux) are capped below what the window scale option would otherwise allow, the effective window is capped at the smaller of the two, check both. - NIC-level settings: segmentation/offload settings (TSO/GSO/LRO) and interface MTU (Maximum Transmission Unit) affect how efficiently the CPU can push bytes onto the wire; a misconfigured or disabled offload setting can bottleneck a fast link at the CPU rather than the network itself, worth ruling out especially if CPU utilization on the sending host is unexpectedly high relative to the achieved throughput.
Worked example
Suppose the link is rated at 1 Gbps with a 200ms round-trip time. The bandwidth-delay product is 1e9 bits/s times 0.2s = 2e8 bits = 25,000,000 bytes (25 MB). If the connection's actual window, even after scaling, tops out at 5 MB (one-fifth of the required 25 MB), the connection can never exceed roughly one-fifth of the link's rated throughput, which lines up suspiciously well with an observed 20%. That's the single most likely explanation to check FIRST, since it directly predicts the exact ratio being observed, before assuming something more exotic like a congestion-control mismatch.
Trade-offs & pitfalls
It's tempting to jump straight to "switch congestion-control algorithms" as the fix, but that's the wrong first move if the real limiter is window size: no algorithm change fixes a window that's structurally too small for the path's bandwidth-delay product. Always compute the bandwidth-delay product FIRST and check whether the achieved throughput lines up with a window-limited explanation before reaching for an algorithm change.
You discover that a team plan is technically solid but no longer matches a new business priority from leadership. What steps would you take to realign the plan, communicate the shift to the team, and minimize confusion or morale impact?
Sample Answer
First, I’d validate the new priority with leadership so I understand what changed and why. Then I’d compare it against the current plan and identify which work still supports the new goal, which work should pause, and which work should stop entirely.
Next steps:
- Reframe the plan around the new business objective
- Call out schedule, scope, and staffing impacts clearly
- Align with managers and key partners before announcing broadly
- Communicate to the team in a direct but calm way
I’d be explicit that the change is a business decision, not a judgment on the team’s work. That helps protect morale. I’d also acknowledge the effort already invested and explain what is being preserved, so people don’t feel like their work was wasted.
Finally, I’d reset expectations with stakeholders and set a short checkpoint to reduce confusion. The main goal is to move quickly, but with enough context that the team can reorient without losing confidence.
Worked example
Say a team is three weeks into building an internal analytics dashboard when leadership announces the company is deprioritizing internal tooling in favor of a customer-facing reporting feature. I'd first confirm with the sponsor that the dashboard is genuinely deprioritized rather than just delayed, then compare the two plans: the dashboard's data-pipeline work turns out to be directly reusable for the new feature, so that portion continues, while the dashboard-specific UI work pauses. In the team update I'd say something like: "Leadership has shifted priority to the customer-facing reporting feature this quarter. The pipeline work you've already built carries over directly, so that effort isn't wasted, but we're pausing the dashboard UI until the new feature ships. This is a business-priority change, not a reflection on the work." That gives the team a concrete before-and-after instead of an abstract instruction to "reframe the plan."
Segment a campus network for HR, Finance, Engineering and Guest. Which isolation controls do you apply at layer 2 and which at layer 3, and how does the design handle growth and compliance?
Sample Answer
Direct answer
Segment by what each group must reach, not by org chart. At Layer 2, give each group its own VLANs (virtual LANs, separate broadcast domains), assign users to them with 802.1X authentication (the switch keeps a port closed until the user or device logs in to a RADIUS server, a central server that checks credentials and replies with which VLAN to use), and harden the switch ports. At Layer 3, give each group its own VRF (a private routing table) and let traffic between groups cross only a stateful firewall (it tracks each connection, so replies to allowed traffic are accepted automatically) with a default-deny rule set (anything not explicitly permitted is dropped). Size each group's addresses with doubling headroom as one summarisable block, and keep the evidence auditors need.
Groups, VLANs and addresses
Assumed headcounts and devices per person: Engineering 800 users with 3 devices each, Finance 120 with 2, HR 60 with 2, Guest 400 concurrent clients at peak. Each block is sized for double today's devices, from the campus supernet (the large block that the group blocks are carved from) 10.20.0.0/16. Usable hosts are 2^(32 - prefix) - 2, minus the network and broadcast addresses: /19 gives 2^13 - 2 = 8,190; /23 gives 2^9 - 2 = 510; /24 gives 2^8 - 2 = 254; /22 gives 2^10 - 2 = 1,022.
| Group | VLANs | Block | Usable | Today | At 2x |
|---|---|---|---|---|---|
| Engineering | 301 to 316 (one per closet) | 10.20.0.0/19 | 8,190 | 2,400 (29%) | 4,800 (59%) |
| Finance | 120 | 10.20.32.0/23 | 510 | 240 (47%) | 480 (94%) |
| HR | 110 | 10.20.34.0/24 | 254 | 120 (47%) | 240 (94%) |
| Guest | 190 | 10.20.36.0/22 | 1,022 | 400 (39%) | 800 (78%) |
One block means one line. 10.20.0.0/19 covers 10.20.0.0 to 10.20.31.255, because its third octet runs from 0 to 31, so a single firewall rule or single route entry (permit source 10.20.0.0/19) names all of Engineering. If the closet subnets were scattered across the address plan, the rule would need one line per subnet and another for each new closet. Growing inside the block adds no line.
Engineering's /19 holds 32 subnets of /24; 16 are used today at 150 hosts each (59% full), and the other 16 (10.20.16.0/24 to 10.20.31.0/24) are the growth reserve, so doubling keeps every closet at 59%. Keeping each closet subnet small keeps broadcast domains small. 10.20.35.0/24 is left free so HR can grow to a /23 and stay one summary route. Finance has no free neighbour: at doubling its block is 94% full, so its growth path is a second block from the reserved 10.20.64.0/18 (64 /24s), at the cost of a second line in rules. Infrastructure uses 10.20.254.0/24 (128 /31 point-to-point links) and 10.20.255.0/24 for loopbacks.
Layer 2 controls
802.1X with VLAN assignment is what actually separates the groups at the port. The remaining rows harden the switch so the separation cannot be bypassed (rogue servers, forged ARP, VLAN hopping, which is a user tricking a switch into placing their traffic in another VLAN).
| Control | Purpose | How you prove it |
|---|---|---|
| 802.1X with RADIUS-assigned VLAN (the RADIUS reply carries three standard fields, Tunnel-Type=VLAN, Tunnel-Medium-Type=802 and Tunnel-Private-Group-ID set to the VLAN number, which together tell the switch "put this port in VLAN 120"; RFC 3580) | Port joins the VLAN of the user's group, not of the wall socket | Log in as a Finance and an HR user on the same port and confirm each lands in its own VLAN |
| MAC authentication bypass for printers and phones (the switch uses the device's MAC address as its credential) | Devices that cannot do 802.1X still land in a fixed VLAN | Plug in a printer, confirm its VLAN and that an unknown MAC gets none |
| DHCP snooping | Only trusted ports may send DHCP server replies; builds an IP-to-MAC binding table | Run a rogue DHCP server on a user port and confirm it is dropped |
| Dynamic ARP inspection | Checks ARP packets against the snooping bindings on untrusted ports | Send a forged ARP reply from a user port and confirm the drop is logged |
| Trunk hygiene: pruned allowed-VLAN lists, unused native VLAN (the VLAN whose frames cross a trunk without a tag), no dynamic trunk negotiation, unused ports shut | Stops VLAN hopping and VLAN sprawl | Compare each trunk's allowed list against the VLANs that closet needs |
| Guest wireless with client isolation (the access point refuses to forward traffic between wireless clients) | Guests cannot reach each other or internal networks | Two guest laptops fail to ping each other |
Layer 3 controls
- VRF per group at the distribution layer, with each group's SVIs (switch virtual interfaces, the gateway address for a VLAN) in its own VRF. Between VRFs there is no route unless the firewall provides it.
- Stateful firewall between VRFs: default deny, logged, rules by group summary block (one line per group, thanks to the summarised blocks: a rule for 10.20.0.0/19 covers every present and future Engineering subnet inside it).
- Guest VRF has only a default route to the internet zone, its own DNS and DHCP, and a rate limit.
- Policy matrix: with four groups there are 12 ordered pairs of source and destination group. One is allowed, 11 are denied.
| From | To | Rule | Pairs |
|---|---|---|---|
| HR | Finance | Allow only the payroll interface (one destination, TCP 443) | 1 allowed |
| HR | Engineering, Guest | Deny | 2 denied |
| Finance | HR, Engineering, Guest | Deny | 3 denied |
| Engineering | HR, Finance, Guest | Deny | 3 denied |
| Guest | HR, Finance, Engineering | Deny | 3 denied |
That is 1 allowed and 2 + 3 + 3 + 3 = 11 denied, 12 in all. HR, Finance and Engineering may reach shared services (DNS, DHCP, directory) through the services zone and the internet through the internet zone; both are outside these 12 pairs. Guest does not use the services zone: as the Guest VRF bullet says, it gets only its own DNS and DHCP and a default route to the internet zone, and it cannot reach the directory.
Growth and compliance
- Growth: new group = new VRF, a VLAN, and a /24 from the 10.20.64.0/18 reserve, plus one firewall rule block. Because blocks are summarised, rules do not multiply with VLAN count.
- Compliance: segmentation earns its audit value only when it is shown to work. If Finance handles payment card data, or a regulation limits who can see HR records, an auditor will ask for evidence. Keep the policy matrix under change control, firewall deny logs, 802.1X authentication logs, a quarterly review of the allow list with an owner, and the results of a periodic test that tries to connect from each group to each other group and records the denial.
- Failure mode to design against: a firewall outage cuts inter-group traffic, so run the pair as high availability, and test that failover keeps existing sessions.
Tell me about a time you had to lead or coordinate a high-risk network change across multiple regions or cloud providers. How did you plan validation, rollback, stakeholder communication, and post-change verification?
Sample Answer
Direct answer
The strongest version of this story shows a change that was risky specifically because it touched more than
one provider or region at once, a plan that made rollback boring rather than heroic, and a communication
cadence that kept stakeholders informed without needing to be asked. Structure the story as situation, task,
action, and result, and be specific about what made the change genuinely high-risk rather than describing an
ordinary change with dramatic language.
Structured elaboration
What makes a network change high-risk across regions or providers specifically. A single-region change
usually fails in one place and is visible immediately. A cross-region or cross-provider change (migrating a
BGP peering session, the point where two routers exchange dynamic routes over BGP, Border Gateway Protocol,
the protocol that dynamically advertises which networks are reachable through which path; changing a route
table, the list of rules a network uses to decide where to send traffic, shared across a transit gateway, a
cloud-provider hub that connects multiple networks or connections together; or cutting over a DNS-based
failover policy) can fail asymmetrically: one side applies cleanly while the other does not, and the
resulting inconsistent state is often harder to diagnose than a clean, total failure would have been.
Planning validation. Before the change window, validate in a lower environment that mirrors the
production topology closely enough to catch the specific failure mode you are worried about, and define, in
advance, the exact signal that means "this is working" (a specific health check passing, a specific latency
threshold held) rather than deciding in the moment whether things look fine.
Rollback. A rollback plan written after the change starts going wrong is not a rollback plan. Write and,
where feasible, rehearse the rollback steps before the change window starts, decide the trigger condition
that starts a rollback in advance (not a judgment call made under pressure), and know how long rollback
itself takes, since that duration is part of the actual risk budget for the change.
Stakeholder communication. Before the window: who needs to know this is happening, and what will they
observe if it goes well versus poorly. During the window: a lightweight, predictable update cadence, not
silence and not a flood of noise. After: a summary that includes what was validated post-change, not just
"it's done."
Post-change verification. Validate against the same explicit success signal defined during planning, not
a general "everything looks fine," and hold the validation window open long enough to catch a delayed
failure mode (a route that looks fine for the first ten minutes but breaks under a load pattern that only
occurs at a specific time of day).
Worked example
Situation: a company needed to migrate its primary on-prem-to-cloud BGP peering from a single VPN-based
session to a redundant pair spanning both a VPN and a newly provisioned dedicated circuit, without an
extended outage window, because the affected path carried live customer traffic around the clock. Task: plan
and execute the cutover with rollback available at every step. Action: validated the new dedicated-circuit
session in a staging environment against the same route-table structure production used, defined the
specific success signal in advance (BGP session state on the new circuit reaching Established, plus a
synthetic transaction completing end to end across it), scheduled the cutover for the lowest-traffic window
based on historical data, and shifted traffic in two stages, first a small percentage to validate the new
path under real load, then the remainder, with an automatic revert-to-VPN trigger defined for either stage if
the success signal was not observed within a fixed number of minutes. Communicated the plan and rollback
trigger to the on-call rotation and to the customer-facing team beforehand, sent one update at the start of
the cutover and one at each stage completion, and held the validation window open for the following business
day before declaring the migration complete. Result: the first-stage traffic shift surfaced a route
preference misconfiguration that the automatic trigger caught and reverted before it affected the full
traffic volume; the corrected configuration was validated again in staging, and the second attempt completed
without triggering a rollback, with the extended dedicated circuit carrying the full production path from
that point on.
Trade-offs and pitfalls
Do not present a change as high-risk without naming the specific asymmetric-failure mode that made it so; a
generic "it was scary" story reads as inexperience with what actually makes cross-provider changes different
from ordinary ones. Equally, do not skip past the part where the plan actually caught a real problem, a
story where nothing ever goes wrong reads as either lucky or embellished; a rollback trigger that fired once
and worked is a stronger answer than a flawless first attempt.
Explain how deploying services across multiple availability zones and regions changes network capacity planning. Cover the effect on latency, inter-region bandwidth, and replication costs, and how you'd factor cross-AZ and cross-region transfer fees into your sizing decisions.
Sample Answer
How the topology changes network capacity planning
Within a single AZ, inter-service network hops are effectively free to plan for: bandwidth is high, latency is low and stable, and cloud providers typically do not charge for traffic that stays within the same AZ. The moment a service spreads across AZs or regions, network capacity planning has to account for three things that were not there before: added latency, the bandwidth the replication/traffic itself needs, and the transfer fees the provider charges for that traffic.
Latency
Cross-AZ traffic within a region adds a small, fairly predictable latency (commonly single-digit milliseconds), usually fine to absorb into a synchronous call path. Cross-region traffic adds latency driven by physical distance (commonly tens to well over a hundred milliseconds), generally too slow for a synchronous request path, so it has to be planned for as asynchronous replication instead. That changes the bandwidth math: you are now sizing for a steady replication stream, not a single request/response pair.
Inter-region bandwidth and replication costs
Size cross-region bandwidth for the actual replication volume, not the request volume: for a database, that is write throughput times average write size times the number of regions receiving a copy; for event/log shipping, it is the raw event volume. Build in headroom above the steady-state rate, because if a target region falls behind (a network blip, a slow follower), it needs to catch up faster than steady-state without permanently falling further behind, which needs spare bandwidth on top of the baseline.
Factoring transfer fees into sizing
Cloud providers typically charge nothing, or very little, for traffic within an AZ, a small per-GB fee for cross-AZ traffic within a region, and a meaningfully larger per-GB fee for cross-region traffic. When sizing, that means: minimize what actually needs to leave the region (only replicate what genuinely needs a remote copy, not everything by default), track the recurring transfer cost as an ongoing line item that scales with traffic growth, not a one-time setup cost, and weigh it in the same trade-off as compute cost when deciding how many regions to run in and how much to replicate to each.
Write a Python script using Nornir that connects to devices in an inventory, runs a command to retrieve interface counters, and writes a CSV with columns device, interface, in_octets, out_octets. Assume devices support a standard 'show interface' output and that Nornir is already configured with credentials and inventory. Provide key code snippets and explain concurrency choices.
Sample Answer
Direct answer
Define one Nornir task that runs show interfaces through the netmiko_send_command task, parses each interface block with regular expressions (the ready-made TextFSM template for Cisco IOS show interfaces has no byte-counter fields), and returns plain rows. Run it across the inventory with the threaded runner (Nornir's default way of working through hosts: a pool of threads, each handling one device at a time, so many devices are contacted in parallel), then write the CSV once, from the main thread, after all hosts have finished. A host that fails stays isolated: Nornir records the failure on that host and carries on with the rest.
In the CSV, in_octets and out_octets are byte counts: an octet is 8 bits, which is one byte, and networking tools use the older word. The device prints them as bytes, so the script maps bytes to octets.
The script (collect_counters.py)
import csv
import re
import sys
from nornir import InitNornir
from nornir.core.task import Result, Task
from nornir_netmiko.tasks import netmiko_send_command
INTF_RE = re.compile(r"^(\S+) is .*, line protocol is ", re.M)
IN_RE = re.compile(r"^\s+(\d+) packets input, (\d+) bytes", re.M)
OUT_RE = re.compile(r"^\s+(\d+) packets output, (\d+) bytes", re.M)
def parse_interfaces(raw):
"""Split 'show interfaces' output into one block per interface."""
heads = list(INTF_RE.finditer(raw))
rows = []
for i, m in enumerate(heads):
end = heads[i + 1].start() if i + 1 < len(heads) else len(raw)
block = raw[m.start():end]
in_m, out_m = IN_RE.search(block), OUT_RE.search(block)
if in_m and out_m:
rows.append({"interface": m.group(1),
"in_octets": int(in_m.group(2)),
"out_octets": int(out_m.group(2))})
return rows
def get_counters(task: Task) -> Result:
sub = task.run(task=netmiko_send_command, command_string="show interfaces")
return Result(host=task.host, result=parse_interfaces(sub.result))
def main(config_file="config.yaml", out_path="counters.csv"):
nr = InitNornir(config_file=config_file)
agg = nr.run(task=get_counters)
with open(out_path, "w", newline="") as fh:
writer = csv.DictWriter(fh, fieldnames=["device", "interface", "in_octets", "out_octets"])
writer.writeheader()
for host, multi in agg.items():
if multi.failed:
continue
for row in multi[0].result:
writer.writerow({"device": host, **row})
for host in agg.failed_hosts:
print(f"FAILED {host}: {agg[host][-1].exception!r}", file=sys.stderr)
return len(agg.failed_hosts)
if __name__ == "__main__":
sys.exit(1 if main() else 0)
What the three regular expressions do, line by line: INTF_RE finds each interface header such as GigabitEthernet0/1 is up, line protocol is up and captures the interface name, which marks where one interface's block starts. IN_RE and OUT_RE then search inside one block for 83 packets input, 14855 bytes and 15513 packets output, 2510810 bytes and capture the packet count and the byte count. parse_interfaces cuts the text into blocks at the headers, so the counters are never attributed to the wrong interface. The file is opened only after nr.run returns, because the run is the part that works in parallel.
The assumed inventory and configuration are three small files (inv/groups.yaml is an empty file). config.yaml points Nornir at the inventory and sets the runner:
# config.yaml
inventory:
plugin: SimpleInventory
options:
host_file: inv/hosts.yaml
group_file: inv/groups.yaml
defaults_file: inv/defaults.yaml
runner:
plugin: threaded
options:
num_workers: 20
inv/hosts.yaml lists the devices (addresses from the documentation range):
# inv/hosts.yaml
sw1: {hostname: 192.0.2.1}
sw2: {hostname: 192.0.2.2}
sw3: {hostname: 192.0.2.3}
inv/defaults.yaml holds values every host inherits:
# inv/defaults.yaml
platform: cisco_ios
username: automation
Why it is built this way
- Task and sub-task: a task is a function Nornir runs once per host.
get_counterscallsnetmiko_send_commandwithtask.run, which starts a sub-task (a task run inside another) that opens the SSH session and runs the command, so the connection and the command are one unit of work per host.sub.resultis the raw text, and the returnedResultobject holds the parsed rows for that host. - Parsing: each interface block starts at a header line (
GigabitEthernet0/1 is up, line protocol is up) and contains83 packets input, 14855 bytesand15513 packets output, 2510810 bytes. The regular expressions take the byte counts asin_octetsandout_octets. An interface block missing either line is skipped, not written with a guess. - Credentials: the inventory carries the username; keep the password out of the file, for example
nr.inventory.defaults.password = os.environ["NET_PASSWORD"]afterInitNornir, or a secrets manager. - Writing the CSV: the file is written after
nr.runreturns, in the main thread, because worker threads are running tasks at the same time and concurrent writers to one file interleave rows.
Concurrency choices
num_workerssets how many hosts run at once (the configuration above sets 20 explicitly). The work is waiting on the network, so threads are a good fit; CPU-heavy parsing is small here.- The default Nornir runner is the threaded one. A thread is a lightweight worker inside one program, so while one worker waits for a slow switch the others keep going. Choose
num_workersfrom what the management plane can take, not from how many CPUs you have: very high values can exhaust management sessions or load the AAA (authentication, authorization and accounting, the login and permission service) server. Start low, raise it while watching failures. - Per-device error isolation comes from Nornir: an exception on one host marks only that host failed (
agg.failed_hosts), the other hosts' results still arrive.aggis the aggregated result of the run, one entry per host, andagg.failed_hostslists the hosts whose task raised an error. The script prints each failure with the underlying exception and returns a non-zero exit status, so a scheduler notices.raise_on_errordefaults toFalse, which is what allows partial results.
Evidence: run with the device call replaced
A real show interfaces needs a device, so this run replaces netmiko_send_command with a function that returns canned output for sw1 and sw2 (the captured text includes the lines quoted above) and raises TimeoutError for sw3. The parse, CSV, concurrency and error-isolation paths ran for real (Nornir 3.6.0, nornir-netmiko 1.0.1, Python 3.12); the SSH call to a switch did not.
# test_run.py
from nornir.core.task import Result
import collect_counters as cc
RAW = {
"sw1": """GigabitEthernet0/0 is up, line protocol is up (connected)
Hardware is iGbE
83 packets input, 14855 bytes, 0 no buffer
15513 packets output, 2510810 bytes, 0 underruns
GigabitEthernet0/1 is administratively down, line protocol is down (disabled)
0 packets input, 0 bytes, 0 no buffer
0 packets output, 0 bytes, 0 underruns
""",
"sw2": """GigabitEthernet0/0 is reset, line protocol is down (notconnect)
324 packets input, 48614 bytes, 0 no buffer
703 packets output, 62737 bytes, 0 underruns
""",
}
def fake_send(task, command_string):
if task.host.name not in RAW:
raise TimeoutError("simulated unreachable device")
return Result(host=task.host, result=RAW[task.host.name])
cc.netmiko_send_command = fake_send
failed = cc.main()
print("failed hosts:", failed)
print(open("counters.csv").read())
Printed output:
FAILED sw3: TimeoutError('simulated unreachable device')
failed hosts: 1
device,interface,in_octets,out_octets
sw1,GigabitEthernet0/0,14855,2510810
sw1,GigabitEthernet0/1,0,0
sw2,GigabitEthernet0/0,48614,62737
Two of three hosts succeeded, giving three data rows; the failed host contributes no rows.
Parsing BGP neighbor and prefix data
The same parsing question comes up for any show command, so here is a second one, BGP (the routing protocol that exchanges routes between networks; a neighbor is a peer router, and a prefix is one announced address block). The interface-counter script above uses regular expressions because the ready-made template has no byte fields; for output that a ready-made template does cover, use the template. For structured BGP output, collect show ip bgp summary and parse it with TextFSM and the ntc-templates package (or use use_textfsm=True on the Netmiko command). The template for Cisco IOS yields one record per neighbor with the fields bgp_neighbor, neighbor_as, up_down and state_or_prefixes_received. The template's ROUTER_ID value is marked Required and is captured from the BGP router identifier header line, so the parser must be given the whole command output, not only the table rows: on the two neighbor rows alone it returns an empty list. Run on this sample:
BGP router identifier 192.0.2.1, local AS number 65001
BGP table version is 12, main routing table version 12
2 network entries using 496 bytes of memory
Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd
192.0.2.2 4 65002 120 118 12 0 0 01:39:02 2
192.0.2.6 4 65003 0 0 1 0 0 never Idle
it printed 192.0.2.2 65002 01:39:02 2 and 192.0.2.6 65003 never Idle. AS is the neighbor's autonomous system number (the identifier of the network it belongs to), Up/Down is how long the session has been up (never if it has not come up), and Idle is the state of a session that is not established. The last column holds a prefix count when the session is established and a state name (here Idle) when it is not, so a numeric check separates up from down. Where a platform can return structured output itself (an API or NETCONF), prefer that over parsing text, because the format does not change with the software version.
Trade-offs and pitfalls
- Counters on large devices can be large; keep them as integers, and remember counters wrap or reset on reload, so a single sample is not a rate (take two samples and divide by the interval).
- Regular expressions depend on the platform's text: re-test them when software versions change, and add a template or regex per platform in a mixed fleet.
- Set a connection timeout and log the failing device name, so a stuck host does not hold the run silently.
You have a recurring 30-minute one-on-one with someone you mentor. Walk through how you'd structure the agenda to balance day-to-day blockers, skill development, and career conversation, and how that structure should evolve over a quarter.
Sample Answer
Direct answer
A recurring 30-minute 1:1 works best with a light, predictable structure (a quick check-in, blockers, a skill or growth item, and a career or forward-looking question), but the real skill is protecting the last two from being crowded out by whatever operational fire is loudest that week, and shifting the balance of the agenda as the relationship matures over the quarter.
Structured elaboration
A default structure for 30 minutes
| Segment | Rough time | Purpose |
|---|---|---|
| Check-in | 3-5 min | Surface anything urgent, gauge how they're actually doing |
| Blockers / operational | 8-10 min | Whatever's actively in their way right now |
| Skill or growth item | 8-10 min | One concrete thing they're building toward, not a status update |
| Forward-looking / career | 5-7 min | Where this is headed, not just what's happening this week |
Guarding against the common failure mode
A well-known failure pattern: the 1:1 happens reliably every week, on time, with all the segments technically present, but the career and growth segments become shallow ritual ("anything on your mind for growth?" "nope, all good") while blockers quietly eat the real time. The fix isn't just having a slot on the agenda, it's asking a specific, forward-looking question each cycle rather than an open-ended one, and being willing to occasionally protect that segment even when there's a real blocker competing for the time.
Diagnosing what's actually going on, not just tracking status
Part of the value of a recurring 1:1 is using it to figure out whether a struggle you're observing is a skill gap or a mindset or behavioral issue, because the two need different responses. Someone who's struggling because they don't yet know how needs teaching and practice; someone who's struggling because of avoidance, overconfidence, or a mismatch in how they're approaching the work needs a more direct conversation about the pattern itself, not more technical instruction. A 1:1 is a good place to probe for which one you're actually looking at before assuming.
An alternative structure for hands-on technical work
For roles where the most valuable use of the time is genuinely technical, a 1:1 doesn't have to follow the career-conversation template at all. Structuring it around live debugging together, walking through a real problem with explicit hypotheses ("I think it's X, here's how we'd check") and tracking which ones got ruled out, can be a more valuable use of 30 minutes than a generic status-and-goals agenda, especially early in a relationship when trust and technical credibility are still being built.
Evolving the structure over a quarter
- Early on, more of the time typically goes to blockers and establishing trust; the person needs to know the meeting is safe and useful before career conversations will be genuine rather than performative.
- As confidence builds, the balance should shift toward growth and forward-looking conversation, and the blockers segment should shrink because there's simply less friction to clear.
- If that shift isn't happening by mid-quarter, that's itself a signal worth naming directly rather than just continuing to run the same agenda.
Worked example
Situation
Early in a mentoring relationship, our 1:1s were almost entirely blockers: real, legitimate ones, but every week's slot filled up before we got near growth or career topics.
Action
I made an explicit change: reserved the last five minutes for a specific forward-looking question every time, stated as a fixed rule rather than something to get to if there was time, and moved lower-urgency blockers to async channels so they didn't have to consume the live time by default.
Result
By partway through the quarter, the ratio had genuinely shifted: blockers took less of the time because fewer new ones were coming up, and the growth and forward-looking segments started generating real, substantive conversation instead of the same shallow "all good" answer each week.
Trade-offs & pitfalls
- Mistaking a full agenda for a working one. Hitting every segment on the template doesn't mean the 1:1 is actually working if the career and growth segments are consistently shallow.
- Applying the same generic structure to a technical, debugging-heavy role. Forcing a career-conversation template onto a context where live technical problem-solving would be more valuable wastes the time on both sides.
- Not distinguishing skill gap from mindset issue. Responding to a mindset or behavioral pattern with more technical coaching, or the reverse, burns the time without addressing what's actually going on.
- Never revisiting the structure. A rigid agenda that never evolves as the mentee matures signals the relationship isn't actually progressing, even if the meeting keeps happening.
A financial firm's multicast market feed stopped reaching subscribers after a topology change. Describe how you would validate IGMP/MLD membership, PIM-SM or PIM-DM state, RPF checks, Rendezvous Point configuration, and how to inspect multicast routing tables ('show ip mroute') and IGMP snooping. Include steps to capture and analyze multicast traffic for subscribers.
Sample Answer
Direct answer
A multicast market-data feed failing after a topology change is diagnosed by walking the multicast control plane in order: confirm IGMP/MLD group membership is still correctly registered from subscribers, confirm the PIM (Protocol Independent Multicast) state that builds the actual distribution tree, and confirm the RPF (Reverse Path Forwarding) check, the mechanism that prevents loops, still passes given the NEW topology, since a topology change specifically can break RPF even when everything else about multicast configuration is untouched.
Structured elaboration
- Confirm IGMP/MLD membership is still current from subscribers:
show ip igmp groups(or the MLD/IPv6 equivalent) on the router closest to subscribers confirms they're still correctly signaling interest in the multicast group; if a topology change moved subscribers to a DIFFERENT router or interface, and IGMP snooping/membership wasn't correctly re-established on the new path, subscribers may appear to have "unsubscribed" even though their own IGMP behavior didn't actually change. - Confirm PIM state reflects the NEW topology correctly:
show ip pim neighborandshow ip mroute(multicast routing table) confirm whether PIM has correctly rebuilt its adjacencies and distribution tree given the topology change; PIM-SM (Sparse Mode) versus PIM-DM (Dense Mode) behave differently here, PIM-SM depends on correctly locating and maintaining state through a Rendezvous Point (RP), so also confirm the RP's own reachability and configuration are still correct post-change. - Confirm the Rendezvous Point configuration specifically, for PIM-SM: if the topology change altered the path to the RP (or, in a worse case, the RP itself is now unreachable due to the topology change), the entire PIM-SM tree-building process for NEW or re-established memberships can fail, even if existing state briefly persisted before finally timing out.
- Check the RPF check specifically, since topology changes are its classic failure trigger: multicast forwarding validates that traffic for a given source arrives via the INTERFACE that unicast routing would use to reach that source (the RPF check), specifically to prevent loops; a topology change that altered the UNICAST routing path to the multicast source, without a corresponding update propagating correctly through the multicast routing state, can cause the RPF check to fail, silently dropping otherwise-valid multicast traffic even though IGMP and basic connectivity look fine.
- Inspect multicast routing tables directly (
show ip mroute) for the specific group, checking the incoming interface list against what unicast routing NOW says the correct path to the source should be, which directly reveals an RPF mismatch if one exists. - Capture and analyze multicast traffic directly at points along the expected (and actual) distribution path to confirm definitively whether traffic is being correctly forwarded, dropped due to RPF failure, or simply never reaching a specific segment due to an IGMP/PIM state issue further upstream.
Worked example
show ip igmp groups confirms subscribers are still correctly signaling membership for the affected multicast group; that part of the chain is healthy. show ip mroute for the specific group shows the multicast routing entry's incoming interface no longer matches what show ip route (unicast routing, for the source's address) says the correct path should be, following the topology change; this mismatch is a direct RPF failure, meaning traffic arriving via the actual (now-correct, per updated unicast routing) path is being REJECTED by multicast forwarding because it doesn't match the STALE incoming-interface expectation still held in the multicast routing table. Clearing and allowing the mroute entry to correctly rebuild against the current, accurate unicast routing state resolves the RPF mismatch and restores the feed.
Trade-offs & pitfalls
It's easy to focus purely on IGMP (since that's the layer closest to and most visible from the subscriber's perspective) and miss that the topology change's actual impact is on RPF, a mechanism most engineers who don't work with multicast daily aren't as familiar with reasoning about; always check the multicast routing table's incoming-interface expectation against CURRENT unicast routing specifically after any topology change, since RPF mismatches are exactly the kind of failure a topology change is prone to introduce, even when every other piece of multicast configuration remains completely untouched.
What's the difference between availability and reliability for a distributed service? Give an example, like an HTTP API versus a background worker, where the two would be measured and prioritized differently.
Sample Answer
Direct answer
Availability is whether the service is up and responding right now, the percentage of time requests get a correct response. Reliability is whether the service does the correct thing every time over a longer horizon, even if that means taking longer or failing loudly rather than silently. A service can be highly available (always responds) while being unreliable (frequently returns wrong or incomplete results), and vice versa.
How they're measured differently
- Availability: uptime percentage, request success rate (successful responses over total requests), and latency, all measured in real time against a rolling window.
- Reliability: job or transaction success rate over time, data-loss incidents, mean time between failures, and correctness checks like reconciliation counts, none of which are visible from a single point-in-time health check.
Worked example: an HTTP API versus a background worker
An HTTP API's job is to respond fast and stay up, so availability is the priority metric. Suppose the API calls three dependencies in sequence to serve a request: an auth service at 99.95% availability, a database at 99.9%, and a cache at 99.99%. Because a single request needs all three to succeed, the composed availability is the product of the three:
Aserial=0.9995×0.999×0.9999≈0.99840That's under three nines even though every individual dependency is at or above three nines, because failures compound across a serial chain. In annual downtime terms:
downtimeserial=(1−0.99840)×525,600≈840.6 min/yrcompared to a single 99.9% dependency on its own:
downtimesingle=(1−0.999)×525,600≈525.6 min/yrChaining three otherwise-strong dependencies serially costs over 300 extra minutes of downtime a year versus just one of them alone. This is why an API-focused architect pushes hard on redundancy at each hop. To see how strong that lever is even when the underlying component is weaker, consider a hypothetical, cheaper cache tier, deliberately worse than the 99.99%-rated cache used above, where each individual replica only hits 99% availability on its own: two independent, parallel replicas of that weaker cache layer already beat any single component in the chain, the strong 99.99% cache included:
Aparallel=1−(1−0.99)2=0.9999A background worker processing a queue of jobs, by contrast, doesn't need to respond within milliseconds; what matters is that every job eventually completes correctly, with no silent data loss, which is a reliability property, not an availability one. If the worker is down for ten minutes and then resumes and correctly processes every job that queued up during that window, availability took a hit but reliability didn't; if the worker stays "up" the whole time but drops or duplicates 0.01% of jobs due to a bug, availability looks perfect while reliability has quietly failed.
Trade-offs & pitfalls
Optimizing for availability alone can mask reliability problems: a service that always responds quickly, even by returning stale or wrong data rather than waiting for a correct answer, looks perfect on an uptime dashboard while silently corrupting downstream state. The practical approach is deciding, per component, which property is actually load-bearing: user-facing APIs generally prioritize availability with graceful degradation for correctness-adjacent risk, while systems of record and background processing prioritize reliability, often accepting higher latency or even temporary unavailability rather than risk an incorrect or lost write.
Design the metric names and labels for interface metrics across 1,000 network devices so operators can query by device role, region and interface type without exploding the number of series. What would you refuse to label by?
Sample Answer
The rule behind everything below: each distinct combination of label values creates a separate time series (one metric name plus one exact set of label values, stored as a stream of timestamped numbers), so every label multiplies series count. The number of distinct values a label can take is its cardinality, and Prometheus guidance is not to use labels with high cardinality (many or unbounded values, such as user IDs or email addresses). It also says to put units in the metric name in base units, with _total on counters.
Metric names (base units, _total for counters, no label names in the name):
network_interface_receive_bytes_totalandnetwork_interface_transmit_bytes_totalnetwork_interface_receive_errors_totalandnetwork_interface_transmit_errors_totalnetwork_interface_receive_discards_totalandnetwork_interface_transmit_discards_totalnetwork_interface_oper_status(a gauge carrying the SNMP interface operational status value)
That is 7 series per interface, plus onenetwork_interface_infoseries. An info series is a metric whose value is always 1 and whose labels carry slow-changing text such as speed and description, so the text is stored once instead of on every series.
Labels on every series (all bounded):
| Label | Source | Values |
|---|---|---|
device | inventory hostname | about 1,000 |
interface | interface name (ifName) | about 48 per device |
device_role | inventory | a short fixed list (spine, leaf, access, edge, wan) |
region | inventory | a short fixed list |
if_class | derived from interface name | physical, lag, loopback, svi |
device_role and region come from the source of truth (the inventory system) and are attached as target labels (labels Prometheus adds to everything it scrapes from one device), not read from the device, so a renamed switch cannot change them silently. Operators then query by role and region without extra series, because those labels are constant per device:
sum by (region, device_role) (rate(network_interface_receive_bytes_total{if_class="physical"}[5m])) * 8
Reading it: a counter only ever grows, so rate(...[5m]) turns it into a per-second rate averaged over the last 5 minutes (a counter that rises by 7,500,000 bytes every minute gives 125,000 bytes per second). * 8 converts bytes to bits, so the result is bits per second. sum by (region, device_role) adds the per-interface rates into one total per region and role.
Series count. 1,000 devices x 48 interfaces x 8 series = 384,000, which stays 384,000 however many role and region combinations exist, because those labels do not split a device's series.
What I refuse to label by:
- VLAN ID on interface series. If 10% of ports (4,800 of 48,000) carry 100 VLANs each, those ports alone grow from 4,800 x 8 = 38,400 series to 4,800 x 100 x 8 = 3.84 million, and the fleet total becomes 3,840,000 + 345,600 (the other 43,200 ports, unchanged) = 4,185,600, about 10.9x the original 384,000.
- Client or peer IP address, or MAC address. If every port were an access port with 30 MACs, that is 384,000 x 30 = 11.52 million series, and they churn as clients move. That is flow-data territory, not metrics.
- Flow 5-tuple fields (the five values that identify a flow: source IP, destination IP, source port, destination port, protocol). Unbounded; send to the flow store.
- Free-text interface description or alias. Operators edit these, and every edit starts a new series and orphans the old history. Keep the text on the info series and join when you need it. A join matches each rate series to the one info series with the same
deviceandinterfaceand copies a label across:rate(network_interface_receive_bytes_total{device="leaf-12"}[5m]) * 8 * on (device, interface) group_left (ifAlias) network_interface_info. Multiplying by the info series' value of 1 leaves the number unchanged, andgroup_left (ifAlias)attaches the description to the result. (Evaluated withpromtool test ruleson illustrative series: a counter rising 7,500,000 bytes a minute returned 1,000,000 bit/s withifAlias="uplink-to-spine1"attached; the hostname and description are illustrative.) - Firmware version or serial on every series. Put it on a per-device info series.
Enforcement in the scrape config (validated with promtool check config):
- job_name: snmp_interfaces
sample_limit: 2000
metric_relabel_configs:
- source_labels: [__name__]
regex: ifHCInOctets
target_label: __name__
replacement: network_interface_receive_bytes_total
- regex: ifAlias|ifDescr
action: labeldrop
Reading the snippet line by line:
job_name: snmp_interfacesnames the scrape job (a scrape is one fetch of a device's metrics). This job carries the counters; the info series comes from a separate job without thelabeldroprule below, otherwise the text labels would be stripped from it too.metric_relabel_configsis a list of rewrite rules applied to every scraped sample after the scrape and before it is stored.- Rule 1:
source_labels: [__name__]is what to read (__name__is the hidden label that holds the metric name),regex: ifHCInOctetsis what to match (the 64-bit received-bytes counter from the IF-MIB, the standard SNMP interface table),target_label: __name__is what to write, andreplacementis the new value. With noactiongiven the rule does a replace, so the metric is renamed. - Rule 2:
action: labeldropdeletes every label whose name matchesifAlias|ifDescr(either name) from every sample, so editable free text can never become a label. sample_limit: 2000is the safety net in the next paragraph.
sample_limit makes Prometheus treat a scrape with more samples than the limit (counted after metric relabeling) as failed. 2,000 is about 5x the roughly 400 samples a 48-port device produces, so a device that suddenly exports a runaway table fails loudly instead of flooding storage. The exporter's own metric names (here the IF-MIB names) are renamed at ingestion, so dashboards depend on one scheme whichever protocol (SNMP or streaming telemetry) feeds it.
Review rule. Any new label needs a stated bound on its values and an owner; anything unbounded is a log or flow field instead.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Network Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs