Amazon Staff-Level Network Engineer Interview Preparation Guide
Amazon's interview process for Staff-level Network Engineer consists of initial recruiter engagement, multiple technical phone screens to assess infrastructure knowledge and architectural thinking, and a comprehensive onsite loop (5-7 rounds) evaluating technical depth, system design capability, operational excellence, security expertise, and alignment with Amazon's Leadership Principles. Staff-level candidates are expected to demonstrate mastery of large-scale network infrastructure, strategic thinking about technology decisions, and ability to influence cross-functional teams.
Interview Rounds
Recruiter Screening
What to Expect
Initial engagement combining recruiter phone screen and follow-up. The recruiter will validate your background, confirm interest in the Staff-level role, discuss career motivations, and assess cultural fit. This round also confirms understanding of the role's scope (network infrastructure design, multi-team collaboration, strategic initiatives). Recruiter may discuss compensation expectations and timeline. This is your opportunity to ask about team structure, current infrastructure challenges, and growth opportunities.
Tips & Advice
Be authentic and enthusiastic. Clearly articulate why you want to work at Amazon and on this specific team's networking challenges. Mention 1-2 specific achievements demonstrating scale and impact. Ask thoughtful questions about the team's current infrastructure priorities and pain points. Remember the recruiter is your advocate—be professional and personable. Avoid vague answers; provide concrete examples. Position yourself as both a deep technical expert and a team player capable of cross-functional influence.
Focus Topics
Questions About Role & Team
Prepare 3-4 thoughtful questions about current infrastructure challenges, team structure, reporting relationships, and strategic priorities the team is addressing.
Practice Interview
Study Questions
Alignment with Amazon Leadership Principles
Prepare brief examples showcasing principles like 'Customer Obsession' (building for reliability and performance), 'Operational Excellence,' and 'Think Big' in network infrastructure context.
Practice Interview
Study Questions
Career Trajectory & Motivation
Articulate your progression to Staff-level, why you're interested in Amazon specifically, and how this role aligns with your career goals in network engineering.
Practice Interview
Study Questions
Technical Background & Scale Experience
Highlight projects involving large-scale network infrastructure, multi-team coordination, or strategic infrastructure decisions. Emphasize metrics: deployment frequency, uptime improvements, cost optimization.
Practice Interview
Study Questions
Technical Phone Screen - Networking Fundamentals & Troubleshooting
What to Expect
45-60 minute technical assessment with a senior network engineer or architect from the team. Focus on deep protocol knowledge, troubleshooting methodology, and diagnostic skills. You'll discuss real-world networking problems, explain protocol behavior, and walk through structured troubleshooting approaches. Interviewer will present scenarios (packet loss, latency issues, routing problems, DNS failures) and evaluate your diagnostic reasoning, tool knowledge, and ability to isolate root causes systematically. This round assesses foundational expertise expected at Staff level.
Tips & Advice
Structure troubleshooting answers using a systematic approach: 1) Confirm the problem scope, 2) Gather data (interface stats, routing tables, packet captures), 3) form hypotheses, 4) test systematically, 5) implement solutions. Reference specific tools (tcpdump, netstat, traceroute, dig, ip command-line utilities) and explain what data they reveal. For any protocol discussion, explain both normal behavior and failure modes. Discuss how you'd approach the same problem at scale in production (monitoring, alerting, impact minimization). Show comfort with both vendor-specific CLIs and open-source tools. Don't rush—thinking aloud demonstrates your problem-solving process. If unsure, acknowledge it and explain how you'd investigate. Reference the job description scenarios (MTU mismatches, DHCP issues, VLAN routing, firewall rules, NAT configuration) covered in your preparation.
Focus Topics
Common Production Failure Scenarios
MTU mismatches in tunnels/VPNs, DHCP exhaustion, DNS resolver failures, port connectivity issues, NAT/firewall blockages, interface errors, ARP conflicts, VLAN misconfiguration, asymmetric routing.
Practice Interview
Study Questions
Routing & Path Analysis
Understanding routing decisions, path selection, convergence behavior, role of metrics and priorities. Troubleshooting black holes, asymmetric routing, suboptimal paths. Experience with dynamic routing protocols.
Practice Interview
Study Questions
OSI Layer Protocols & Behavior
Deep knowledge of Ethernet, IP (IPv4/IPv6), TCP/UDP, ICMP, ARP, DNS, routing protocols (BGP, OSPF, RIP). Understand normal vs abnormal behavior, timeout behaviors, retransmission strategies.
Practice Interview
Study Questions
Advanced Troubleshooting Methodology
Structured approach to isolating network problems: scoping symptoms, gathering telemetry, forming hypotheses, testing theories, implementing solutions. Ability to differentiate layer 2 vs 3 issues, endpoint problems vs infrastructure problems.
Practice Interview
Study Questions
Network Diagnostic Tools & Interpretation
Mastery of packet capture analysis (tcpdump, Wireshark), routing table inspection (show ip route, BGP routes), connection state tools (netstat, ss), DNS troubleshooting (dig, nslookup), performance analysis (iperf, mtr), flow analysis.
Practice Interview
Study Questions
Technical Phone Screen - Infrastructure Architecture & Design Patterns
What to Expect
45-60 minute technical interview with an architect or senior engineer, focusing on design thinking and large-scale infrastructure patterns. You'll discuss how you would architect network solutions for specific constraints, trade-offs between reliability/performance/cost, scaling strategies, and infrastructure evolution. Expect questions like 'How would you design a global network to minimize latency?', 'What's your approach to network segmentation for a multi-tenant platform?', or 'How do you evolve legacy network infrastructure while maintaining operations?' This round assesses your ability to think strategically about network systems, a critical Staff-level competency.
Tips & Advice
Approach design problems by first clarifying requirements and constraints (scale, geographic distribution, compliance, existing infrastructure, budget). Articulate trade-offs explicitly (redundancy vs complexity, performance vs cost, standardization vs flexibility). Discuss phased approaches to implementation. Reference real decisions you've made and why they were optimal given constraints. Include operational considerations: monitoring, alerting, runbooks, capacity planning. Discuss what you'd measure to validate the design works. For Staff-level, emphasize strategic thinking: how does this network design enable business capabilities? How would you evolve it as requirements change? Mention cross-team collaboration and how you'd communicate tradeoffs to stakeholders. Show awareness of cloud networking (hybrid architectures, cloud integration) as it's increasingly relevant to modern network design.
Focus Topics
Cloud & Hybrid Network Integration
Connecting on-premises infrastructure to cloud environments, VPC design, hybrid connectivity (Direct Connect, VPN), multi-cloud networking considerations, edge computing implications.
Practice Interview
Study Questions
Operational Excellence & Evolution Strategy
Phased infrastructure evolution, backward compatibility, deprecation strategies, zero-downtime transitions, monitoring and observability design, capacity planning, cost optimization.
Practice Interview
Study Questions
Network Security Architecture
Segmentation strategies (VLANs, micro-segmentation, zero-trust models), DDoS mitigation, firewall architectures, encryption in transit, access controls, compliance frameworks (ISO 27001, SOC 2), threat modeling.
Practice Interview
Study Questions
Network Architecture Principles & Trade-offs
Designing network systems with understanding of reliability vs complexity, performance vs cost, standardization vs flexibility. Justifying architectural decisions based on requirements. Phased implementation and risk mitigation strategies.
Practice Interview
Study Questions
Scalability & High-Availability Design
Designing networks that scale geographically or in capacity. Redundancy patterns, failover strategies, load balancing approaches, active-active vs active-passive, convergence time considerations, zero-downtime scaling.
Practice Interview
Study Questions
Onsite Round 1 - Network Architecture System Design
What to Expect
55-60 minute deep-dive system design interview with a senior architect. You'll be presented with a large-scale network design problem (e.g., 'Design a global network for a company with data centers on 4 continents and millions of users,' or 'Design network infrastructure for a multi-tenant SaaS platform requiring strict data isolation'). You'll whiteboard or discuss your approach, covering requirements analysis, architectural components, routing strategy, redundancy, security segmentation, monitoring, and scalability. Interviewer will probe trade-offs, ask 'what-if' questions, and challenge your assumptions. This round deeply evaluates architecture maturity and ability to handle enterprise-scale problems.
Tips & Advice
Start by clarifying all requirements and constraints—don't assume. Sketch your architecture, labeling key components (core routers, access layers, firewalls, security zones, monitoring infrastructure). Walk through traffic flows for different scenarios (north-south, east-west, disaster recovery). Explicitly discuss how you handle redundancy, convergence time, and failure scenarios. Address operational concerns: monitoring, alerting, troubleshooting approach, capacity planning. Discuss security from the start, not as an afterthought. Invite feedback: 'What aspects would you probe further?' or 'Are there constraints I'm missing?' For Staff-level, emphasize strategic impact: how does this design support business goals? How would you phase implementation? What are the risks and mitigation? How do you communicate tradeoffs to leadership? Be ready to dive deep on any component—if you mention BGP, expect protocol-level questions. Show awareness of industry patterns and best practices, but justify why they apply here. Discuss both greenfield and brownfield (evolving existing) scenarios.
Focus Topics
Scalability & Capacity Planning Architecture
Designing for known and unknown growth, non-disruptive scaling strategies, capacity forecasting integration, modular design enabling incremental expansion, cost-efficient scaling patterns.
Practice Interview
Study Questions
Security Segmentation & Zero-Trust Architecture
Micro-segmentation strategies, security zone design, identity-based access (zero-trust), encryption architecture, compliance integration, threat model-driven design.
Practice Interview
Study Questions
Operational Architecture & Observability
Designing monitoring and alerting into architecture, telemetry collection strategy, operational dashboards, automation points, troubleshooting accessibility, runbook integration.
Practice Interview
Study Questions
High-Availability & Disaster Recovery Architecture
Multi-site failover, active-active design patterns, convergence optimization, recovery time objectives (RTO) and recovery point objectives (RPO) considerations, split-brain prevention, cross-region synchronization.
Practice Interview
Study Questions
Enterprise-Scale Network Architecture Design
Designing networks supporting hundreds of thousands to millions of endpoints, multi-site/multi-region deployments, complex traffic patterns. Design must handle growth, geographic distribution, and business resilience requirements.
Practice Interview
Study Questions
Onsite Round 2 - Advanced Networking Technologies & Protocol Deep Dive
What to Expect
55-60 minute technical interview with a senior network engineer focusing on deep protocol knowledge, advanced technologies, and complex scenarios. You'll discuss advanced routing protocols (BGP, OSPF, IS-IS), MPLS, segment routing, network programmability, SDN concepts, emerging technologies, and how to apply them in modern infrastructure. Expect detailed questions about protocol behavior under stress, convergence optimization, and unusual edge cases. You may be asked to design solutions using specific technologies or evaluate technology choices for particular problems. This round assesses your technical depth and awareness of networking evolution.
Tips & Advice
Be conversant with modern routing protocols and their nuances. Understand BGP well—it's the internet's backbone and Amazon's networks rely on it heavily. Know OSPF for internal routing. Understand MPLS use cases (traffic engineering, VPNs, fast reroute). Be familiar with segment routing as an emerging technology. Discuss SDN concepts and infrastructure-as-code for networking. For any technology mentioned, be prepared to discuss: when you'd use it, design tradeoffs, operational complexity, scalability implications. Share experiences with these technologies, including challenges overcome. Show awareness of networking trends: programmable networks, intent-based networking, AI/ML for network optimization, disaggregated networking. However, balance theory with practical applicability—Staff-level engineers are pragmatists. Discuss how you stay current with networking evolution. If asked about unfamiliar tech, acknowledge it honestly and explain how you'd evaluate it. For Staff-level, emphasize how technology choices support business and operational goals.
Focus Topics
Emerging Technologies & Evolution
Segment routing (SR), software-defined WAN (SD-WAN), network telemetry and analytics, AI/ML applications in networking, zero-trust security evolution, edge computing networking implications.
Practice Interview
Study Questions
MPLS & Traffic Engineering
MPLS forwarding paradigm, label distribution protocols (LDP, RSVP-TE), traffic engineering use cases, fast reroute (FRR), VPN applications (L3VPN, L2VPN), segment routing as modern alternative.
Practice Interview
Study Questions
Interior Routing Protocols (OSPF, IS-IS)
Link-state protocol mechanics, area design, fast convergence techniques, OSPF types and traffic engineering, IS-IS multilevel design, comparison of OSPF vs IS-IS tradeoffs.
Practice Interview
Study Questions
Network Programmability & SDN Concepts
Infrastructure-as-code principles, API-driven networking, OpenFlow/NETCONF/YANG, controller-based architectures, disaggregated networking, intent-based networking, automation frameworks.
Practice Interview
Study Questions
Border Gateway Protocol (BGP) Deep Dive
BGP path selection, convergence behavior, best path algorithm, communities and attributes, route filtering, failover handling, multi-AS design, BGP security (RPKI, route validation), optimization techniques (aggregation, summarization).
Practice Interview
Study Questions
Onsite Round 3 - Network Security, Compliance & Operations
What to Expect
55-60 minute interview focusing on network security architecture, compliance requirements, and operational excellence. You'll discuss designing secure networks (segmentation, access controls, encryption, DDoS mitigation), meeting regulatory requirements (ISO 27001, SOC 2, HIPAA, PCI-DSS), security monitoring and incident response, and operational best practices (change management, automation, documentation). Expect scenarios: 'Design a network for a healthcare provider with strict data residency,' or 'How do you evolve firewall rules as business needs change?' You'll be evaluated on balancing security with operational efficiency, understanding business and compliance drivers, and building security into infrastructure from the start.
Tips & Advice
Approach security as an architectural concern, not an afterthought. Discuss defense-in-depth strategies: multiple security layers, redundancy in security controls, fail-secure defaults. Demonstrate understanding of compliance frameworks and how they drive network design. For regulatory requirements, show you understand business drivers: data residency for GDPR, encryption for HIPAA, segmentation for PCI-DSS. Discuss monitoring and alerting for security events. Address incident response: detection, containment, investigation, recovery. Show familiarity with security tools and best practices. Discuss the tension between security and usability/performance, and how to resolve it through design. For operational excellence, discuss change management, testing strategies, rollback procedures. Emphasize automation to reduce human error and improve consistency. Share examples of security improvements you've implemented and measured. At Staff-level, discuss how you've influenced security culture and mentored engineers on secure design. Be candid about security tradeoffs and limitations of any approach.
Focus Topics
Compliance Frameworks & Regulatory Requirements
ISO 27001, SOC 2, HIPAA, PCI-DSS, GDPR implications for network design. Data residency requirements, audit trails, logging requirements, network segmentation for compliance, documentation and evidence collection.
Practice Interview
Study Questions
Operational Excellence & Change Management
Change management processes, testing strategies for changes, rollback procedures, documentation practices, automation for consistency, monitoring and alerting, incident response procedures, runbooks, knowledge management.
Practice Interview
Study Questions
Encryption & Data Protection in Transit
TLS/SSL deployment, certificate management, encryption protocols (IPsec, TLS 1.3), key management, encrypted tunnels (VPN, WireGuard), encryption for data center interconnect, limitations and performance considerations.
Practice Interview
Study Questions
Network Segmentation & Zero-Trust Architecture
VLAN strategy, DMZ design, micro-segmentation principles, least-privilege access control, identity-based networking, zero-trust model implementation, internal vs external threat models, segmentation validation.
Practice Interview
Study Questions
Firewall Architecture & Access Control
Firewall placement and design, stateful vs stateless filtering, next-generation firewall capabilities, rule design and management, DDoS mitigation strategies, NAT and port translation, VPN termination.
Practice Interview
Study Questions
Onsite Round 4 - Operational Excellence & Troubleshooting Leadership
What to Expect
55-60 minute interview with a senior operations engineer or infrastructure leader assessing your approach to operational excellence, troubleshooting at scale, and incident management. You'll discuss challenging production incidents you've handled, your troubleshooting methodology, how you've improved operational processes, automation investments, and monitoring strategy. Expect scenarios: 'Walk us through your worst production incident and how you'd prevent it,' or 'How do you evolve network monitoring as infrastructure grows?' This round evaluates maturity in running production systems, learning from failures, and building reliable, maintainable infrastructure. Staff-level candidates are expected to mentor others in these practices.
Tips & Advice
Prepare 2-3 detailed incident stories showing your troubleshooting methodology, decision-making under pressure, and learnings. Structure incident discussions: what happened (symptoms), what did you check first, how did you isolate root cause, what was the fix, how did you prevent recurrence? Show systematic thinking and how you involve others. Discuss monitoring strategy: what do you measure, what triggers alerts, how do you avoid alert fatigue? For automation, explain what you've automated and why—prioritize high-impact, repetitive tasks. Discuss documentation: runbooks, architecture diagrams, change logs. Emphasize learning from failures through blameless post-mortems. At Staff-level, discuss how you've grown team capabilities in these areas through mentoring and process improvement. Show awareness of balancing innovation with stability. Discuss metrics: mean time to recovery (MTTR), change failure rate, deployment frequency. Be specific about improvements you've driven and their business impact.
Focus Topics
Documentation & Knowledge Management
Architecture documentation, runbook creation and maintenance, change logs, incident postmortem documentation, architectural decision records (ADRs), keeping documentation current, making documentation accessible and useful.
Practice Interview
Study Questions
Automation & Operational Efficiency
Identifying high-impact automation opportunities, building vs buying tools, automation frameworks, reliability and testing of automation, reducing manual toil, enabling faster deployment, documentation and knowledge capture.
Practice Interview
Study Questions
Incident Management & Post-Incident Learning
Incident classification and escalation, response procedures, communication during incidents, blameless post-mortem culture, root cause analysis, prevention planning, tracking and follow-up, team learning from failures.
Practice Interview
Study Questions
Monitoring, Alerting & Observability Design
Comprehensive monitoring strategy covering availability, latency, throughput, error rates, resource utilization. Alert design avoiding false positives. Logging and telemetry collection. Dashboards for different audiences. SLO/SLI definition. Trend analysis and capacity planning integration.
Practice Interview
Study Questions
Production Troubleshooting Methodology & Leadership
Systematic troubleshooting approaches, data collection and analysis, hypothesis formation and testing, time pressure management, communicating with stakeholders, escalation paths, teaching troubleshooting skills to junior engineers.
Practice Interview
Study Questions
Onsite Round 5 - Amazon Leadership Principles & Behavioral Assessment
What to Expect
55-60 minute behavioral interview, potentially with multiple interviewers or a dedicated behavioral round. Focused on assessing alignment with Amazon's 16 Leadership Principles through structured behavioral questions and project discussions. You'll discuss specific situations where you demonstrated principles like 'Customer Obsession' (designing for user reliability), 'Operational Excellence,' 'Learn and Be Curious,' 'Earn Trust,' 'Hire and Develop the Best' (mentoring), 'Think Big,' 'Invent and Simplify,' 'Are Right, a Lot' (decision-making), 'Dive Deep,' 'Deliver Results,' 'Frugality,' 'Bias for Action,' 'Ownership,' 'Strive for Tenacity,' and 'Earn Respect.' For Staff-level, emphasis is on demonstrated leadership influence across teams.
Tips & Advice
Prepare 4-5 detailed project stories using the STAR method (Situation, Task, Action, Result) that showcase Amazon Leadership Principles. Each story should highlight specific principles and quantifiable impact. For Staff-level, focus on: influencing cross-functional teams, mentoring engineers, driving organizational change, strategic thinking, handling ambiguity, making difficult tradeoffs. Connect your network engineering work to Amazon's business: 'This reliability improvement enabled faster deployments, reducing customer time to value,' or 'Network segmentation redesign simplified compliance validation for our auditors.' Be specific about metrics: cost saved, time reduced, reliability improved, team capability increased. Discuss failures and learnings—vulnerability at Staff-level shows maturity. When asked about disagreements with peers, show 'agree and commit' culture: you heard other perspectives, made best decision for business, and committed to execution. Avoid generic answers; provide concrete examples. Practice concise storytelling—you have ~2-3 minutes per story. Listen carefully to questions and answer what's asked, not a prepared story. Show genuine passion for reliability, customers, and team development. At Staff-level, discuss your influence on culture and how you've developed talent. For 'Earn Trust,' discuss confidentiality and reliability in following through on commitments.
Focus Topics
Amazon Leadership Principle: Dive Deep
Understanding network systems at depth rather than superficially. Auditing designs and processes. Asking probing questions to understand root causes. Not accepting incomplete information or hand-wavy explanations.
Practice Interview
Study Questions
Amazon Leadership Principle: Think Big
Envisioning transformative network infrastructure that enables new capabilities. Driving strategic initiatives beyond incremental improvement. Considering long-term technology evolution. Proposing solutions at appropriate scope for business needs.
Practice Interview
Study Questions
Amazon Leadership Principle: Deliver Results
Driving initiatives to completion despite obstacles. Managing competing priorities and tradeoffs. Holding self and team accountable to commitments. Celebrating achieved goals while identifying further improvements.
Practice Interview
Study Questions
Amazon Leadership Principle: Operational Excellence
Driving processes, automation, monitoring that enable reliable operations. Continuously improving operational effectiveness. Mentoring team in operational discipline. Building cultures of operational rigor.
Practice Interview
Study Questions
Amazon Leadership Principle: Hire and Develop the Best
Recruiting and mentoring talented engineers. Creating growth opportunities for team members. Teaching troubleshooting, design thinking, and operational excellence. Developing next generation of network leaders.
Practice Interview
Study Questions
Amazon Leadership Principle: Customer Obsession
Designing network infrastructure with deep understanding of customer/business needs. Driving improvements that directly impact user experience or business operations. Considering reliability, latency, and cost from customer perspective.
Practice Interview
Study Questions
Frequently Asked Network Engineer Interview Questions
A cloud platform must give each of about 10,000 tenants its own isolated L3 network on shared Linux hosts. How would you build the host networking, handle addressing and tenant routing, isolate east-west traffic, and keep performance and monitoring manageable?
Sample Answer
Direct answer
Give each tenant its own VRF (virtual routing and forwarding instance: a separate routing table) on every host that runs one of its workloads, and carry tenant traffic between hosts inside VXLAN (Virtual Extensible LAN) tunnels with one VNI (VXLAN Network Identifier) per tenant. Use BGP (Border Gateway Protocol) EVPN (Ethernet VPN, a BGP address family that advertises tenant routes and MAC addresses so hosts do not have to flood and learn: the older switch method of sending unknown traffic everywhere and remembering who answers) as the control plane, run by FRR (an open-source routing daemon) on each host. Plain VLANs cannot do the job: a VLAN ID is a 12-bit number, which allows 4094 usable values (0 and 4095 are reserved), and RFC 7348 says that limit is inadequate for multi-tenant environments, while a VNI is a 24-bit value that allows up to 16 million segments. Ten thousand tenants is already 2.4 times the whole VLAN space.
The core of the design is the two lines of control: a VRF per tenant for routing isolation and a VNI per tenant for overlay isolation. The firewall, MTU, monitoring and offload sections below make that design safe and operable at 10,000 tenants.
The path a packet takes through the host
flowchart LR
W["Workload: network namespace or VM"] --> V["veth into the tenant VRF"]
V --> F["nftables forward hook: default drop"]
F --> B["Bridge + VXLAN device, VNI per tenant"]
B --> U["Underlay IP fabric, MTU 1550 or more"]
U --> R["Remote host: VNI, VRF, workload"]
C["FRR BGP EVPN"] -.-> B
| Layer | Linux object | Job |
|---|---|---|
| Workload | Network namespace (netns) or VM tap (a virtual network port for a VM), connected to the host by a veth (a pair of virtual Ethernet ports, like a cable: what goes in one end comes out the other) | A netns gives the workload its own interfaces, routes, firewall and sysctls. A VRF only gives one shared network stack a second routing table, so the workload gets the netns and the host side gets the VRF. |
| Tenant router | VRF device (ip link add vrf-A type vrf table 1001) | The VRF is the tenant's router on this host. Interfaces are attached with ip link set dev NAME master vrf-A, and their connected routes move into the VRF's table (kernel VRF documentation). |
| Tenant tunnel | Bridge enslaved to the VRF plus a VXLAN device with nolearning | Carries the tenant's L3 VNI. Learning is off because EVPN supplies the MAC and route entries instead of flood-and-learn. |
| Address of the host in the underlay | A VTEP (VXLAN tunnel endpoint) IP such as a loopback | Never visible to tenants and never in a tenant VRF. |
One VXLAN device per tenant would mean up to 10,000 devices on a host. FRR's EVPN documentation describes a single VXLAN device mode (ip link add vxlan0 type vxlan dstport 4789 local <VTEP IP> nolearning external vnifilter, then bridge vni add dev vxlan0 vni 100, plus a bridge VLAN id mapped to each VNI with tunnel_info id, plus a VLAN interface on the bridge enslaved to the VRF). In plain words: external makes one device handle many VNIs; vnifilter plus bridge vni add lists which VNIs this device accepts; and tunnel_info id maps a local bridge VLAN id to a global VNI, so the bridge can translate between the two. Each tenant on the host therefore borrows one local VLAN id, and a VLAN id is a 12-bit number with 4094 usable values. That VLAN id only has local meaning (it never leaves the host; the VNI is what travels), so one host can hold at most 4094 tenants this way, which is far above how many tenants have a workload on one host. Hosts create a tenant's VRF and VNI mapping when its first workload lands and remove it with the last one, so the object count per host follows resident tenants and not the 10,000 total.
Addressing and tenant routing
- Tenant address space. Each tenant picks its own CIDR (for example a /24 per subnet carved from the tenant's block). Overlap between tenants is legal because every VRF is a separate table; the platform's IP address management (IPAM) only enforces uniqueness inside one tenant. The demo below gives tenants A and B the same 10.0.1.0/24 and 10.0.2.0/24.
- Underlay. The fabric's own addresses (VTEP IPs, BGP peering) live in the default VRF from a range tenants never see.
- Routing between hosts. Use symmetric routing: the source host routes the packet inside the tenant VRF, encapsulates it with that VRF's L3 VNI, and the destination host decapsulates and routes again. Symmetric routing means both the sending and receiving host do a routing lookup in the tenant's VRF, with the tunnel in between carrying a single per-tenant VNI. Tenant prefixes travel as EVPN type-5 (IP prefix) routes (a route type that says "this address block is reachable behind this host"). In FRR the VRF-to-VNI mapping is a
vrf vrf1stanza containingvni 100;advertise-all-vnigoes insideaddress-family l2vpn evpnof the defaultrouter bgp <ASN>instance (ASN is the autonomous system number); andadvertise ipv4 unicastgoes insideaddress-family l2vpn evpnof the per-tenantrouter bgp <ASN> vrf vrf1instance. FRR documents that prefixes are not exported as type-5 routes until thatadvertiseline exists. - Route targets. FRR derives route targets automatically: import is the wildcard
*:VNIand export is(AS & 0xFFFF):VNI. A route exported for VNI 5001 is therefore imported only by the VRF that owns VNI 5001, which is the control-plane half of tenant isolation. A route target is a label attached to a BGP route that says which VRFs may import it. Worked example (illustrative AS 65001, VNI 5001):65001 & 0xFFFFis 65001, because 65001 is below 65536 and fits in two bytes, so the export label is65001:5001. Tenant A's VRF, which owns VNI 5001, imports anything labelled*:5001(any AS, VNI 5001), so it accepts this route. Tenant B's VRF owns VNI 5002 and imports*:5002, so it ignores the route. The*wildcard on the AS side is what lets every host import routes from any other host's AS, while the VNI half keeps tenants apart. - Leaving the cloud. Overlapping tenants cannot share one flat internet or on-premises route table, so each tenant's egress goes through a per-tenant NAT or gateway attachment, and any shared service (DNS, metadata, a managed database) is reached by an explicit, audited route import into that one VRF, never by leaking whole tables.
Isolating east-west traffic
Isolation comes from stacking independent layers so one mistake does not expose a tenant:
- Routing. A VRF has no route to another tenant's prefixes. The demo shows tenant A's workload receiving no reply for 10.9.9.2, a prefix that exists only in tenant B.
- Overlay. Frames carry a VNI, and a host only accepts VNIs it has configured.
- Host firewall. An nftables table (nftables is the Linux kernel's packet-filtering framework) hooked on
forwardwithpolicy drop, onect state established,related acceptline, and explicit per-tenant allow rules. Only routed traffic crosses this hook, so two workloads of one tenant that share a bridge need their own bridge-level filtering, and a test should confirm which path same-host traffic takes.
table inet tenant_fw {
chain forward {
type filter hook forward priority 0; policy drop;
ct state established,related accept
iifname "vrf-A" ip daddr 10.0.2.0/24 icmp type echo-request accept
counter comment "dropped by default"
}
}
Pitfall found while testing this on a Linux 7.0 kernel: a first version matched iifname "wv1A" (the host side of the workload veth) and silently dropped everything, including the allowed ping. For VRF-routed traffic the forward hook reported the VRF device (vrf-A) as the input interface. With the rule matching vrf-A, tenant A's ping passed and tenant B's identical ping was dropped. Check the interface name your own kernel presents before trusting a rule.
Performance
- MTU (maximum transmission unit). VXLAN adds about 50 bytes of outer Ethernet, IP, UDP and VXLAN headers (RFC 7348), so tenants that get a 1500-byte MTU need at least 1500 + 50 = 1550 on every underlay link, or a jumbo underlay. The demo sets 1550 on the underlay link, the kernel set the VXLAN device to 1500, and a 1500-byte packet with the do-not-fragment bit passed.
- ECMP spreading. RFC 7348 recommends deriving the outer UDP source port from a hash of the inner packet, which lets the fabric's equal-cost multipath (ECMP) spread tenant flows over all uplinks.
- Offload. Offload means letting the network card (NIC) do work the CPU would otherwise do. The kernel segmentation documentation lists UDP-tunnel segmentation types (such as SKB_GSO_UDP_TUNNEL; GSO is generic segmentation offload, splitting large packets into wire-size ones late or in hardware), so tunnel traffic can be segmented and checksummed in the NIC. Confirm each candidate NIC with
ethtool -kand compare iperf3 throughput with and without the tunnel before committing to hardware. - Control plane. Per-host cost follows resident tenants, as above. Watch the BGP table size with
show bgp l2vpn evpn summaryduring the scale test.
Keeping monitoring manageable
| Signal | Where it comes from | Why it matters |
|---|---|---|
| Per-tenant traffic | ip -s link show vrf-A (receive side only) plus the host-side veth counters of the tenant's workloads, summed per tenant by the collector | Bytes and packets per tenant in both directions, one series per tenant |
| Control plane | show bgp l2vpn evpn summary, show vrf vni, show evpn mac vni (FRR) | Session state, VNI-to-VRF mapping, learned MACs |
| Policy drops | nftables counters | Distinguishes a blocked flow from a broken path |
| Isolation | A scheduled negative probe: tenant A must NOT reach a canary only tenant B owns | Catches a leak that no throughput graph shows |
Four counters (receive and transmit bytes and packets) per tenant is 4 x 10,000 = 40,000 time series for the whole fleet. One measured caveat shapes where those counters come from: in the demo run the VRF device's TX counters stayed at 0 for forwarded traffic and only its RX counters moved, so the VRF device alone gives one direction. The other direction comes from the host-side veth of each workload, where RX is what the workload sent and TX is what was delivered to it. The VXLAN device is not a per-tenant source when one device carries every VNI. The collector reads the per-workload counters locally and exports only the per-tenant sum; per-workload series are exported for one tenant only while debugging.
Worked example, executed
The script builds two hosts as namespaces joined by an underlay link, gives tenants A (VNI 5001) and B (VNI 5002) overlapping addresses, and installs by hand the neighbor, forwarding-database and route entries that BGP EVPN would install. Save the script below as tenants.sh in the current directory and run it as root on Linux with iproute2 and iputils-ping, for example docker run --rm --privileged -v "$PWD":/w debian:stable-slim bash -c "apt-get update -qq && apt-get install -y -qq iproute2 iputils-ping && bash /w/tenants.sh".
#!/usr/bin/env bash
# Run as root in a Linux environment with iproute2 and iputils-ping (for example a privileged container).
# Two hypervisor "hosts" (h1, h2) as namespaces joined by an underlay link.
# Each tenant gets: a VRF, a bridge + VXLAN device carrying its L3 VNI, and one workload namespace per host.
set -euo pipefail
ip netns add h1; ip netns add h2
ip link add u1 type veth peer name u2
ip link set u1 netns h1; ip link set u2 netns h2
ip -n h1 addr add 172.16.0.1/30 dev u1; ip -n h2 addr add 172.16.0.2/30 dev u2
ip -n h1 link set u1 mtu 1550 up; ip -n h2 link set u2 mtu 1550 up # 1500 inner + 50 VXLAN overhead
# tenant_setup <name> <vni> <vrf-table> <router-mac-h1> <router-mac-h2> <transit-net>
tenant_setup() {
local t=$1 vni=$2 tbl=$3 m1=$4 m2=$5 n=$6
for h in 1 2; do
# r is the other host (h=1 gives 2, h=2 gives 1); ${!mac} reads the variable whose NAME is stored in mac (m1 or m2)
local r=$((3 - h)); local mac=m$h rmac=m$r
ip -n h$h link add vrf-$t type vrf table $tbl; ip -n h$h link set vrf-$t up
ip -n h$h link add br-$t type bridge
ip -n h$h link set br-$t address ${!mac} master vrf-$t up
ip -n h$h link add vx-$t type vxlan id $vni dstport 4789 local 172.16.0.$h nolearning
ip -n h$h link set vx-$t master br-$t up
ip -n h$h addr add 10.255.$n.$h/32 dev br-$t
# STAND-IN FOR EVPN (two lines): a type-2/type-5 route would tell this host the remote router's MAC and VTEP.
# neigh add ... nud permanent: a fixed ARP entry (remote router IP -> remote router MAC) that never ages out.
# bridge fdb add ... dst: forwarding entry saying frames for that MAC go inside VXLAN to the remote VTEP IP.
ip -n h$h neigh add 10.255.$n.$r lladdr ${!rmac} dev br-$t nud permanent
bridge -n h$h fdb add ${!rmac} dev vx-$t dst 172.16.0.$r self static
# tenant workload: veth into the VRF, namespace w<h>-<t>
ip netns add w$h-$t
ip link add wv$h$t type veth peer name wp$h$t
ip link set wp$h$t netns w$h-$t; ip link set wv$h$t netns h$h
ip -n h$h link set wv$h$t master vrf-$t up
ip -n h$h addr add 10.0.$h.1/24 dev wv$h$t
ip -n w$h-$t addr add 10.0.$h.2/24 dev wp$h$t; ip -n w$h-$t link set wp$h$t up; ip -n w$h-$t link set lo up
ip -n w$h-$t route add default via 10.0.$h.1
# STAND-IN FOR EVPN type-5: the other host's subnet, reachable via its router IP, installed in this tenant's VRF only.
# onlink: accept the next hop as directly reachable on br-<tenant> even though no address is configured on that subnet.
ip -n h$h route add 10.0.$r.0/24 vrf vrf-$t via 10.255.$n.$r dev br-$t onlink
done
}
tenant_setup A 5001 1001 02:00:00:00:0a:01 02:00:00:00:0a:02 1
tenant_setup B 5002 1002 02:00:00:00:0b:01 02:00:00:00:0b:02 2
for h in h1 h2; do ip netns exec $h bash -c 'echo 1 > /proc/sys/net/ipv4/ip_forward'; done
# tenant B alone owns a second subnet on h2
ip -n h2 addr add 10.9.9.1/24 dev wv2B
ip -n w2-B addr add 10.9.9.2/24 dev wp2B
ip -n h1 route add 10.9.9.0/24 vrf vrf-B via 10.255.2.2 dev br-B onlink
echo "--- A w1 -> A w2 (same addresses exist in tenant B)"
ip netns exec w1-A ping -c2 -W1 10.0.2.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- B w1 -> B w2"
ip netns exec w1-B ping -c2 -W1 10.0.2.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- B w1 -> 10.9.9.2 (prefix that exists only in tenant B)"
ip netns exec w1-B ping -c1 -W1 10.9.9.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- A w1 -> 10.9.9.2 (no such route in tenant A's VRF)"
ip netns exec w1-A ping -c1 -W1 10.9.9.2 2>&1 | grep -o '[0-9]* packets transmitted, [0-9]* received' || true
echo "--- full-size inner packet, DF set (1472 + 28 = 1500 bytes)"
ip netns exec w1-A ping -c1 -W1 -M do -s 1472 10.0.2.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- MTUs on h1"
for dev in u1 vx-A; do ip -n h1 -o link show $dev | awk '{sub(/@.*/, "", $2); sub(/:$/, "", $2); print $2, $4, $5}'; done
echo "--- tenant VRF route tables on h1"
echo "vrf-A:"; ip -n h1 route show vrf vrf-A
echo "vrf-B:"; ip -n h1 route show vrf vrf-B
Output:
--- A w1 -> A w2 (same addresses exist in tenant B)
2 packets transmitted, 2 received
--- B w1 -> B w2
2 packets transmitted, 2 received
--- B w1 -> 10.9.9.2 (prefix that exists only in tenant B)
1 packets transmitted, 1 received
--- A w1 -> 10.9.9.2 (no such route in tenant A's VRF)
1 packets transmitted, 0 received
--- full-size inner packet, DF set (1472 + 28 = 1500 bytes)
1 packets transmitted, 1 received
--- MTUs on h1
u1 mtu 1550
vx-A mtu 1500
--- tenant VRF route tables on h1
vrf-A:
10.0.1.0/24 dev wv1A proto kernel scope link src 10.0.1.1
10.0.2.0/24 via 10.255.1.2 dev br-A onlink
vrf-B:
10.0.1.0/24 dev wv1B proto kernel scope link src 10.0.1.1
10.0.2.0/24 via 10.255.2.2 dev br-B onlink
10.9.9.0/24 via 10.255.2.2 dev br-B onlink
How to read the script: tenant_setup runs once per tenant and builds, on each of the two hosts, a VRF, a bridge and VXLAN device carrying that tenant's VNI, and a workload namespace connected by a veth. The shell trick ${!mac} means "the value of the variable whose name is stored in mac", so on host 1 it reads m1, the argument holding host 1's router MAC; r=$((3 - h)) is arithmetic that gives the other host's number. The only steps that stand in for BGP EVPN are the neigh add ... nud permanent line (a fixed ARP entry for the remote router, never expiring), the bridge fdb add line (which tunnel destination to use for that MAC) and the final route add ... onlink line (the remote subnet, inside this tenant's VRF only; onlink accepts the next hop as directly reachable). In a real deployment EVPN advertises those three facts by itself. Everything else is the same.
Reading it: both tenants reach their own 10.0.2.2 through the same address plan, tenant A has no route to the prefix owned by tenant B, and each VRF holds only its own routes. The static entries stand in for EVPN, and the single-VXLAN-device and FRR lines above come from the FRR documentation and were not run here.
Trade-offs and pitfalls
- EVPN on the host versus a central controller. A controller-driven overlay such as OVN (Open Virtual Network, an Open vSwitch based controller that uses Geneve tunnels, a tunnel format similar in purpose to VXLAN) gives distributed firewalling and L2 features with one control point. EVPN on the host keeps the fabric on standard BGP that a network team already operates and avoids a controller as a single dependency. Recommend EVPN when the team is network-led and the model is routed (L3) per tenant; switch to the controller model if tenants need stretched L2 subnets and live migration with identical addresses.
- VRF is not a security boundary by itself. It separates routing only. A host compromise, a kernel bug or a mistaken route import defeats it, which is why the firewall and the negative probe exist.
- MTU black holes. If ICMP "fragmentation needed" is blocked in the underlay, small pings work while large transfers stall.
- Overlapping prefixes at every shared boundary (NAT, peering, shared services) need per-tenant handling, or the first two tenants who overlap collide there.
A circuit breaker is flapping: it trips every time the error rate blips to 5% for about a minute, then closes, then trips again a few minutes later. Walk through why this is probably happening and what you'd change about the breaker's configuration to fix it.
Sample Answer
Direct answer
Flapping like this, tripping on a brief 1-minute blip and closing again a few minutes later, is almost always caused by an evaluation window that's too short and a threshold that's checked as a raw instantaneous rate instead of a smoothed one, with no hysteresis between the open and closed conditions. The fix is to require the elevated error rate to persist for a sustained period before tripping (smoothing or a minimum-duration requirement), and to require a different, lower threshold sustained for a while before closing again (hysteresis), so the breaker doesn't oscillate around a single threshold value that noisy traffic keeps crossing in both directions.
Why this specific symptom happens
A breaker that evaluates a short window (say a single 10 to 30 second bucket) against a fixed absolute threshold (say 3 to 5 percent) will trip the instant any one bucket crosses that number, regardless of whether the elevated rate is a real sustained problem or a one-minute noise blip from a handful of slow requests. Once it trips and the cooldown expires, it closes again because the very next window looks normal, and then trips again a few minutes later the next time normal traffic variance happens to produce another short blip above the same threshold. The breaker isn't wrong that error rate crossed 5 percent, it's wrong that a single brief crossing is sufficient evidence of a real outage worth failing traffic away from a service that's actually healthy most of the time.
Fixes, in order of how much they change behavior
1. Smoothing the signal. Replace the raw per-window rate with an exponential moving average (EMA), so a short spike gets damped rather than immediately crossing the threshold:
EMAt=αxt+(1−α)EMAt−1With α=0.2, a baseline error rate of 1 percent, and a 1-minute burst of 5 percent sampled every 10 seconds (6 samples):
EMA1EMA2EMA3EMA4EMA5EMA6=0.2(0.05)+0.8(0.01)=0.018=0.2(0.05)+0.8(0.018)=0.0244=0.2(0.05)+0.8(0.0244)=0.02952=0.2(0.05)+0.8(0.02952)=0.033616=0.2(0.05)+0.8(0.033616)=0.036893=0.2(0.05)+0.8(0.036893)=0.039514Against a 3.5 percent trip threshold, the raw signal crosses it on sample 1 (5 percent, instantly), but the EMA doesn't cross 3.5 percent until sample 5, roughly 50 seconds into the burst. That 50-second delay is the point: it means a true 1-minute-and-done blip barely trips the breaker at all (it only just crosses right as the burst is ending), while a genuinely sustained failure keeps climbing well past threshold and trips decisively.
2. Hysteresis between open and close conditions. A breaker actually cycles through three states: closed (normal, calls flow through), open (tripped, calls are rejected outright without even trying the dependency), and half-open (a brief trial period after opening where a small number of requests are deliberately let through to test whether the dependency has actually recovered) before it's allowed back to closed. Use a different, lower threshold to close than to open, and require it sustained for a minimum duration, not a single good sample: for example, trip open at EMA > 3.5 percent sustained for 2 evaluation periods, but only close from half-open back to closed once EMA stays below 2 percent (not 3.5 percent) for 2 consecutive periods. This asymmetry is what actually stops the flap-then-immediately-reopen cycle, because closing requires meaningfully cleaner traffic than the level that caused the trip, not just traffic that's dipped fractionally under the same number.
3. Minimum sample size. Require a floor on request volume before evaluating the rate at all (for example at least 200 requests in the window); at low traffic volumes a handful of failed requests can swing the percentage wildly even though the absolute failure count is tiny, and no smoothing scheme fixes a rate computed from too small a denominator.
4. Gradual re-entry. When the breaker closes, ramp traffic back in (5 percent, then 20 percent, then full) rather than snapping straight back to 100 percent, so a dependency that's only marginally recovered doesn't get immediately re-tripped by the full traffic load the instant it reopens.
Trade-offs & pitfalls
Every one of these fixes trades detection speed for stability: a smoothed, hysteresis-gated breaker takes longer to trip on a real outage than a naive instant-threshold one, which is the correct trade for a dependency where false trips are expensive (unnecessary failover, alert fatigue), but the wrong trade for something where even a few seconds of cascading failure is unacceptable, so the constants here (alpha, thresholds, minimum duration) should be tuned against the actual cost asymmetry for that specific dependency, not copied from another service. A subtler pitfall is that smoothing and retries interact: if callers retry failed requests, each retry counts as an additional data point in the error-rate window, so a retry storm during a real degraded period can itself inflate the EMA further and trip the breaker faster than the underlying failure rate alone would justify, which is usually the desired outcome (retries are evidence something is actually wrong) but is worth knowing explicitly rather than discovering by surprise. Validating any change to these constants should happen by replaying real historical traffic traces (including past incidents and known-noisy periods) through the new policy offline, comparing false-trip rate and time-to-detect-real-outage against the old policy, before rolling the new thresholds out as a canary.
Write a tcpdump/BPF filter that captures TCP packets destined to port 443 that contain a TLS "ClientHello" record (TLS record type 0x16) in the first TLS record. Provide the BPF filter expression and explain why it works (explain offsets for IP and TCP headers and the TLS record header check).
Sample Answer
Approach: match on destination port first as a cheap precondition, then reach past the fixed Ethernet/IP/TCP headers into the first byte of the TLS record layer to check its record-type value.
tcp dst port 443 and (tcp[((tcp[12:1] & 0xf0) >> 2)] = 0x16)
This compiles to valid BPF bytecode against a real TCP/IP packet, verified directly with tcpdump -d.
Explanation of the offsets
tcp dst port 443: a standard port match, and also a cheap precondition guaranteeing the packet has a full TCP header before any byte-offset math runs against it. Note what this filter does not need to do manually: because it uses thetcp[]protocol qualifier rather than rawether[]/ip[]byte offsets, BPF has already located the start of the TCP header for us, skipping past the Ethernet header and the IP header (itself variable-length if IP options are present) automatically. Every offset written below is relative to that TCP header start, not to the start of the packet.tcp[12:1]: the single byte at offset 12 into the TCP header. In the TCP header layout, the upper nibble of byte 12 is the Data Offset field, the TCP header's own length measured in 32-bit words.& 0xf0isolates that upper nibble (masking off the lower 4 bits, which belong to reserved/flag bits).>> 2then converts "number of 4-byte words, already positioned in the top nibble of the byte" into "header length in bytes" directly, since a byte's top nibble shifted right by 2 is arithmetically equivalent to the word count multiplied by 4. To see why a right-shift, which normally divides, ends up scaling by 4 here: because the Data Offset value sits in the byte's top nibble, its raw value is already the word count times 16 (occupying the high 4 bits multiplies the underlying number by 2^4 = 16). The>> 2then divides that by 4 (2^2), leaving the word count times 4; and since each 32-bit word is 4 bytes, word count times 4 is exactly the header length in bytes. The shift really is dividing, it simply starts from a value that was 16 times too large.tcp[((tcp[12:1] & 0xf0) >> 2)]therefore reads the byte located exactly at "TCP header length in bytes" past the start of the TCP header, which is the first byte of whatever follows it: here, the TLS record layer.= 0x16checks that byte equals0x16, the TLS record type value for Handshake, the record type that carries a ClientHello (and every other TLS handshake message).
Edge cases
- TCP options. The header is often longer than the minimum 20 bytes because of options like timestamps or selective acknowledgment. That's exactly why the filter computes the offset dynamically from the Data Offset field instead of hardcoding a fixed byte position; a hardcoded offset would silently miss any connection using TCP options, which is most modern traffic.
- Only matches the first record boundary. This filter checks the very first TLS record right after the TCP header, so it correctly matches a ClientHello sent as the first bytes of the first segment, but it will not match one whose relevant byte falls at a different offset because of prior segmentation. It also cannot distinguish a ClientHello from a ServerHello or any other handshake sub-type, since all handshake messages share record type
0x16; disambiguating specifically requires checking one further byte in (0x01for ClientHello), which most simple BPF filters skip because deeply nested offset math becomes awkward in classic BPF's limited instruction set. - IPv6 is a hard limitation, not just a caveat.
tcpdump's own documentation states plainly that arithmetic expressions against transport-layer headers, liketcp[n], only work against IPv4 packets and do not work against IPv6 at all. This was confirmed empirically as well: the same filter compiled and tested against a synthetic IPv6 packet matched nothing. Filtering TLS records over IPv6 requires a different approach entirely, usingip6 protochain 6to locate the TCP header past IPv6's variable-length extension-header chain, combined with manually computed byte offsets from a fixed link-layer base rather than thetcp[]shorthand used here.
You've just joined a new team and inherited a technical system that's messy or poorly understood: undocumented infrastructure, no CI, an unfamiliar codebase, or an unclear architecture. As the new owner, outline your first 30/60/90-day plan: how you'll learn the system and assess risk, the quick wins you'll ship early to build trust, the medium-term fixes you'll drive, and how you'll know you're making real progress without destabilizing production.
Sample Answer
Direct answer
Inheriting a messy, poorly understood system as its new owner calls for a 30/60/90-day plan with a deliberate shape: the first 30 days are almost entirely learning and risk assessment with no risky changes, the next 30 add a small number of low-risk quick wins that build trust while you're still learning, and the final 30 start the real medium-term fixes, with progress measured the whole time by a small set of health signals you define early, not by how much code has changed.
Structured elaboration
Days 1 to 30, learn and assess risk: map what actually exists (services, dependencies, who else touches this system), identify where undocumented behavior or missing tests create the biggest blind spots, and deliberately make no risky changes yet, the goal is an accurate risk map, not visible progress. Talk to whoever used to own it or adjacent teams that depend on it, since institutional knowledge not written down anywhere is exactly the risk this phase exists to surface.
Days 30 to 60, quick wins: pick 2 to 3 small, genuinely low-risk fixes discovered during the learning phase, ideally things that are visibly annoying to the team or reduce a clear, contained risk, such as a missing alert on an already-known failure mode or a manual step that's easy and safe to automate. These build trust and credibility precisely because they're small enough to be confident about, not because they're impressive.
Days 60 to 90, medium-term fixes: start the larger, riskier work the learning phase identified as actually mattering, such as adding real test coverage to the highest-risk undocumented path, or setting up CI (continuous integration, automated build-and-test on every change) if none exists. Sequence it so the riskiest changes happen with the safety net of whatever quick wins already shipped (monitoring, a rollback path), rather than being the very first thing touched.
Measuring real progress without destabilizing production: track a small number of concrete health signals defined in the first 30 days, an error or incident rate baseline, and a rough measure of how much of the system now has any test coverage versus none, and use those same signals throughout, not a shifting definition of progress. Real progress means those numbers move in the right direction while the incident rate does not get worse, not just that visible work is happening.
Worked example
In week 1, discover the system has no CI and three services with no documented owner besides "the previous person," and that on-call has been fielding an average of 4 pages a week related to one specific, undocumented failure mode. That failure mode and its frequency becomes the baseline health signal to track. Days 30 to 60, the first quick win is adding an alert with a documented runbook for that exact failure mode, not fixing its root cause yet, just making it visible and handled faster when it fires, which is low-risk (adding observability, not changing behavior) and immediately useful to whoever's on-call. Days 60 to 90, the medium-term fix addresses the failure mode's actual root cause, now done with the new alert already in place as a safety net if the fix doesn't fully resolve it. Progress at day 90 is measured against the same baseline: page frequency for that failure mode should be visibly down from the original 4-a-week baseline, not just "we shipped a fix."
Trade-offs and pitfalls
Trying to fix real problems in the first 30 days, before you actually understand the system, is the most common way a new owner destabilizes production early and burns the trust they were trying to build. The opposite mistake, spending 60 or 90 days purely learning with nothing shipped, reads as inaction to a team that needs to see the new owner is capable, the quick-wins phase exists specifically to avoid that trap. Picking a "quick win" that turns out to not be low-risk, because the learning phase missed a dependency, is the sharpest version of this failure, which is why the quick wins should specifically come from what the first 30 days actually surfaced, not from a generic list of common fixes applied without local context. And measuring progress by activity (commits, pull requests merged) instead of the health signals defined up front lets real problems keep degrading even while it looks like a lot is happening.
Design the WAN for 200 branches that need access to several public clouds and to central sites. How do you decide between local breakout and backhaul, where does central policy live, and how are paths selected?
Sample Answer
Direct answer
Use local breakout (the branch sends internet and SaaS traffic, meaning software you reach over the internet such as email or a CRM, straight out its own internet link) for trusted traffic, an encrypted overlay (virtual tunnels built across ordinary internet or carrier circuits, so the WAN behaves like a private network) to regional hubs (larger sites where branch tunnels terminate) and cloud gateways for private applications, and central backhaul (hauling a branch's traffic across the WAN to a central site before it goes anywhere else) only for traffic that must be inspected centrally. Central policy lives in a controller (the management system that pushes templates, segments and application rules) while each branch router (the CPE, customer premises equipment, the device at the branch end of the WAN) keeps forwarding if the controller is unreachable. Paths are chosen per application from live loss, latency and jitter (how much the delay varies from packet to packet) measurements on every tunnel.
What a new branch does first: it powers on and reaches the controller over any internet link, pulls its template, builds an encrypted tunnel on each transport to both hubs, starts a probe on every tunnel, and only then begins steering applications onto the best measured path.
Sizing the problem
Assumptions (replace with measurements): 200 branches, 40 users each, 1.5 Mbps per user in the busy hour, so 60 Mbps per branch and 12 Gbps in total. Assume 70% of that is internet and SaaS.
- Backhaul everything: 12 Gbps through the hubs.
- Breakout for internet and SaaS: only 30% reaches the hubs, 3.6 Gbps. With two active hubs that is 1.8 Gbps each normally and 3.6 Gbps on one hub after the other fails, which is 36% of a 10G hub link.
- Tunnel count: a full mesh of 200 branches needs 200 x 199 / 2 = 19,900 tunnels. Hub-and-spoke with 2 hubs and 2 transports per branch (broadband plus LTE or a second broadband) needs 200 x 2 x 2 = 800 tunnels, 400 per hub.
- Latency penalty for backhaul: a branch 35 ms RTT from its hub, with the hub 8 ms from the SaaS provider, takes 43 ms compared with 12 ms direct. That is 31 ms more per round trip, and a page that needs 10 sequential round trips is 310 ms slower.
Breakout or backhaul: the decision rule
| Traffic class | Path | Reason |
|---|---|---|
| Trusted SaaS and general internet | Local breakout with a local firewall or cloud-delivered security | Avoids the detour and hub bandwidth |
| Private apps in clouds and data centers | Overlay tunnel to the nearest hub or cloud gateway | Never touches the internet |
| Sensitive or audited flows | Backhaul to central inspection | Inspection capacity is central |
| Voice and video | Best measured path, breakout if the SaaS endpoint allows | Latency and jitter matter most |
Where policy lives and how paths are selected
- Controller (two or three instances, in different locations): holds templates, segmentation (VRFs, virtual routing tables, one per business unit so their routes never mix) and application policy. Branches pull config; if the controller is down, the data plane (the part that actually forwards packets, as opposed to the control part that decides policy) keeps running on the last config and only changes are blocked.
- Probes: every tunnel carries a probe, 800 tunnels at 1 probe per second is 800 packets per second in total, which is negligible. Each probe measures loss, latency and jitter.
- Application policy: a class maps to an SLA, for example voice must stay under 1% loss, 150 ms latency and 30 ms jitter (assumed thresholds, tune to your codec). If the current path breaks the SLA, the application moves to the next path. Add a hold-down (a wait period, for example several minutes of good measurements, before moving back to the original path) so it does not flap back at once.
- Liveness: tunnel BFD (bidirectional forwarding detection, a fast hello) at 300 ms with a multiplier of 3 detects loss in 0.9 s (interval times multiplier, per RFC 5880).
BGP route policy at hubs and cloud gateways
BGP (the border routing protocol, which exchanges address prefixes between networks) runs only at the hubs and cloud gateways, not at all 200 branches. Of the three policies below, the first decides which hub carries traffic; the other two are guardrails and a maintenance tool. Branch prefixes are allocated from one block so a summary is possible: branch i gets 10.64.i.0/24 for i from 0 to 199, all inside 10.64.0.0/16, which holds 256 /24s so 56 remain free. The single /16 can go to neighbours that need no per-branch detail, but the cloud gateways and the other hub receive the individual /24s (200 routes, a trivial load). A hub that kept advertising the /16 after losing its tunnel to branch 7 would keep attracting branch 7's traffic and drop it (a black hole), whereas a withdrawn /24 lets traffic move to the other hub.
- Prefer the primary hub with a higher local preference (200 versus 100). Local preference is a number a router attaches to a route it learns, and when two routes reach the same prefix the higher number wins (RFC 4271). Tag routes with a community per region (a community is a label attached to a route that other routers can match on, for example to apply a regional policy). Traced example: a cloud gateway learns 10.64.7.0/24 from hub-1 and from hub-2. Its import policy sets local preference 200 on the hub-1 copy and 100 on the hub-2 copy, so it sends traffic for branch 7 through hub-1 and keeps hub-2 as the standby.
- Filter what you accept from each neighbour to the prefixes it owns, and set a maximum-prefix limit (a cap on how many routes a neighbour may send; past it the session is shut) so a misconfigured neighbour cannot flood the table (RFC 7454 recommends both).
- To drain a hub for maintenance, send the GRACEFUL_SHUTDOWN community (65535:0), a well-known label meaning "this path is about to go away". Receivers lower the local preference of those paths (RFC 8326 recommends 0), so in the example above hub-1's routes drop from 200 to 0, hub-2 wins, and traffic moves before the session drops.
Failover and monitoring at 100+ sites
Each branch has two transports and tunnels to two hubs. Losing one transport is a path change, losing a hub moves traffic to the second hub, both within the BFD detection time plus route recomputation. Export flow records and tunnel measurements from every branch to a central collector and alarm on SLA breaches per site, not per link.
Trade-offs and pitfalls
- Breakout moves the security perimeter to every branch. Without a firewall or cloud-delivered inspection at the branch you have 200 internet edges.
- If both hubs sit behind the same provider, a provider event takes both. Place them on different providers.
- A lost or stolen branch router must be revocable from the controller the same day, so give every device its own credential and rotate keys on a schedule.
- Hubs are sized for the failure case (3.6 Gbps on one hub), not the normal case.
As a Solutions Architect, compare dedicated interconnect options (e.g., AWS Direct Connect, Azure ExpressRoute, GCP Dedicated Interconnect) versus internet-based VPNs for hybrid connectivity. Discuss throughput, latency, SLA/predictability, security, operational complexity, and cost trade-offs. Provide guidance on which to choose for predictable high-volume data versus low-volume ad-hoc traffic.
Sample Answer
Direct answer
Choose dedicated interconnect (a private, physical circuit connecting your site directly into the cloud provider's network, bypassing the public internet), such as AWS Direct Connect, Azure ExpressRoute, or GCP Dedicated Interconnect, for predictable, high-volume, latency-sensitive traffic, and internet-based site-to-site VPN for low-volume or ad-hoc traffic where the fixed cost of a dedicated circuit is not justified. The crossover point is a straightforward cost calculation once both options' actual pricing is known, not a rule of thumb, because dedicated interconnect trades a fixed monthly cost for a lower per-gigabyte rate, and at low enough volume that trade loses.
Structured elaboration
Comparison table
| Dimension | Dedicated interconnect | Internet VPN |
|---|---|---|
| Throughput | Fixed port speeds, commonly 1, 10, or 100 Gbps, consistently available up to the provisioned speed | Bounded by the VPN gateway's own capacity and by the variable quality of the underlying internet path; effective throughput can be well below the gateway's rated maximum under congestion |
| Latency | Consistent and typically lower, since traffic does not traverse the public internet's variable routing | Variable, subject to internet routing changes, congestion, and the specific ISPs in the path |
| SLA (service-level agreement, a contractual guarantee on performance or uptime) and predictability | Provider-backed SLA covering the dedicated circuit's availability and performance | No SLA over the underlying public internet path itself; only the VPN gateway endpoints carry a provider SLA |
| Security | Traffic stays off the public internet by construction, still typically layered with encryption for defense in depth | Encrypted via IPsec (a protocol suite that encrypts and authenticates traffic between two network gateways) by design, but traverses the public internet, meaning the security posture depends entirely on the encryption, not on path isolation |
| Operational complexity | Longer provisioning lead time, often weeks, involving a carrier or colocation partner, but simpler to operate day to day once live | Fast to stand up, often hours, but production behavior needs more active monitoring since the underlying path quality is not guaranteed |
| Cost shape | Fixed monthly cost, port and circuit fees, plus a lower per-gigabyte data-transfer rate | No fixed circuit cost; a standard, typically higher, per-gigabyte data-transfer or egress rate, usage-based |
Cost-factor checklist for dedicated interconnect
When pricing an interconnect option, name every line item rather than only the headline port fee: the port allocation cost for reserving the physical port capacity itself, the monthly port charge billed regardless of how much of it is used, the data-transfer or egress rate for traffic actually sent over the circuit, usually discounted relative to standard internet egress but not free, and cross-connect fees charged by the colocation facility for the physical cable connecting your equipment to the cloud provider's, separate from anything the cloud provider itself bills. Missing any one of these line items when estimating cost is the most common reason a dedicated interconnect ends up more expensive than projected.
Worked cost comparison and the crossover point
Using representative, illustrative rates, confirm current list prices with the specific provider before budgeting: an internet VPN's data transfer runs at roughly $0.08 per GB in this example, while a dedicated interconnect carries a $300 monthly port fee plus a discounted $0.02 per GB transfer rate. Setting the two monthly costs equal, 0.08 times GB equals 300 plus 0.02 times GB, gives 0.06 times GB equals 300, so GB equals 5,000, or 5 TB per month. Below 5 TB per month, the VPN is cheaper because the interconnect's fixed port fee has not been earned back yet; above 5 TB per month sustained, the interconnect is cheaper, and the gap widens linearly with volume. The actual decision rule is to compute this crossover with real quoted numbers for the specific volume in question, rather than defaulting to either option by habit.
Three concrete usage scenarios
- Ad-hoc, low-volume: a quarterly batch export of a few hundred gigabytes to a partner's cloud environment. Internet VPN is the right call, since standing up a dedicated circuit for traffic that runs a few days a quarter would sit almost entirely idle, paying the fixed port fee for capacity that is rarely used.
- Predictable, high-volume: continuous database replication moving several terabytes per day between an on-prem data center and a cloud region. Dedicated interconnect is the right call, since this is exactly the sustained-volume, latency-sensitive case where the fixed port fee is earned back quickly, well above the 5 TB per month crossover in the example above, and the consistent latency matters for keeping replication lag bounded.
- Latency-sensitive but moderate-volume: a hybrid application where a subset of transactions must complete within a tight latency budget, but total data volume is moderate. Here the decision is not purely cost-driven; even below the cost crossover point, the SLA and latency predictability of a dedicated interconnect may be worth paying for if a VPN's variable latency risks violating the application's own latency requirement, a case where the guaranteed-performance argument overrides the raw cost comparison.
Worked example
A company evaluates options for a workload projected at 8 TB per month of sustained replication traffic. Over the internet VPN at $0.08 per GB, that is 8,000 times $0.08, or $640 per month. Over the dedicated interconnect, $300 plus 8,000 times $0.02, or $300 plus $160, equals $460 per month, cheaper by $180 per month and widening every month volume grows, on top of the latency and SLA benefits, making the interconnect the clear recommendation for this specific volume.
Trade-offs and pitfalls
The most common mistake is comparing only the headline per-gigabyte rate or only the fixed port fee in isolation, instead of the total monthly cost at the workload's actual projected volume, which the crossover calculation above is built specifically to avoid. A second is choosing internet VPN for a genuinely latency-sensitive workload purely because it is cheaper at the current volume, without weighing scenario 3's point that SLA and predictability sometimes justify the interconnect below the cost crossover. A third is forgetting the cross-connect and port-allocation fees when budgeting an interconnect, only to find the actual monthly bill higher than the headline port-fee-plus-data-rate estimate suggested.
Compare host-based agent microsegmentation with network-based approaches (software-defined networking, VLANs, next-gen firewalls). Discuss security effectiveness, deployment complexity, policy expressiveness, and how well each approach handles encrypted east-west traffic in a hybrid environment.
Sample Answer
Host-based agent microsegmentation enforces policy from inside the workload itself, so it sees encrypted east-west traffic (service-to-service traffic between workloads, as opposed to north-south traffic between users and services) the same way it sees anything else, because the agent sits at the endpoint of the encryption, not in the middle of it. Network-based approaches, software-defined networking (SDN), VLANs, next-generation firewalls (NGFW), enforce from network infrastructure, which loses most of its policy-relevant visibility once traffic is encrypted with mutual TLS (mTLS, where both sides cryptographically prove their identity) unless the device terminates and re-encrypts the connection, a costly move that reintroduces the "trusted intermediary" model zero trust is trying to remove.
Comparing the two approaches
| Axis | Host-based agent | Network-based (SDN, VLAN, NGFW) |
|---|---|---|
| Security effectiveness | Policy tied to actual workload identity; survives IP changes and workload movement | Policy tied to network location (IP, subnet, VLAN); brittle when workloads are ephemeral or move |
| Deployment complexity | Needs an agent or kernel hook on every host or workload; harder in heterogeneous or legacy fleets | Centralized on network devices; no per-host rollout, but requires traffic to actually route through the enforcement point |
| Policy expressiveness | Can reference workload, process, or identity attributes directly | Mostly limited to network- and transport-layer attributes unless combined with deep packet inspection, which breaks on encrypted traffic |
| Encrypted east-west traffic | Naturally compatible; the agent sits at the endpoint before encryption or after decryption | Falls back to metadata-only visibility, or requires terminating TLS in the middle |
Worked example
Two workloads communicate over mTLS. A host-based agent on each side can enforce "workload A may call workload B on this specific method," because it evaluates policy where the plaintext request is still visible to the local process, even though the wire traffic between the hosts is fully encrypted. A network-based NGFW sitting between them, by contrast, sees only an encrypted TLS stream between two IP addresses on a port; without terminating and re-encrypting the connection, it can only enforce "IP A may talk to IP B on this port," a coarser and more brittle rule, especially once workloads get rescheduled to new IP addresses, which happens constantly with containers.
Trade-offs and pitfalls
Host-based agents require an install and maintenance footprint on every workload type in a hybrid environment, VMs, containers, and, hardest of all, legacy or appliance-style bare-metal systems that may not support running arbitrary agents, so pure host-based coverage is rarely complete in a real heterogeneous estate. Network-based controls remain the fallback for whatever can't run an agent. A mature design usually layers both: identity-based enforcement wherever an agent can run, and coarser network-based segmentation as a second, less precise safety net for everything else.
Two of your top customers want mutually exclusive behaviors from the product, and both say they will leave if they do not get theirs. How do you decide?
Sample Answer
Direct answer
I would not decide on revenue size or on who is louder. First I find out whether the two demands are truly exclusive or only the two requested solutions are, by getting to the goal behind each request. If the goals can both be met, I design for both. If they truly cannot, I pick the customer that fits the product we are building, tell both of them the decision and the reasoning before it ships, and plan for the loss we have chosen to risk.
Step 1: Test the ultimatum
"We will leave" is a claim, not data. Check the renewal date, how deeply the customer uses the product (active users, integrations they built), what alternatives they have and what switching would cost them. A customer two months from renewal with a competitor pilot is a different situation from one with eighteen months left on contract. Then ask each customer: what are you trying to get done when this behavior matters, and what happens if you do not get it?
Step 2: Move from the behavior up to the goal (illustrative example)
- Customer A, a finance firm, demands that edits to shared dashboards need manager approval before going live.
- Customer B, a media company, demands that analysts publish edits instantly.
- Asking why: A needs auditability (to show an auditor who changed what, and when). B needs speed during breaking news. Neither wants the other's world.
- Design that serves both: a workspace-level setting for publish mode (approval required or instant), with the edit history always on. A gets approvals plus the audit trail. B gets instant publishing, with after-the-fact review available.
- Cost: two modes to test and support and one more settings screen. Compare that with losing both accounts.
Step 3: If the behaviors are truly exclusive, decide on criteria
| Question | Why it matters |
|---|---|
| Which customer is closer to the ideal customer profile (ICP, the kind of customer the company is built to serve best)? | Strategy, not size, should steer the product |
| What do other customers in that segment want? | Two voices are not a market; check request counts and usage data |
| Which option is cheaper to build and easier to reverse? | Reversible choices can be made faster |
| Which keeps the product coherent? | A product that does two contradictory things tends to do neither well |
| Revenue and strategic value (reference customer, expansion) | A tiebreaker, not the decider |
My rule: choose on ICP fit and breadth of demand, and use revenue only to break ties.
Step 4: Communicate
Before any announcement, the product lead and the account's executive sponsor call each customer. Say what we heard, what we decided, why, and what we offer instead (a workaround, an API or integration, a date to revisit). Assume the two customers may compare notes, so both hear the same reasoning.
Pitfalls
- Splitting the difference into a halfway behavior that satisfies neither.
- Building two forks of the product.
- Letting one ultimatum teach every customer that ultimatums work.
What would flip my call
If the setting is cheap, build both. If one customer sits outside the ICP, accept losing them. If usage data shows the "we will leave" is a bargaining position, I treat it as a preference, not a deadline.
Explain TCP Selective Acknowledgment (SACK): how SACK blocks are represented in the TCP options, and how SACK lets a sender avoid retransmitting segments the receiver already has after a single loss event. What does a sender do differently once SACK is enabled versus a sender using only cumulative ACKs?
Sample Answer
Direct answer
Selective Acknowledgment (SACK) lets a receiver tell the sender exactly which non-contiguous blocks of data it has ALREADY received, so after a loss the sender only has to retransmit the specific missing segment(s), not everything that came after it.
Structured elaboration
Without SACK, TCP uses cumulative acknowledgment: an ACK only confirms "I have received everything up through this byte, contiguously." If segment 3 of a 10-segment flight is lost but segments 4 through 10 all arrive fine, the receiver can only ACK up through the end of segment 2, it has no way to tell the sender "I actually already have 4 through 10, I'm just missing 3." A sender using only cumulative ACKs, upon detecting the loss, may end up retransmitting segments 3 through 10 (everything the receiver hasn't cumulatively acknowledged), even though 4 through 10 were never actually lost.
With SACK enabled (negotiated via a permitted option in the handshake, then carried on subsequent ACKs), the receiver's ACK can include SACK blocks, explicit ranges of sequence numbers it holds that are NOT contiguous with the main acknowledged run, in the example above, a SACK block spanning segments 4 through 10. Now the sender knows precisely that only segment 3 needs retransmitting.
Worked example
Say a sender has segments with sequence ranges [1000-1500), [1500-2000), [2000-2500), ... up to [4500-5000), and segment [2000-2500) is lost in transit while everything else arrives. Without SACK: the receiver's ACKs stay pinned at ack=2000 (the last contiguous byte received) even as segments up through 5000 keep arriving; the sender, upon detecting the loss (via duplicate ACKs all saying ack=2000), knows only that SOMETHING after 2000 needs resending and, in older/naive implementations, could resend everything from 2000 onward. With SACK: the same duplicate ACKs at ack=2000 now also carry a SACK block like sack=2500-5000, telling the sender explicitly that only the single segment [2000-2500) is actually missing, so it retransmits exactly that one segment and nothing else.
Trade-offs & pitfalls
SACK is most valuable on connections with a large amount of data in flight (a large window relative to segment size) and where losses are isolated rather than in a solid burst, since that's exactly the scenario where "retransmit everything after the gap" wastes the most bandwidth compared to "retransmit only the gap." On a connection with a tiny window, or where an entire flight is lost at once (nothing left to selectively acknowledge), SACK provides little advantage.
Explain how you would evaluate and improve the team's incident runbooks and onboarding process so new hires become operationally effective faster. Describe specific measurement approaches (e.g., time-to-first-response), sample runbook changes that deliver value, and how you would enforce runbook ownership and regular updates.
Sample Answer
Approach summary
I’d treat runbooks and onboarding as measurable products: define metrics, iterate on content, and enforce ownership so new hires reach operational independence faster.
Measurement approaches
- Time-to-first-response (TTFR): time from alert to first actionable step taken by on-call — measures discoverability and clarity.
- Time-to-mitigation (TTM) and Mean Time to Resolve (MTTR): show runbook effectiveness.
- Ramp-to-autonomy: weeks until hire can independently handle N class-1 and M class-2 incidents in simulation.
- Runbook usage and success rate: % of incidents where runbook was followed and outcome (helpful / incomplete).
- Drill performance: score on simulated incidents using runbooks.
Sample runbook changes
- TL;DR at top: symptom checklist, impact, severity decision matrix.
- Clear pre-reqs: login, sudo, VPN, jumpbox, device IPs.
- Step-by-step playbook with exact CLI commands and example outputs; include copy-paste commands and command safety notes.
- Decision tree / flowchart for branching paths (e.g., link flap vs. routing issue).
- Telemetry links: direct links to dashboards, packet captures, show commands, and relevant RFCs.
- Rollback and escalation steps with pager names/RCA owner.
- Tagging and search metadata (device-type, vendor, service).
- Small “one-pager” cheat-sheets for common tasks (BGP reset, ACL checks).
Enforcing ownership & updates
- Assign a runbook owner (primary + secondary) per service/segment in team roster; owners are responsible for accuracy and runbook PRs.
- Integrate runbook updates into change control: any network change requires runbook validation/PR before deployment.
- Quarterly review cadence enforced by calendar invite; owners present evidence (drill results, recent incidents).
- Automate lints: CI checks that runbooks contain required sections (TL;DR, pre-reqs, commands, telemetry links).
- Include runbook-driven drills in onboarding: new hires must complete simulated incidents with >80% score before independent on-call.
- Recognition: include runbook quality in performance reviews and on-call postmortems to close the loop.
This combination makes runbooks actionable, measurable, and owned — reducing TTFR and ramp time while improving network reliability.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Network Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs