Apple Network Engineer (Mid-Level) Interview Preparation Guide
Apple's network engineer interview process typically follows a structured evaluation path designed to assess technical depth, architectural thinking, hands-on troubleshooting ability, and cultural fit. For mid-level candidates, the process balances assessment of independent technical competency with the ability to design and own medium-sized infrastructure projects. The interview loop includes initial recruiter screening, 1-2 technical phone screens focusing on networking fundamentals and problem-solving, and 4-5 onsite rounds covering technical depth, system architecture, real-world troubleshooting scenarios, and behavioral assessment.
Interview Rounds
Recruiter Screening
What to Expect
Initial conversation with Apple recruiter to assess background, motivation, role fit, and logistics. Recruiter will discuss your experience with network design, infrastructure projects, and familiarity with enterprise or cloud-scale networking. This round also covers location flexibility, timeline, and team alignment. The recruiter will evaluate your communication skills and ability to articulate technical accomplishments in business terms.
Tips & Advice
Frame your experience around ownership and impact. Rather than listing tools used, explain the problems you solved and outcomes achieved (e.g., 'I redesigned the WAN to reduce latency by 30%, improving application performance for 50,000 employees'). Be clear about your motivation for the role and what attracts you to Apple's infrastructure challenges. Ask thoughtful questions about the team structure, current infrastructure priorities, and growth areas. Demonstrate familiarity with Apple's business (product scale, geographic distribution, privacy commitments) to show genuine interest.
Focus Topics
Familiarity with Apple's Scale and Challenges
Demonstrate awareness of Apple's infrastructure footprint: billions of devices, global data centers, cloud scale, privacy-first architecture, retail networks, and supply chain systems. Reference specific infrastructure challenges or design patterns relevant to Apple's business.
Practice Interview
Study Questions
Motivation for Apple and Infrastructure Interest
Explain why you're interested in Apple specifically, what attracts you to their infrastructure challenges, and how this role aligns with your career goals. Reference Apple's scale, product ecosystem, or engineering culture.
Practice Interview
Study Questions
Communication of Technical Impact
Practice translating technical work into business impact. Prepare 2-3 examples of infrastructure projects where you can articulate the technical challenge, your approach, and measurable outcome (cost reduction, performance improvement, security enhancement, reliability gains).
Practice Interview
Study Questions
Career Progression and Network Engineering Background
Articulate your journey in network engineering, key projects, and progression from junior to mid-level responsibilities. Explain how you've grown from support roles to owning architecture and design decisions.
Practice Interview
Study Questions
Technical Phone Screen 1: Networking Fundamentals and Troubleshooting
What to Expect
First technical phone screen conducted by a senior network engineer or engineering manager. This round assesses your foundational knowledge of networking protocols, your troubleshooting methodology, and ability to think through network problems systematically. Expect detailed questions about layer 2/3 protocols, routing, switching, network troubleshooting tools, and a real-world scenario where you must diagnose a connectivity or performance issue. Questions are designed to evaluate depth of practical experience and your ability to apply networking fundamentals to solve real problems.
Tips & Advice
Be precise with terminology and avoid hand-waving explanations. For each technical question, explain not just the 'what' but the 'why'—demonstrate conceptual understanding, not just memorization. When presented with a troubleshooting scenario, walk through your diagnostic approach step-by-step: gather information (show output of diagnostic commands), form hypotheses, eliminate possibilities, and root-cause analysis. Draw diagrams or describe network topology clearly. If you don't know something, say so directly but then discuss how you'd investigate it. Mid-level candidates should demonstrate independence in troubleshooting, not reliance on vendor support or colleagues. Prepare real examples from your background; interviewers will ask follow-up questions to verify hands-on experience.
Focus Topics
Network Security Fundamentals (Firewalls, ACLs, DDoS Mitigation)
Firewall design (stateful vs. stateless), access control lists, network segmentation for security, understanding common attacks (DDoS, reconnaissance, spoofing), rate limiting, and security best practices. Not deep cryptography, but infrastructure-level security.
Practice Interview
Study Questions
Network Performance and Capacity Planning
Understanding network metrics (throughput, latency, jitter, packet loss), tools for performance measurement, capacity planning methodology, identifying bottlenecks, QoS design, and traffic engineering. How to analyze network behavior under load.
Practice Interview
Study Questions
IP Addressing, Subnetting, and IPv6
Mastery of IPv4 subnetting, CIDR notation, address aggregation, IPv6 addressing architecture, IPv6 routing, and migration strategies. Understanding private vs. public address space, NAT implications, and address planning for growth.
Practice Interview
Study Questions
Switching, VLAN Design, and Network Segmentation
Spanning tree protocol, VLAN design patterns, trunk configuration, switch security (port security, DHCP snooping, DAI), redundancy design, and network segmentation strategies. Understanding when and how to segment networks for security and performance.
Practice Interview
Study Questions
Network Troubleshooting Methodology and Tools
Systematic troubleshooting approach: defining the problem scope, gathering evidence (ping, traceroute, tcpdump, netstat, show commands), hypothesis formation, and root cause analysis. Familiarity with packet capture analysis, syslog analysis, SNMP monitoring, and understanding common failure modes (black hole routing, MTU issues, asymmetric paths, congestion).
Practice Interview
Study Questions
OSI Model and Routing Protocols (BGP, OSPF, EIGRP)
Deep understanding of routing fundamentals, protocol selection for different network scenarios, BGP attributes and route selection, OSPF design (areas, costs, convergence), multipath routing, and route summarization. Be ready to explain why you'd choose one protocol over another.
Practice Interview
Study Questions
Technical Phone Screen 2: Network Design and Architecture
What to Expect
Second technical phone screen with a senior infrastructure architect or network engineering lead. This round shifts focus from troubleshooting individual problems to designing and architecting network solutions. You'll be presented with a network design scenario (e.g., 'Design a network for a new data center', 'Design WAN for a global company', 'Design a network for a hybrid cloud environment') and expected to walk through your design thinking: requirements gathering, technology selection, scalability, redundancy, security, and operational considerations. This assesses your ability to think at the architecture level, a key mid-level responsibility.
Tips & Advice
Start by clarifying requirements and constraints before designing. Ask questions: What's the scale? What's the budget? What are availability requirements? Are there security constraints? This shows maturity. Then walk through your design systematically: topology, technology choices, redundancy strategy, security layers, monitoring approach, and operational model. Explain tradeoffs—no design is perfect; discuss what you're optimizing for (cost, performance, simplicity, security) and what you're accepting as tradeoffs. Use diagrams (draw on paper or describe clearly). For mid-level, focus on practical design rather than bleeding-edge technology. Show that you consider operational aspects: monitoring, troubleshooting, change management, disaster recovery. Reference specific technologies you've worked with, but also explain why you'd choose alternatives if this were a different environment.
Focus Topics
Network Monitoring, Observability, and Operations Design
Designing monitoring and alerting strategies, baselines and anomaly detection, SNMP/NetFlow analysis, syslog aggregation, designing for operational visibility, and tooling selection. How to operationalize a network at scale.
Practice Interview
Study Questions
WAN Design and Hybrid Cloud Networking
Wide-area network design for distributed organizations, SD-WAN concepts, multipath routing for WAN, cloud connectivity (Direct Connect, ExpressRoute, Interconnect), hybrid cloud network design, and optimization for application performance across geographies.
Practice Interview
Study Questions
Cloud Networking (AWS, GCP, Azure) and Hybrid Scenarios
Virtual private clouds, subnet design, security groups/NACLs, cloud routing, VPN and dedicated connections, network segmentation in cloud, cloud-native networking (container networking, service meshes). Designing networks that span multiple cloud providers or hybrid on-premises/cloud.
Practice Interview
Study Questions
Redundancy, High Availability, and Disaster Recovery Design
Designing for no single points of failure, redundant paths, failover mechanisms, and disaster recovery. Understanding RPO/RTO requirements, backup connectivity, and testing DR plans.
Practice Interview
Study Questions
Data Center Networking and Spine-Leaf Architectures
Modern data center network design, spine-leaf topology, East-West vs. North-South traffic, overlay networks (VXLAN, NVGRE), network virtualization, and considerations for cloud and containerized environments. Understanding why traditional hierarchical designs don't scale for DC networks.
Practice Interview
Study Questions
Network Architecture Design Principles
Designing scalable, resilient network architectures. Understanding layered network design (access, distribution, core), redundancy patterns, fault isolation, and design for operational simplicity. Principles like defense-in-depth, loose coupling, and designing for observability.
Practice Interview
Study Questions
Onsite Technical Interview: Advanced Networking Protocols and Deep Dives
What to Expect
First onsite technical interview with a senior network engineer. This round goes deep into specific networking technologies and protocols relevant to large-scale infrastructure. Expect detailed questions about a specific protocol or technology you claim expertise in (e.g., BGP behavior in large networks, MPLS, multicast, QoS implementation, network virtualization). You may also be asked about recent industry changes or how you'd approach learning a new technology. This assesses technical depth and your ability to discuss complex topics with peers.
Tips & Advice
Come with a specific area of deep expertise that you can discuss in detail. If you claim BGP expertise, be ready for tough questions about convergence, failover, route filtering, and failure scenarios. Prepare to discuss not just how something works but why design decisions were made that way (historical context), current limitations, and where the industry is heading. If asked about a technology you're less familiar with, explain your learning approach: what resources you'd consult, how you'd test it, how you'd implement safely. This shows maturity. Mid-level engineers are expected to have depth in their specialty but also intellectual curiosity about unfamiliar areas.
Focus Topics
Multicast Networking
Multicast basics, IGMP, PIM protocols (Sparse Mode, Dense Mode), rendezvous points, multicast VPNs, and use cases for multicast. Operational challenges in multicast deployment.
Practice Interview
Study Questions
Quality of Service (QoS) and Traffic Management
QoS models (IntServ, DiffServ), marking and classification, queuing disciplines, congestion management, rate limiting, and implementing QoS end-to-end. Practical QoS configuration on routers and switches.
Practice Interview
Study Questions
MPLS, Traffic Engineering, and Advanced Routing
MPLS fundamentals, label switching, RSVP-TE, traffic engineering design, FEC (Forwarding Equivalence Class), and using MPLS for resilience and optimization. When and why MPLS is used in modern networks.
Practice Interview
Study Questions
Network Virtualization and Overlay Networks (VXLAN, NVGRE, etc.)
Overlay network concepts, VXLAN header structure and operation, control plane options (multicast, BGP EVPN), underlay network requirements, integration with physical network, scaling considerations, and operational debugging of overlay networks.
Practice Interview
Study Questions
Border Gateway Protocol (BGP) in Production Networks
Deep understanding of BGP operation, route selection process, path attributes, route filtering and manipulation, convergence behavior, failover mechanisms, BGP security (route filtering, RPKI, prefix hijacking prevention), and multi-AS design patterns. Real-world BGP failures and recovery.
Practice Interview
Study Questions
Onsite Technical Interview: Real-World Problem Solving and Incident Management
What to Expect
Second onsite technical interview, typically with an engineering manager or a senior engineer who has handled major incidents. This round presents a complex real-world scenario: a production network issue requiring diagnosis and resolution, possibly a major incident or a complex performance problem. You're given context and asked to walk through how you'd investigate, communicate, and resolve it. This assesses your troubleshooting depth, incident management skills, and ability to stay calm under pressure. A core responsibility for mid-level engineers is owning incident response and complex troubleshooting.
Tips & Advice
For the scenario, ask clarifying questions about symptoms, scope, and impact. Walk through your diagnostic approach step-by-step, explaining what you'd check and why. Use a logical, systematic methodology rather than guessing. Explain which diagnostic commands you'd run and how you'd interpret the output. If you get stuck, explain your next steps ('I'd escalate to vendor support' or 'I'd check the configuration change log'). Discuss communication during the incident—keeping stakeholders updated, maintaining documentation, coordinating with other teams. Mid-level engineers should show not just technical problem-solving but also the soft skills of incident management. Be honest about limitations and when you'd need help from specialists.
Focus Topics
Incident Management and Communication
Managing production incidents: establishing communication channels, keeping stakeholders informed of progress, coordinating across teams, documenting during the incident, and post-incident review process. Understanding severity levels and escalation paths.
Practice Interview
Study Questions
Configuration Management and Change Troubleshooting
Using configuration management tools, version control for network configs, identifying changes that caused issues, rollback procedures, and testing changes safely. Understanding how configuration changes propagate through networks.
Practice Interview
Study Questions
Performance Analysis and Bottleneck Identification
Identifying performance bottlenecks (CPU, memory, bandwidth, queue depth), using performance monitoring tools, baseline vs. current comparison, and understanding interactions between layers (network affecting application performance). Optimization strategies.
Practice Interview
Study Questions
Packet Analysis and Protocol-Level Debugging
Using packet capture tools (tcpdump, Wireshark), interpreting packet traces, understanding TCP/IP behavior at the packet level, diagnosing packet loss, reordering, or latency issues. Knowing which packets to capture and how to filter and analyze large trace files.
Practice Interview
Study Questions
Systematic Troubleshooting and Root Cause Analysis
Structured troubleshooting methodology: defining problem scope, gathering symptoms, forming hypotheses, testing systematically, and documenting root causes. Avoiding common pitfalls (jumping to conclusions, not isolating variables). Tools and approaches for different problem types.
Practice Interview
Study Questions
Onsite Behavioral and Culture Fit Interview
What to Expect
Final onsite round focused on behavioral assessment, cultural fit, and team dynamics. Usually conducted by an engineering manager or senior peer. Expect questions about how you work in teams, handle conflict, communicate with non-technical stakeholders, approach learning, and align with Apple's values. You'll discuss your career growth, how you've handled failure, examples of collaboration, and what you're looking for in a team. This round evaluates soft skills, maturity, and whether you'd thrive in Apple's engineering culture.
Tips & Advice
Use the STAR method (Situation, Task, Action, Result) for behavioral questions. Give specific examples, not generalizations. Show evidence of ownership ('I owned this project end-to-end') and collaboration ('I worked closely with security and operations teams'). Be honest about failures and what you learned. Discuss how you've mentored juniors (mid-level should show some mentorship). Ask thoughtful questions about team structure, current challenges, and growth opportunities. Prepare to discuss your leadership style and how you'd approach cross-team collaboration. Apple values engineers who care about the craft, continuously improve, and consider the impact of their work on users.
Focus Topics
Resilience, Handling Failure, and Growth Mindset
An example of a significant failure or mistake, what you learned, and how you grew from it. How you stay calm under pressure during outages. Your approach to continuous improvement and constructive feedback.
Practice Interview
Study Questions
Handling Ambiguity and Continuous Learning
How you approach new technologies or unfamiliar problems. Examples of learning new skills for a role, adapting to changes, and driving improvement in areas of uncertainty. Your learning style and resources.
Practice Interview
Study Questions
Mentoring, Knowledge Sharing, and Growing Junior Engineers
Evidence of mentoring or helping junior engineers, sharing knowledge, and raising the team's capabilities. Examples of documentation you've written, training you've delivered, or engineers you've helped develop. Your approach to raising technical standards.
Practice Interview
Study Questions
Ownership and End-to-End Project Delivery
Evidence of owning projects from design through implementation and production optimization. Describing how you took initiative, made decisions autonomously, and delivered impact. Examples of projects where you drove results with minimal guidance.
Practice Interview
Study Questions
Cross-Functional Collaboration and Communication
Examples of working effectively with security, operations, application teams, and other infrastructure groups. How you communicate technical decisions to non-technical stakeholders. Handling disagreements and building consensus. Bridge-building between teams.
Practice Interview
Study Questions
Frequently Asked Network Engineer Interview Questions
Tell me about a time you had to give someone you were mentoring difficult or critical feedback. How did you deliver it, and what happened afterward?
Sample Answer
Direct answer
Difficult feedback to a mentee works best delivered privately, tied to a specific, observed behavior and its concrete impact, not to the person's character, and followed up on to confirm the message landed and something changed. The delivery mechanics matter less than getting three things right: timing (soon after the behavior, not saved up), specificity (a real example, not a vague pattern), and follow-through (checking back in, not treating the conversation itself as the fix).
Structured elaboration
Before the conversation
- Get the facts straight: what exactly happened, what was the impact, and is this a one-off or a pattern. Vague feedback ("you need to be more careful") is unusable; a mentee can't act on a mood, only on a specific instance.
- Decide the stakes. Not all critical feedback carries the same weight:
- A routine performance gap (missed a deadline, sloppy code style) can wait for the next scheduled 1:1.
- An ethical or safety concern (something in the person's work poses real risk if it ships) changes the calculus: it needs to happen immediately, privately, and is often paired with a concrete containment step (pause the change, get a second reviewer), not just a conversation.
- Decide the medium: a private 1:1, not written feedback and not in front of the team, unless the finding also needs to be logged for safety or compliance reasons.
Delivering it
- Lead with the specific behavior and its impact, not a label: "this change would have introduced X" is usable; "this was careless" is not.
- Ask before you assert. The mentee may know something you don't (a constraint you weren't aware of); asking "walk me through the reasoning" often surfaces that before you've over-committed to a judgment.
- Separate the person from the work. The message is "this output has a problem," not "you are the problem."
When the power dynamic is reversed
Feedback isn't always flowing to someone junior. Giving critical feedback to a mentee who is more senior, more tenured, or simply more confident than you requires the same content but a different frame: lead with genuine respect for their experience, be more explicit that you're not questioning their general competence, and expect (and plan for) more pushback. Defensiveness here is a normal reaction to a status threat, not necessarily a sign the feedback was wrong; the skill is staying anchored to the specific evidence instead of either escalating or backing down.
After the conversation
- Confirm shared understanding before ending: ask them to restate what they heard.
- Agree on a concrete next step and a checkpoint to revisit it, not just "let's see how it goes."
- Follow up. Feedback that isn't revisited quietly signals it wasn't actually important.
Worked example
Situation
During a code review, I found a race condition in an interrupt service routine (an ISR, the block of code that runs automatically when a hardware event interrupts normal execution) a mentee had written: a shared buffer was being written from the ISR without disabling interrupts around the critical section (the stretch of code that touches shared data and must not be interrupted mid-update), so a preemption (the interrupt firing and pausing the main code at an unpredictable moment) at the wrong moment could corrupt data intermittently and unpredictably.
Why this wasn't routine feedback
This wasn't a style nitpick. Left unaddressed it was a latent, hard-to-reproduce bug that could surface in the field. That pushed it from "note it for next time" to "we talk today, and the change doesn't merge until it's fixed."
The conversation
I asked the mentee to walk me through what happens if the interrupt fires mid-write, rather than telling them the bug outright. They found the failure mode themselves partway through the explanation, so the fix landed as their own understanding rather than my correction. We then talked through the general pattern (anything touching state shared between an ISR and main-line code needs an explicit critical section) so it would generalize past this one bug.
Follow-through
I asked them to check two other places in the codebase where similar shared state existed, as a way to prove the concept had stuck rather than just fixing the one instance. Both had the same latent issue.
Result
The immediate bug was fixed before merge, and the mentee started flagging similar patterns unprompted in their own future changes, which was the real signal the feedback had generalized rather than just been complied with once.
Trade-offs & pitfalls
- The feedback sandwich dilutes the message. Padding critical feedback between two compliments is a common junior instinct; it often causes the actual point to get lost. Genuine positive feedback is worth giving, but on its own merits, not as camouflage for the critical part.
- Waiting to "collect examples" delays too long. A senior mentor gives feedback close to the event; batching several issues into one big conversation later makes it feel like an ambush and makes each point harder to act on.
- Not distinguishing skill gap from something more serious. A performance gap and an ethical or safety issue call for different urgency and different documentation; treating a safety issue as routine coaching is itself a failure mode worth naming.
- Confusing "they got defensive" with "I was wrong." Especially with a more senior or tenured mentee, defensiveness is a predictable reaction to a status threat. A junior mentor backs off; a senior one stays anchored to the specific evidence while still leaving room for the other person to be right about something they missed.
You inherit (or newly join and discover) a system you're now responsible for that is in poor shape: undocumented, fragile, lacking tests or monitoring, and causing frequent failures or disruption (for example, an unreliable CI pipeline, a flaky automation repository, a poorly documented service, a drifting cloud environment, or a stale backlog). Describe the concrete steps you would take in the first week to stabilize things and establish ownership, and outline a phased remediation plan over the following weeks or months, including milestones and how you'd measure progress.
Sample Answer
Direct answer
In the first week, stop the bleeding and claim the system as yours in writing; over the following weeks and months, work through a small number of phases, understand failure patterns, fix the worst recurring one, then build back tests, monitoring, and documentation, each with a milestone you can show and a number that proves things are actually improving.
Structured elaboration
- First week, stabilize and establish ownership: pull the history of recent incidents or failures to find the actual pattern, not the loudest anecdote, put even crude monitoring or alerting in place if none exists, and write a short doc stating what the system does, that you now own it, and what the plan is; send that to stakeholders so "who owns this" stops being a question.
- The following weeks, phased remediation with milestones: a first phase of roughly the next 3 weeks fixes the single highest-frequency failure mode, the quick win that buys you credibility and breathing room for the rest of the plan; a second phase of roughly the next 4 weeks adds test coverage and documentation for the core paths people actually depend on daily, not full coverage everywhere; a third phase covering the remaining weeks up to about 3 months hardens the rest, removes workarounds people have built to route around the system's flakiness, and puts a real runbook in place.
- Measuring progress: track a simple, visible leading indicator, most naturally failures or incidents per week, from a clear baseline, so stakeholders see a trend line rather than taking your word for "it's better now."
Worked example
I inherit a system with a baseline of about 6 disruptive failures a week and no documentation of why. Week 1: reviewing the last 2 months of incident history shows over half of them trace back to one specific recurring cause; I add a basic alert on that specific failure mode so at least we're notified instead of surprised, and send a short ownership note to stakeholders. By week 4, end of phase 1, the highest-frequency failure mode is fixed, and the weekly failure count drops from about 6 to about 3. By week 8, end of phase 2, core-path tests and basic documentation are in place, and the count drops further to about 1 a week. By week 12, end of the roughly 3-month plan, the remaining workarounds are removed and a runbook is in place, and the system is stable at close to 0 failures a week, down from the original baseline of 6.
Trade-offs and pitfalls
A common mistake is spending the first week building the perfect long-term architecture instead of just stopping the most disruptive recurring failure and telling people you own it; credibility comes from visible short-term progress, not a beautiful plan no one has seen work yet. Another failure is skipping the ownership communication step, leaving stakeholders unsure whether problems are still being routed to the previous owner or a whole team. Watch also for declaring success once the failure count drops without also removing the manual workarounds people built around the old flakiness; those workarounds often hide the next failure mode until you take them away.
Application servers and the primary database sit on the same network, and during peak traffic the link between them saturates, driving up query latency. Walk through how you'd confirm the network really is the binding constraint (and not something else), what you'd try first to buy headroom quickly, and what longer-term architectural change you'd make so this doesn't keep recurring as traffic grows.
Sample Answer
Confirming the network really is the binding constraint
I wouldn't take the symptom (link saturation, rising query latency) at face value without ruling out other causes first. I'd check network interface utilization on both the app server and database side during the exact peak window, and correlate its onset with the onset of the latency increase, while also checking CPU and disk I/O on both tiers over the same window. For example: NIC (network interface) throughput at 950 Mbps sustained out of a 1 Gbps link (95% utilized), with TCP retransmissions (packets being resent because they weren't acknowledged in time, a sign the network link is congested) rising right when p99 latency (the 99th percentile response time, the slowest 1% of requests) crosses 300ms, while app server CPU sits at 40% and database CPU at 55%, both well under their own ceilings. That combination, network metric pegged and rising retransmits at the same moment latency spikes, while compute stays comfortable, is what actually confirms the network link as the binding constraint rather than assuming it from the topology alone.
Buying headroom quickly
Without touching the architecture, I'd reduce the bytes crossing that link: stop over-fetching (select only the columns actually needed instead of every column), compress payloads, batch queries to cut per-query overhead, and move some read traffic to a local cache or replica so it never has to cross the saturated link at all. These changes can ship in days and buy real headroom while a longer-term fix is planned.
The longer-term architectural fix
The quick fixes reduce load on the shared link, but the underlying issue is that traffic growth keeps colliding with a fixed-bandwidth path between two tiers that are architecturally coupled. Longer term I'd look at moving the app and database tiers onto a higher-bandwidth or lower-latency path (a dedicated link, tighter physical or network placement to cut hop count), and reducing chattiness structurally, colocating read replicas closer to the app servers, or adding a caching layer so most reads never round-trip to the database at all.
What happens next
Once the network stops being the ceiling, growth will eventually push the next resource, likely database CPU or the app tier itself, into becoming the new binding constraint. This isn't a one-time fix, it's the same identify-the-constraint exercise that needs to run again at the next growth checkpoint.
Design a closed-loop system where streaming telemetry from network devices triggers automated remediation for problems such as interface flaps or BGP session loss. Describe the components and how you decide that an event is real and a fix is safe to run.
Sample Answer
Direct answer
Build it as a pipeline with a decision gate (a check that must pass before the next step) and a brake (a limit that stops the system acting too often) in front of every action: devices stream state over gNMI (the gRPC network management interface) into a bus (a message queue that decouples producers from consumers); a normalizer and correlator (a component that groups events sharing one cause) turn raw updates into one incident per cause; a policy engine decides whether the incident is real and whether a known-safe runbook (a written, tested fix procedure for one kind of problem) applies; an executor runs it under locks and rate limits; and a verifier confirms the symptom cleared, otherwise it rolls back and pages a human. Automatic action is allowed only for event classes with a tested runbook, never for "anything that looks bad".
Components
- Telemetry collection. A gNMI subscription in STREAM mode delivers long-lived updates. Only two values of each state matter for this design. Interface
oper-status(a leaf under/interfaces/interface/state) has several values, but the loop cares about DOWN and UP; the BGPsession-stateleaf walks through IDLE, CONNECT, ACTIVE, OPENSENT and OPENCONFIRM while a session is coming up, and the loop cares about ESTABLISHED (healthy) versus anything else. For state that changes rarely like these, use ON_CHANGE (the device sends an update only when the value changes, so a flap shows up as a burst of updates), optionally with aheartbeat_intervalso the value is re-sent periodically and a silent stream is noticed. For counters such as errors, which change constantly, use SAMPLE with asample_interval(the device sends the current value every interval, for example every 10 seconds). Thesync_responsemessage marks that the initial state of every subscribed path has been sent, so the pipeline knows its picture is complete. In time order: the subscription starts, the device sends the current value of every path,sync_responsemarks the end of that first dump, then ON_CHANGE updates arrive as things happen, SAMPLE updates arrive every interval, and a heartbeat re-sends an unchanged value so silence can be told apart from stability. If a platform does not offer ON_CHANGE for a leaf, subscribe with SAMPLE and accept the slower detection. - Stream bus and normalizer. Events are stamped with device, path and time and mapped to a common schema (device, object, kind, value).
- Correlator. Groups events that share a cause. A BGP session loss within seconds of an interface going down on the same device is a child of that interface incident; it does not trigger its own remediation.
- Policy engine. Checks the gates below and picks a runbook from an allowlist keyed by event class.
- Executor. Runs the runbook (for example an Ansible job) with a per-device lock, a rate budget, and a pre-check of current state.
- Verifier and audit. Re-reads telemetry after the action, closes the incident or reverts and pages, and logs the event, the decision and the evidence.
Deciding that an event is real
- Persistence: the condition holds past a debounce window (a short wait during which brief blips are ignored), or it repeats (three DOWN updates in 60 seconds is a flap).
- Corroboration: a second independent signal (the neighbor side also reports the link down, or the counters agree). A telemetry stream that goes quiet is a collector problem, not a link problem; the heartbeat separates the two.
- Not caused by us or by a change: the device is not in a maintenance window and the event is not within the hold-down (a quiet period after our own action during which the resulting events are ignored) of our own previous action.
Deciding that a fix is safe
- The event class has a runbook that has been tested and an owner.
- Blast radius (how much of the network a mistake can hurt) is bounded: one device or link per incident, and a global budget (here 3 actions per hour) after which the system stops acting and pages.
- Preconditions pass (for an interface shut: an alternate path with enough headroom to take the traffic), and the runbook is idempotent (running it twice leaves the same result as running it once).
- A kill switch (a manual off control) stops all automatic action instantly, and a circuit breaker (an automatic off control, like an electrical one) stops it when two consecutive actions fail verification.
- Actions with no known-safe fix, such as a lone BGP session loss with no interface event, become a ticket with the evidence attached.
Worked example
A tiny simulator of the correlation and gating logic, with fixed events (times in seconds). Each event is (time, device, kind, object, new value), where kind if is an interface and bgp is a session. Three constants set the rules: FLAP_WINDOW is how many seconds of history count toward a flap (60), FLAP_MIN is how many DOWN events inside that window make it a flap (3), and CORRELATE_S is how close in time a BGP drop must be to an interface DOWN on the same device to be treated as its consequence (10 seconds):
from collections import defaultdict
events = [
(0, "leaf1", "if", "Ethernet1", "DOWN"),
(2, "leaf1", "bgp", "10.1.0.0", "IDLE"), # session rides that interface
(9, "leaf1", "if", "Ethernet1", "UP"),
(14, "leaf1", "if", "Ethernet1", "DOWN"),
(21, "leaf1", "if", "Ethernet1", "UP"),
(30, "leaf1", "if", "Ethernet1", "DOWN"),
(31, "leaf1", "bgp", "10.1.0.0", "IDLE"),
(60, "leaf2", "bgp", "10.2.0.0", "IDLE"), # lone BGP loss, no interface event
]
FLAP_WINDOW, FLAP_MIN, CORRELATE_S = 60, 3, 10
MAX_ACTIONS_PER_HOUR = 3
downs = defaultdict(list); incidents = []
for t, dev, kind, key, val in events:
if kind == "if" and val == "DOWN":
downs[(dev, key)].append(t)
recent = [x for x in downs[(dev, key)] if t - x <= FLAP_WINDOW]
if len(recent) >= FLAP_MIN and not any(i["root"] == (dev, key) for i in incidents):
incidents.append({"root": (dev, key), "t": t, "children": []})
if kind == "bgp" and val != "ESTABLISHED":
parent = next((i for i in incidents if i["root"][0] == dev and t - i["t"] <= CORRELATE_S), None)
near_if = any(d == dev and 0 <= t - x <= CORRELATE_S for (d, _), xs in downs.items() for x in xs)
if parent: parent["children"].append(key)
elif not near_if: incidents.append({"root": (dev, key), "t": t, "children": []})
actions = 0
for i in incidents:
kind = "interface flap" if i["root"][1].startswith("Ethernet") else "bgp session loss"
if actions >= MAX_ACTIONS_PER_HOUR: verdict = "BUDGET EXHAUSTED: page a human"
elif kind == "interface flap": verdict = "ELIGIBLE: shut interface, then verify"; actions += 1
else: verdict = "NOT ELIGIBLE: no known-safe fix, open ticket with context"
print(i["root"], kind, "children", i["children"], "->", verdict)
Output:
('leaf1', 'Ethernet1') interface flap children ['10.1.0.0'] -> ELIGIBLE: shut interface, then verify
('leaf2', '10.2.0.0') bgp session loss children [] -> NOT ELIGIBLE: no known-safe fix, open ticket with context
Reading the BGP branch line by line: parent looks for an already-open incident on the same device that began at most 10 seconds ago; near_if is true if any interface on this device went DOWN within the last 10 seconds (the nested loop walks every recorded DOWN time x for every interface); if a parent exists the drop is recorded as its child, and only if there is neither a parent nor a nearby interface DOWN does the drop open its own incident.
Eight events become two incidents. The third DOWN at t=30 completes three DOWN events inside 60 seconds (at 0, 14 and 30), which opens the interface incident; the BGP drop at t=31 is attached to it as a child rather than acting separately. The BGP drop at t=2 happened before the incident existed but within 10 seconds of an interface DOWN, so it was not opened as its own incident either; it is the same event the flap explains. The lone loss on leaf2 has no interface cause, so the only safe automatic outcome is a ticket.
Pitfalls
- Remediation that triggers its own alert: the shut interface generates DOWN events. Mark the target under remediation first and suppress its events for a hold-down period.
- Acting on a stale picture: if the stream lagged, re-read the current state just before the action.
- Acting on the symptom the telemetry shows while the cause is upstream, for example shutting leaf uplinks because a spine is failing. Gate on correlation across devices and cap the blast radius, so a fabric-wide event stops the automation and pages a human.
Several workstations on one floor suddenly show 169.254.x.x addresses and users cannot reach anything. What does that tell you, what connectivity still works, and how do you narrow down the cause?
Sample Answer
Direct answer
A 169.254.x.x address means the workstation asked for an address with DHCP (Dynamic Host Configuration Protocol), got no offer, and gave itself an IPv4 link-local address. Microsoft calls this Automatic Private IP Addressing (APIPA), and the standard behind it is RFC 3927. The address is only valid on the local network segment: the host has no default gateway, no DNS server, and routers will not forward it, so users cannot reach anything beyond their own VLAN. Because several machines on one floor fail together, suspect something they share on the path between the floor and the DHCP server (the VLAN, the uplink, the relay, or the server's scope) rather than the machines. Step 1 below uses the pattern of who fails to tell which of those it is.
What the address tells you
- The host picks an address at random from 169.254.1.0 to 169.254.254.255 (the first and last 256 addresses of 169.254.0.0/16 are reserved, leaving 65,536 minus 512 = 65,024), then sends three ARP probes (Address Resolution Protocol, "does anyone own this address?") before using it. The mask is /16.
- Link still works at layer 2 (the local-network level of switches and MAC addresses, below IP), so the NIC, cable and switch port are probably fine. The failure is in getting a DHCP reply.
- A host keeps retrying. How often is implementation-dependent (RFC 3927 cites Mac OS checking for a DHCP server every five minutes as an example), so hosts recover on their own, usually within minutes, once DHCP works again.
What still works and what does not
- Works: other hosts in the same VLAN that also hold 169.254 addresses can reach each other by IP (ping, file shares by IP). Hosts and servers with static addresses or unexpired leases keep working, which is why the outage looks patchy.
- Fails: anything off the segment, because there is no gateway and routers must not forward link-local traffic. DNS names fail (no DNS server). Local name resolution by multicast may still resolve names between neighbours.
- Why "suddenly": existing leases survive. With an 8-day lease, a client first tries to renew at T1 (50 percent of the lease, day 4) and rebinds at T2 (87.5 percent, day 7) per RFC 2131. A DHCP outage that began days ago becomes visible only as machines reach those points or are newly plugged in, so ask when the problem actually began.
Narrowing it down, in order
- Scope the blast radius. Is it every machine on that floor or a subset, and one VLAN or several? Do the same model of PC on another floor work? If only new or rebooting machines fail while machines holding valid leases keep working, the problem is the DHCP service, scope or relay (existing leases hide it). If machines that already had a lease also lose connectivity together, the floor's own network (the VLAN, the uplink) has failed as well as DHCP.
- Separate "no DHCP" from "no network". On an affected PC set a temporary static address in the floor's subnet with its gateway and ping the gateway. If that works, layer 2 and the gateway are healthy and only the DHCP path is broken. If it fails, check the access port's VLAN, the trunk's allowed-VLAN list on the uplink (the list of VLANs a switch-to-switch link is permitted to carry, so a VLAN missing from it is cut off from the rest of the network), and port authentication (a network access control system checks a device before giving its port normal access, and can hold it in a restricted VLAN).
- Watch the DHCP exchange from the client. A normal lease is a four-step conversation: the client broadcasts a DISCOVER ("any DHCP server here?"), a server answers with an OFFER (an address on loan), the client broadcasts a REQUEST accepting it, and the server confirms with an ACK. Capture with Wireshark (display filter
dhcp) ortcpdump -ni eth0 'udp port 67 or udp port 68'on Linux. Three outcomes:- No DISCOVER leaves the PC: a client-side or port-authentication problem.
- DISCOVERs leave and no OFFER returns: the problem is between the floor and the server. Go to step 4.
- An OFFER arrives but the host still uses 169.254: rare, a host firewall or NIC problem.
A healthy capture shows DISCOVER, OFFER, REQUEST and ACK in that order within about a second; a broken one shows DISCOVER repeated every few seconds with nothing coming back.
- Check the floor's gateway and switches. The router or layer-3 switch interface for that VLAN must relay DHCP broadcasts to the server (a "helper address" setting, which forwards the client's broadcast as a direct message to the server because broadcasts do not cross routers); check it points at the current server's IP and that the relay was not lost when the VLAN was changed or merged. If DHCP snooping (a switch feature that drops DHCP replies from untrusted ports) is on, check that the uplink toward the server is marked trusted and look at the snooping drop counters and binding table for the VLAN.
- Check the server. Is the DHCP service running, is the scope active and not exhausted, does the server's log show DISCOVERs arriving from this relay, and is a failover partner healthy? On Windows Server these two commands show scope use and which servers are authorised in Active Directory (only an authorised server leases addresses):
Get-DhcpServerv4ScopeStatistics -ComputerName dhcp01.example.com -ScopeId 10.30.0.0
Get-DhcpServerInDC
Get-DhcpServerv4ScopeStatistics reports in-use and free counts and PercentageInUse for the scope; an exhausted scope has no free addresses, so a healthy result for a floor scope is a free count well above zero and a low percentage, while a scope stuck near 100 percent explains new clients getting no offer. Get-DhcpServerInDC should list this server; if it is missing, an Active Directory-joined server will not lease addresses. Worked example: a floor VLAN 10.30.0.0/24 whose scope covers .10 to .250 has 250 - 10 + 1 = 241 addresses. A floor with 150 desk PCs and 120 IP phones needs 270, so 29 devices go without an address once leases are full. The fix is a larger subnet or scope, shorter leases for transient devices, or splitting phones into their own VLAN.
6. Fix and verify. Restore the broken element (relay target, trunk VLAN, snooping trust, service or scope). Force a client with ipconfig /release then ipconfig /renew (Windows) and confirm with ipconfig /all that it shows a normal DHCP lease from the right server and a default gateway, then ping the gateway and a name. Check that the server's lease table now lists the floor's machines, and add a monitor on scope percentage in use so exhaustion alerts before users do.
Pitfalls
A static-address test that works proves the wire is fine but not that DHCP is fine, so always do both. Do not renew leases on the whole floor until the cause is fixed, because hosts with valid leases are the ones still working. A rogue DHCP server gives a wrong address rather than a 169.254 one, so it is a different symptom with a different fix. Linux hosts often show no address at all rather than a 169.254 address unless link-local fallback is configured, so the same fault looks different on mixed fleets.
Your network team has lost credibility with product teams after repeated capacity incidents. What communication rhythm and artifacts would you put in place to rebuild it, and how would you know it is working?
Sample Answer
Direct answer
Trust returns through predictability: product teams need to hear from the network team before problems, see honest numbers, and watch commitments get kept. I would set up a regular rhythm (a weekly review, a monthly capacity forecast, a change calendar and a quarterly planning session) plus a small set of artifacts: a capacity forecast, an incident follow-up commitments list and a change notice. I would judge success by whether product teams stop being surprised and whether our promises land on time.
Communication rhythm
- Weekly, 30 minutes: open network risks and upcoming changes, shared with product engineering leads.
- Monthly: capacity forecast review with product owners, including their growth plans, so demand is known before it becomes an incident.
- Change calendar, always current: a shared calendar of planned network changes that product teams can check at any time and that the weekly review walks through.
- Per incident: update cadence during, then a written follow-up within a fixed number of days listing what we will do and by when.
- Quarterly: planning session where product asks and network answers with what is possible.
Artifacts
- Capacity forecast. For each critical link (a network connection between two sites or systems, with a fixed capacity): current peak usage, growth rate, the threshold at which we act, and how long it takes to add capacity. Short, plain, dated.
- Commitments ledger. Every promise after an incident: owner, due date, status, visible to product teams. Missed dates are flagged by us first.
- Change notice. Plain-language description of a planned change, who could be affected, when, and how to reach us.
Worked example (illustrative)
A link carries 10 Gbps and peaks at 7.2 Gbps, so 72% utilised (utilised means the share of capacity in use: 7.2 / 10). Traffic grows about 5% per month. We act at 80% (8.0 Gbps). After 2 months the peak is 7.2 x 1.05 x 1.05, about 7.94 Gbps, and it crosses 8.0 Gbps a bit past 2 months (about 2.2). If adding capacity takes 3 months to order and install, we are already late, so the forecast tells product teams now: "we need approval this week, or expect risk within about two months". That is the honest early warning that rebuilds credibility, because it turns a surprise incident into a scheduled decision.
What the other two artifacts look like (illustrative)
Commitments ledger rows:
| Commitment | Owner | Due | Status |
|---|---|---|---|
| Add a second uplink to the payments region | N. Okafor | 30 Nov | On track |
| Alert product teams at 70% link use | J. Lim | 15 Oct | Done 12 Oct |
| Publish a failover test result | R. Diaz | 20 Oct | At risk: flagged by us on 8 Oct, new date 3 Nov |
Change notice:
Planned change: replace a router in the east data centre.
When: Saturday 02:00-04:00 UTC. Who may notice: brief loss of about
30 seconds for services in the east region, as traffic reroutes.
What you need to do: nothing, unless you run a batch job at that time.
Contact: network-oncall@example.com
How to know it is working
- Share of commitments delivered on or before the promised date (target: rising quarter over quarter).
- Incidents where product teams learned of the problem from us first, rather than from customers or dashboards.
- Forecast accuracy: how close the predicted peak was to the measured peak.
- A short quarterly survey to product leads ("do you trust our capacity information?"), read together with the numbers above.
Trade-offs and pitfalls
- Do not promise a date you may miss just to look responsive; a smaller kept promise beats a larger broken one.
- More meetings are not the goal. If the weekly review has nothing to say, shorten it.
- Trust returns slowly. One good quarter of kept commitments matters more than a strong kickoff presentation.
You must keep months of flow records at petabyte scale for capacity planning and security queries. Which storage architecture would you choose, how would you lay out the data so the common queries stay fast, and how do costs change as volume grows?
Sample Answer
Direct answer
Make Parquet files (a compressed columnar file format, explained below) on object storage (an S3-compatible service) the system of record, meaning the authoritative copy everything else can be rebuilt from, partitioned by hour (one folder per hour). Add a ClickHouse hot tier (an analytical database holding the newest, most queried data on fast storage) for the newest two weeks of raw records, and rollup tables (pre-summed copies of the data, much smaller than the raw records) for dashboards. Interactive search over months is served by the rollups and by partition-pruned scans (reading only the hour folders a query can match and skipping the rest), not by indexing every record. Choose a search cluster (Elasticsearch or OpenSearch) instead only if the security team genuinely needs second-level search over months of raw records and will pay for it. Cost grows almost linearly with retained bytes (volume discounts are small), so the controls are retention, compression and rollups.
Sizing and layout
- Ingest. 1 PB per 30 days is 1e15 / (30 x 86,400) = 386 MB/s on average.
- Compression assumption. Columnar compression on flow records is assumed to be 10x; measure it on your data, since every figure below scales with it.
- Hot tier. 14 days of raw: 1,000,000 GB / 30 x 14 = 466,667 GB raw, about 46.7 TB at 10x.
- Hour partitions. One hour of raw is 1,000,000 GB / 720 = 1,389 GB, or about 139 GB at 10x. A query that touches one hour and 3 of 20 equal-width columns reads about 20.8 GB; the same three columns over 6 months read 90 TB, which is 4,320x more. Partitioning by time and reading only needed columns is what keeps the common queries fast.
Parquet is a columnar file format (values are stored column by column, so a query reads only the columns it names): a file is split into row groups (blocks of many rows), each holding column chunks, and the footer stores per-column min and max statistics. Readers skip whole row groups whose min/max cannot match the filter and read only the columns requested. Sort rows within each file by exporter then time, so exporter and interface filters prune well, and write files of a few hundred MB.
ClickHouse hot tier (an open-source column-oriented analytical database; its MergeTree tables use a sparse index over the sort key, one entry per 8,192-row granule by default, and support partitions and TTL):
CREATE TABLE flows_raw
(
ts DateTime, exporter IPv4, in_if UInt32, out_if UInt32,
src_ip IPv6, dst_ip IPv6, src_port UInt16, dst_port UInt16,
proto UInt8, bytes UInt64, packets UInt64, sampling_rate UInt32
)
ENGINE = MergeTree
PARTITION BY toDate(ts)
ORDER BY (exporter, in_if, ts)
TTL ts + INTERVAL 14 DAY DELETE;
CREATE TABLE flows_5m
(
bucket DateTime, exporter IPv4, in_if UInt32, out_if UInt32, proto UInt8,
bytes_scaled UInt64, packets_scaled UInt64
)
ENGINE = SummingMergeTree
PARTITION BY toYYYYMM(bucket)
ORDER BY (exporter, in_if, out_if, proto, bucket);
CREATE MATERIALIZED VIEW flows_5m_mv TO flows_5m AS
SELECT toStartOfFiveMinutes(ts) AS bucket, exporter, in_if, out_if, proto,
sum(bytes * sampling_rate) AS bytes_scaled,
sum(packets * sampling_rate) AS packets_scaled
FROM flows_raw
GROUP BY bucket, exporter, in_if, out_if, proto;
What the statements do:
flows_rawholds one row per sampled flow record.sampling_rateis the rate the router used (1000 means one packet in every 1,000 was sampled), stored with each row so bytes can later be multiplied by it to estimate real traffic.ENGINE = MergeTreeis ClickHouse's main table type for large append-only data.PARTITION BY toDate(ts)stores each day in its own partition so old days can be dropped or skipped whole.ORDER BY (exporter, in_if, ts)sorts rows on disk by that key and builds the sparse index on it, so a filter on exporter and interface reads few granules.TTL ts + INTERVAL 14 DAY DELETEdeletes rows 14 days after their timestamp, which implements the two-week hot tier.flows_5mholds the 5-minute rollup.SummingMergeTreeadds up the numeric columns of rows that share the same sort key when ClickHouse merges data in the background (the merge can be incomplete at any moment, hence the advice to query withsum()andGROUP BY).PARTITION BY toYYYYMM(bucket)makes one partition per month.- The materialized view
flows_5m_mvis an insert trigger: each batch inserted intoflows_rawis grouped into 5-minute buckets, multiplied bysampling_rate, and written toflows_5m. Nothing is recomputed at query time.
Loading 10,000 rows of 1,500 bytes with a sampling rate of 1,000 and then querying with SELECT sum(bytes_scaled) FROM flows_5m returned 15,000,000,000, equal to 10,000 x 1,500 x 1,000, so the scaling and the rollup agree; always read the rollup with sum() and GROUP BY. Capacity dashboards read flows_5m, which is tiny next to the raw data.
Security queries such as "every flow touching this IP over 90 days" filter on an address, which the exporter-first sort and the file statistics do not prune. Keep a small side table of (hour, IP) pairs seen. For illustration it holds rows such as (2026-03-01 14:00, 203.0.113.7) and (2026-03-01 15:00, 203.0.113.7), one row per distinct address per hour, so it is a tiny fraction of the raw data. To answer "every flow touching 203.0.113.7 over 90 days", a query first looks that address up in the side table, which returns the hours it appeared in (for a rarely seen address, a handful), then scans only those hour partitions of the raw Parquet files instead of 2,160 hours (90 days x 24). The saving depends on the address: a rarely seen external address appears in a handful of hours, while a busy internal server appears in nearly all 2,160, so for those addresses the side table saves little and the query needs a time or exporter filter as well.
If you choose a search cluster instead
Elasticsearch stores IPv4 and IPv6 addresses in the ip field type and answers CIDR (a block of addresses written as prefix and length, for example 192.168.0.0/16 meaning every address starting 192.168) term queries. Use data streams with index lifecycle management (ILM), whose phases are hot, warm, cold, frozen and delete, and whose rollover action creates a new write index when the current one reaches a size, document-count or age limit. Map src_ip and dst_ip as ip, ts as date, and bytes and packets as numeric types, and keep retention short on the expensive tiers. At petabyte scale every indexed field costs storage and memory and ingest throughput, so index only the fields you filter on.
Cost as volume grows
| Stored (decimal TB) | S3 Standard list price per month | Per TB |
|---|---|---|
| 100 TB | $2,100 | $21.0 |
| 600 TB | $12,298 | $20.5 |
| 1,000 TB | $20,121 | $20.1 |
| 2,000 TB | $39,679 | $19.8 |
| 6,000 TB | $117,910 | $19.7 |
The figures use $0.023, $0.022 and $0.021 per GB-month for the first 50 TB, the next 450 TB and everything over 500 TB, which are the S3 Standard rates for US East (N. Virginia) on AWS's S3 pricing page (checked October 2026); other regions and storage classes are priced differently, and the stored volumes in the table are illustrative. AWS states that S3 usage is metered in binary gigabytes (1 GB is 2^30 bytes) and that 1 TB is 1,024 GB, so the tier breaks sit at 51,200 GB and 512,000 GB and the table converts each decimal TB of data (10^12 bytes) to GB before applying the tiers: 100 TB is about 93,132 GB. The per-TB price falls only about 6% between 100 TB and 6 PB, so cost is nearly proportional to retained bytes. Six months at 1 PB per month raw is 600 TB after 10x compression (about $12,300 per month). Request, query and egress charges come on top.
Interactive search versus object-storage batch
| Question | Search cluster | Parquet on object storage with a column engine |
|---|---|---|
| Latency | seconds, over what is indexed | seconds for pruned scans, minutes for wide ones |
| Cost growth | tracks indexed data and replicas on fast disks | tracks bytes stored; compute only when queried |
| Operational load | shard sizing, mappings, rebalancing | file compaction and partition hygiene |
| Fit | recent days, free-form pivots | months of history, capacity planning, batch forensics |
Pitfalls
- Partitioning by something with huge cardinality (such as the IP address) creates millions of tiny files.
- Keeping raw records for the full retention instead of rolling up.
- Forgetting to store the sampling rate with each record, so history cannot be rescaled.
In a service provider network, how would you protect traffic against link failure using segment routing, and how do you make sure the backup paths do not themselves become congested?
Sample Answer
Direct answer
Protect the links with TI-LFA (Topology-Independent Loop-Free Alternate, standardised in RFC 9855): every router precomputes, for each protected link and each destination, a repair path that follows the post-convergence path (the route the network will use once routing has finished recalculating after the failure) after the IGP (interior gateway protocol, the link-state routing protocol such as IS-IS or OSPF) reconverges, and encodes it as a short list of segments (each identified by a SID, a segment identifier: a number that names a router or a single link) installed in forwarding before anything fails. Then make sure the backup does not congest by treating it as ordinary post-failure traffic: simulate every single-link failure against the demand matrix and engineer capacity, metrics or steering until each failure case stays under a utilisation ceiling. The repair mechanism does not manage capacity at all, so the capacity check is a separate, mandatory step.
How the protection works, in the order it happens
- Detection. The point of local repair (PLR), the router adjacent to the failed link, learns of the failure from loss of signal or from BFD (bidirectional forwarding detection, a lightweight hello protocol designed for fast failure detection).
- Local switchover. The PLR flips the affected prefixes to the precomputed backup entry in its forwarding table. No signalling and no computation happens at failure time.
- The repair list. The backup path is described with segments. In SR (segment routing, where the router that starts a packet's journey, the headend, lists the waypoints in the packet), a Node-SID identifies a router and an Adjacency-SID (Adj-SID) identifies one specific link. TI-LFA uses the P-space (routers the PLR can reach without using the failed link) and the Q-space (routers that can still reach the destination after the failure). If the two spaces share a router, the Node-SID of that router is enough to carry the packet; if they are one link apart, a Node-SID plus an Adj-SID bridges them; wider gaps need more segments. RFC 9855 reports that in studied real networks link protection needed one SID or fewer for over 99 percent of cases, and that two or fewer sufficed for 99 percent of node-protection cases.
- IGP convergence. Meanwhile every router recomputes its shortest paths. Because the repair followed the post-convergence path, traffic does not change route a second time when convergence finishes. This is the reason TI-LFA is preferred over repairs that take a temporary detour.
- Coverage and limits. RFC 9855 guarantees coverage in any two-connected network, and covers link, node and shared risk link group (SRLG, links that share a fibre conduit and fail together) failures. The repair list must fit the headend's maximum SID depth (MSD), the number of segments the hardware can push onto a packet.
Why the backup path can congest, and how to prove it will not
The repair path equals the post-convergence path, so the load it carries during the repair is the load every router will carry once the IGP converges, and that state lasts until the link is fixed, which may be hours. RFC 9855 itself notes that post-convergence paths may introduce congestion if the capacity plan does not account for them. So the question is plain N-1 capacity planning: for each link failure, reroute the demand matrix and read the worst link.
Worked example: two edge routers pe-1 and pe-2, four core routers c1 to c4, every link metric 10, demand of 110 Gbps in each direction between the edge routers, links of 100 Gbps. The topology is drawn below. The design ceiling after any single failure is 80 percent, a planning rule chosen to leave room for hash imbalance and bursts (an assumption, set by your loss and delay targets). The model splits traffic equally over equal-cost paths, which is the average case, not a per-flow guarantee.
+---- c1 ------- c3 ----+
| | | |
pe-1 --+ | | +-- pe-2
| | | |
+---- c2 ------- c4 ----+
Each line is a 100 Gbps link with metric 10; c1-c2 and c3-c4 are the vertical links.
import heapq
from collections import defaultdict
LINKS = { # (a, b): (IGP metric, capacity in Gbps per direction)
("pe-1", "c1"): (10, 100), ("pe-1", "c2"): (10, 100),
("pe-2", "c3"): (10, 100), ("pe-2", "c4"): (10, 100),
("c1", "c3"): (10, 100), ("c2", "c4"): (10, 100),
("c1", "c2"): (10, 100), ("c3", "c4"): (10, 100),
}
DEMANDS = {("pe-1", "pe-2"): 110, ("pe-2", "pe-1"): 110} # Gbps
def loads(links, demands):
adj = defaultdict(list)
for (a, b), (m, _) in links.items():
adj[a].append((b, m)); adj[b].append((a, m))
load = defaultdict(float) # directed (u, v) -> Gbps
for (src, dst), gbps in demands.items():
dist, pq = {dst: 0}, [(0, dst)]
while pq:
d, u = heapq.heappop(pq)
if d > dist[u]: continue
for v, m in adj[u]:
if d + m < dist.get(v, 1e9):
dist[v] = d + m; heapq.heappush(pq, (d + m, v))
inflow = defaultdict(float, {src: gbps})
for u in sorted(dist, key=lambda n: -dist[n]):
nh = [v for v, m in adj[u] if dist[u] == m + dist[v]]
for v in nh:
load[(u, v)] += inflow[u] / len(nh)
inflow[v] += inflow[u] / len(nh)
return load
def report(title, links, limit=0.8):
print(title)
cases = [("no failure", links)] + [
(f"fail {a}-{b}", {k: v for k, v in links.items() if k != (a, b)}) for (a, b) in links]
for name, topo in cases:
l = loads(topo, DEMANDS)
(u, v), g = max(l.items(), key=lambda kv: kv[1] / (topo.get(kv[0]) or topo[kv[0][::-1]])[1])
cap = (topo.get((u, v)) or topo[(v, u)])[1]
flag = "OVER" if g / cap > limit else "ok"
print(f" {name:<14} worst {u}->{v} {g:.0f}/{cap}G = {g/cap:.0%} {flag}")
report("Before: 100G everywhere", LINKS)
CROSS = {("c1", "c2"), ("c3", "c4")}
UP = {k: (m, c if k in CROSS else 200) for k, (m, c) in LINKS.items()}
report("After: 200G except the c1-c2 and c3-c4 cross links", UP)
Before: 100G everywhere
no failure worst pe-1->c1 55/100G = 55% ok
fail pe-1-c1 worst pe-1->c2 110/100G = 110% OVER
fail pe-1-c2 worst pe-1->c1 110/100G = 110% OVER
fail pe-2-c3 worst pe-1->c2 110/100G = 110% OVER
fail pe-2-c4 worst pe-1->c1 110/100G = 110% OVER
fail c1-c3 worst pe-1->c2 110/100G = 110% OVER
fail c2-c4 worst pe-1->c1 110/100G = 110% OVER
fail c1-c2 worst pe-1->c1 55/100G = 55% ok
fail c3-c4 worst pe-1->c1 55/100G = 55% ok
After: 200G except the c1-c2 and c3-c4 cross links
no failure worst pe-1->c1 55/200G = 28% ok
fail pe-1-c1 worst pe-1->c2 110/200G = 55% ok
fail pe-1-c2 worst pe-1->c1 110/200G = 55% ok
fail pe-2-c3 worst pe-1->c2 110/200G = 55% ok
fail pe-2-c4 worst pe-1->c1 110/200G = 55% ok
fail c1-c3 worst pe-1->c2 110/200G = 55% ok
fail c2-c4 worst pe-1->c1 110/200G = 55% ok
fail c1-c2 worst pe-1->c1 55/200G = 28% ok
fail c3-c4 worst pe-1->c1 55/200G = 28% ok
At 100 Gbps every failure of an edge uplink or of a c1-c3 / c2-c4 link puts 110 Gbps on a 100 Gbps link (110 percent), even though the healthy network runs at 55 percent. Upgrading the uplinks and those two core links to 200 Gbps gives a worst case of 110/200 = 55 percent after any failure, with no link at its limit. The cross links c1-c2 and c3-c4 stay at 100 Gbps: the search for the worst link covers them too, and in every case that worst link is at most 55 percent.
Why the numbers come out this way, on the drawing: with no failure, pe-1 reaches pe-2 over two equal 30-metric paths (via c1-c3 and via c2-c4), so 110 Gbps splits 55 and 55. If c1-c3 fails, the path through c1 now costs 40 (pe-1, c1, c2, c4, pe-2), more than the 30 via c2-c4, so after convergence all 110 Gbps uses pe-1 to c2: 110 percent of a 100 Gbps link.
The repair list on this topology (illustrative)
Take the failure of c1-c3, where c1 is the PLR and pe-2 the destination. The P-space is the set of routers c1 can reach along its normal shortest paths without crossing the failed link: it includes c2 (directly connected). The Q-space is the set of routers that can reach pe-2 without crossing the failed link: it includes c2 (via c4) and c4. c2 is in both, so one Node-SID, that of c2, is the whole repair list: c1 sends the packet to c2, and c2 forwards it normally over c4 to pe-2. That is exactly the post-convergence path from c1 (c1, c2, c4, pe-2), so nothing changes route a second time when the IGP converges. If P-space and Q-space had shared no router but a single link had bridged them, the list would be a Node-SID plus an Adj-SID, as in step 3. In the 100 Gbps design the repair is fast, but the load it creates is the 110 percent case above, which is why the capacity step is separate.
When more capacity is not available
A sensible order is cheapest first, and the bullets below follow it: metric changes (configuration only), then queueing (limits damage while you fix the cause), then steering with an SR policy or a Flexible Algorithm (needs SR features), and last central path computation (needs a controller). SRLG data is an input to every one of these steps rather than a step of its own.
- Metric change. Re-run the same simulation with altered IGP metrics so that the post-convergence path splits over more equal-cost paths. The simulation, not intuition, says whether it helped.
- Queueing. DiffServ queues (traffic marked into classes in the IP header, each class served from its own queue with its own priority) make congestion hit best-effort traffic first, which protects voice and trading classes while capacity is added. It limits the damage; it does not remove it.
- Steering the big flows. Move a named class of traffic with an SR policy (an ordered segment list bound to a colour, a number that labels the intent such as low latency, and to an endpoint; RFC 9256) or a Flexible Algorithm (a user-defined IGP computation with its own constraints, RFC 9350) onto a separate topology, so the repair of ordinary traffic does not land on top of it.
- Central path computation. A stateful PCE (path computation element, a server that computes paths, RFC 8231) can move large flows after the IGP settles using PCEP (the PCE communication protocol, the way the PCE sends paths to routers). That fixes the steady state, not the first moments of the repair.
- SRLG data. Without SRLG tags the planner thinks two links in one conduit are independent. Tag them or the simulation understates the failure.
Pitfalls
Checking only no-failure utilisation, simulating only single failures when a maintenance window plus a failure is routine, and forgetting that repair-list depth is a hardware limit on some platforms.
You observe periodic route oscillations for a set of prefixes approximately every 5 minutes. Explain how you would analyze for causes such as BGP MED oscillation, iBGP route-reflector feedback loops, next-hop unreachability flapping, route flap damping interactions, or upstream policy. Specify the logs, BGP update traces, timers, packet captures, and configuration changes you would examine and tests you would run to identify and mitigate the oscillation.
Sample Answer
Direct answer
A steady 5-minute period means something is driven by a timer or a repeating state machine, not by random failures. I would first prove what is oscillating (reachability, best-path choice, or only attributes), group the affected prefixes by what they share (next hop, neighbor AS, origin AS, community), and then match the period against each candidate timer. The five candidates leave different fingerprints, so a few targeted observations separate them before I change anything.
Vocabulary used below. A route reflector (RR) is an iBGP router (BGP between routers of one autonomous system) that re-advertises routes it learned from its clients, so those clients need not peer with every router; an RR and its clients form a cluster. Two attributes stop reflection loops: ORIGINATOR_ID (the router ID of the router that first injected the route into the AS) and CLUSTER_LIST (the CLUSTER_IDs of the clusters the route has been reflected through, newest first); the CLUSTER_ID is the cluster's name, shared by its redundant RRs. A confederation splits one AS into member ASes with the same effect of avoiding a full mesh. MED (Multi-Exit Discriminator) is a number a neighboring AS attaches to say which of its links it prefers you use, lower wins, and it is only comparable between paths from the same neighbor AS (RFC 4271 section 9.1.2.2). MRAI (MinRouteAdvertisementInterval) is the minimum time a router must leave between two UPDATEs about the same prefix to one peer (RFC 4271 section 9.2.1.1). BMP (BGP Monitoring Protocol, RFC 7854) lets a router stream a copy of the routes and updates it receives to a collector.
Step 1: Pin down the pattern (10 minutes of data)
- Pick two or three affected prefixes. Capture every UPDATE and WITHDRAW for them from a passive source: a BMP feed or a route collector if you have one, otherwise a debug of BGP updates filtered to those prefixes on one router, plus syslog timestamps. Record, per event: time, peer, whether it was a withdrawal or an announcement, and the attributes (AS path, next hop, MED, local preference, communities, ORIGINATOR_ID, CLUSTER_LIST).
- Run a packet capture on one affected session so the capture agrees with the router's own view:
tcpdump -nn -ttt -w bgp.pcap 'tcp port 179 and host 192.0.2.1'(the filter compiles as written; replace the address with the peer). Decode the UPDATEs offline. - Classify each event:
- Withdrawal then re-announcement: the prefix really becomes unreachable (next hop loss, upstream withdrawal, session event).
- Announcement only, attributes changed: the new UPDATE silently replaces the old path (an implicit withdrawal, no WITHDRAW message is sent). This is best-path churn. MED, route reflection and policy cause it.
- Group the prefixes. If they all share one next hop, suspect next-hop reachability. If they all come from two neighbor ASes, suspect MED. If they share an origin AS or a community, suspect upstream policy.
- Write down the period precisely (for example 300 s plus or minus 1 s) and the phase relative to other events (the IGP (interior gateway protocol) log, interface errors, scheduled jobs).
- Review configuration changes: diff the running configuration against the last known-good version and read the change log for the onset time of the first oscillation. Look for new or edited route-maps, MED or local-preference settings, damping, route-reflector cluster IDs, next-hop-self, timers, and IGP metric changes. A flap that began right after a policy change is the cheapest hypothesis to test.
Reading the router side: on a reflected path show ip bgp <prefix> prints a line such as Originator: 192.0.2.1, Cluster list: 2.2.2.2 (illustrative addresses). The Originator is the router that first learned the route inside the AS, and the Cluster list shows which clusters reflected it. If the Originator or Cluster list changes between consecutive updates for the same prefix, the route is arriving through different reflection paths each time, which is the reflection signature.
Step 2: Match the period to a timer
RFC 4271 section 10 suggests these defaults, which are the timers worth comparing against 300 seconds: the hold time 90 seconds with keepalive at one third of it, MinRouteAdvertisementInterval 30 seconds on eBGP and 5 seconds on iBGP, and the connect retry timer 120 seconds. Also compare against the IGP SPF (shortest path first) throttle (the delay a link-state IGP waits before re-running its path calculation after a change), the BGP next-hop tracking delay (BGP's watch on the IGP route to each next hop, so it reacts within seconds when one disappears; Cisco default 5 seconds), route flap damping half-life (default 15 minutes), and anything scheduled on the box (IP SLA, tracked objects, cron-style automation). None of the protocol defaults equals 300 seconds. A whole-session reset resets every prefix from that peer, so a flap limited to "a set of prefixes" points away from keepalive and hold-timer expiry and toward attribute or next-hop causes.
As a first suspicion, not proof: MED oscillation has no clock of its own, its rhythm comes from how quickly updates propagate, so a period that stays at exactly 300 s is more likely a configured timer, a scheduled job or an upstream's own policy. Rule those in or out with the capture first, then move to the route-reflection causes.
Step 3: Test each cause, in this order
| Cause | Fingerprint | Test | Mitigation |
|---|---|---|---|
| Next-hop unreachability | Withdrawals for every prefix sharing a next hop, at the moment the IGP route to that next hop disappears | show ip route <next-hop> repeatedly; correlate with IGP adjacency, LSA or LSP logs and interface counters; check whether the next hop is only reachable over a flapping link | Fix the link or IGP instability; BFD on the underlying link; keep next hops on stable loopbacks. Tuning bgp nexthop trigger delay (range 1 to 100 s, default 5 s) changes how fast BGP reacts, it does not remove the cause |
| MED-induced oscillation | The prefix stays reachable but its best path rotates among three or more paths, with the rotation visible as updates and withdrawals of reflected copies between the RRs; no external change | show ip bgp <prefix> on the route reflectors and a client: are there paths from two or more neighbor ASes with different MEDs, behind a route-reflector or confederation design? | RFC 3345 mitigations: make links between clusters cost more in the IGP than links inside a cluster, so each RR prefers its own cluster's exits; stop accepting MED (overwrite it with a constant on ingress); set local preference per advertising AS so MED is not compared across ASes; or, where the router count allows it, use a full iBGP mesh (RFC 3345 lists it and notes it does not scale) |
| Route-reflector feedback | Same prefix reflected between RRs with changing CLUSTER_LIST; or forwarding loops, not only control-plane churn | Read ORIGINATOR_ID and CLUSTER_LIST in show ip bgp <prefix>; check that redundant RRs in one cluster share a CLUSTER_ID (RFC 4456 section 7) and that RR topology follows the physical topology | Correct cluster IDs; follow the RFC 4456 loop-prevention design; align RR hierarchy with IGP topology |
| Route flap damping | Prefix disappears for tens of minutes after repeated flaps; show ip bgp dampening dampened-paths lists it | show ip bgp dampening flap-statistics, then show ip bgp dampening dampened-paths | Use RFC 7196 values (suppress threshold of at least 6000), or exempt the affected prefixes; damping hides a flap, it is not its source |
| Upstream policy | Announcements arrive with periodically changing attributes (AS-path prepend, MED, community) | Compare consecutive UPDATEs from the capture; ask the upstream for their change window; check whether the change is on their schedule | Escalate with the capture as evidence; filter or normalize the attribute on ingress; if only the MED changes, overwrite it |
How a route reflector plus MED makes the best path rotate
MED is only compared between paths learned from the same neighbor AS, and a route reflector chooses from only the paths it has been sent. Together these let two RRs knock each other's preferred path out in turn. The numbers below are the textbook case from RFC 3345 (Figure 1). Cluster 1 has RR Ra with clients Rb and Rc; cluster 2 has RR Rd with client Re. All three clients learn the same prefix from outside: Rb from AS10 with MED 10, Rc from AS6 with MED 1, Re from AS6 with MED 0. Ra's IGP costs are 5 to Rb, 4 to Rc and 13 to Re; Rd's are 12 to Re, 5 to Rc and 6 to Rb.
- Ra holds the AS6 path via Rc (MED 1, cost 4) and the AS10 path via Rb (cost 5). They come from different neighbor ASes, so MED is not compared and the lower IGP cost wins: Ra picks the AS6 path and advertises it to Rd.
- Rd now holds Re's AS6 path (MED 0, cost 12) and Ra's AS6 path (MED 1, cost 5). Same neighbor AS, so MED decides: MED 0 beats MED 1. Rd picks Re's path and advertises it to Ra.
- Ra now sees an AS6 path with MED 0 (via Rd, cost 13). That path eliminates the MED 1 path from Rc, leaving the AS6 MED 0 path (cost 13) against the AS10 path (cost 5). IGP cost now picks AS10, so Ra advertises the AS10 path and withdraws what it sent about AS6.
- Rd now holds Re's AS6 path (cost 12) and Ra's AS10 path (cost 6). IGP cost picks AS10 and Rd withdraws its AS6 advertisement.
- Ra no longer hears the MED 0 path, so the MED 1 path via Rc is back in play and beats AS10 on IGP cost (4 against 5). That is step 1 again, and the cycle repeats with no external change.
RFC 3345 states the conditions: a single-level reflection or confederation design, and MEDs from two or more neighbor ASes that are unique. Its preferred mitigation is the one in the table: had the inter-cluster IGP metrics been much larger than the intra-cluster ones, step 3 would not occur, because Ra would always prefer exits inside its own cluster.
Worked example: how damping would behave on a 5-minute flap cycle
Route flap damping gives each prefix a penalty score. Every flap adds points, the score decays over time and halves every half-life, and once it passes the suppress threshold the router stops using the route; it uses the route again when the score decays below the reuse threshold. Take the default Cisco damping parameters listed in RFC 7196 Table 1: a withdrawal adds 1000, an attribute change adds 500, suppress threshold 2000, reuse threshold 750, half-life 15 minutes, maximum suppress time 60 minutes. The ceiling is therefore 750 x 2^(60/15) = 12000. With one withdrawal every 5 minutes, the penalty just after each event is 1000, 1794, 2424, 2924, 3321, 3635, 3885, 4084 (computed, each step decays the previous value by 0.5^(5/15) and adds 1000). It crosses 2000 at the third flap (minute 10) and the route is suppressed from then on. The steady-state value approaches 1000 / (1 - 0.5^(5/15)) = about 4847, so once the upstream stops flapping the route stays suppressed for 15 x log2(4847/750) = about 40 minutes before reuse. That is the operational lesson: damping converts a 5-minute flap into a 40-plus-minute outage, so if the symptom stays a clean 5-minute cycle, damping is not suppressing it, and if reachability becomes much worse after the first 10 minutes, damping is amplifying it.
Step 4: Mitigate and prove it
Apply the one fix that matches the fingerprint, then verify for at least two full periods plus margin (a 12-minute window shows two cycles): zero UPDATEs for the affected prefixes on the capture, a stable best path in show ip bgp <prefix>, and no growth in flap statistics. If you cannot prove which cause it is, a safe interim step is to stop the loop from propagating (a stable static or aggregate at the edge), not to widen timers across the whole network.
Pitfalls
- Changing MRAI or damping first can hide the symptom while the root cause remains.
- MED is compared only between routes learned from the same neighboring AS (RFC 4271 section 9.1.2.2), so "always compare" settings change the whole decision process and need a network-wide change window.
- Treating route reflection and MED as two unrelated suspects: RFC 3345 shows the oscillation needs both (a route-reflection or confederation design plus unique MEDs from two or more ASes).
Design the network portion of a disaster recovery (DR) solution where primary systems run on-premises and DR runs in a cloud region or different cloud provider. Requirements: RTO under 1 hour, RPO 15 minutes, secure connectivity options, DNS failover, and consistent network and security policies. Provide a cutover plan, steps to pre-provision connectivity, and validation/testing approaches.
Sample Answer
Direct answer
A disaster recovery (DR) network design with a recovery time objective (RTO) under one hour and a recovery point objective (RPO) of 15 minutes, where the primary runs on-prem and DR runs in a cloud region or a different cloud provider, needs a warm-standby posture, not cold and not fully active-active: the network path, DNS failover mechanism, and security policy must already exist and be continuously validated before an incident, because provisioning any of that from scratch during an actual failure consumes most or all of the one-hour budget on its own. The core design decision is pre-provisioned, always-on private connectivity to the DR site plus a DNS failover mechanism that does not depend on a human noticing the outage, combined with an application tier designed to be stateless (or to externalize its session state) so that failed-over compute in the DR site can serve requests correctly on first contact rather than needing to rebuild in-memory state it never had.
Structured elaboration
Pre-provisioned connectivity, not connectivity built during the incident
- Primary path: a dedicated private circuit (AWS Direct Connect, Azure ExpressRoute, or GCP Dedicated/Partner Interconnect, depending on the DR target) between the on-prem data center and the DR environment, provisioned and tested well before any incident, carrying ongoing replication traffic during normal operation so it is proven working, not a cold link nobody has used.
- Backup path: a site-to-site virtual private network (VPN) over the public internet as a fallback if the dedicated circuit itself is the thing that failed, since a DR plan whose only path to the DR site shares a failure domain with the primary's own connectivity is not actually a DR plan.
- Consistent addressing: the DR environment's network should use an IP addressing scheme that does not require renumbering the application or its configuration during failover, either matching the on-prem scheme via an extended layer-3 design where feasible, or having every configuration reference a name, not a hard-coded address, so cutover is a DNS or routing change, not a configuration-file edit under pressure.
DNS failover
Health-checked DNS failover (a managed DNS service, for example Route 53 or Cloud DNS, running active health checks against the primary and automatically shifting the record to the DR endpoint on failure) removes the dependency on a human noticing the outage and manually triggering a change, which matters directly against the one-hour RTO: every minute spent on human detection and decision-making is a minute not available for the technical recovery steps. A short time-to-live (TTL) on the relevant DNS record (short enough that client-side caching does not meaningfully delay propagation, commonly under a few minutes) is a prerequisite for this to actually deliver fast failover rather than a failover that is technically triggered quickly but takes a long time to reach clients.
Consistent network and security policies
The DR environment's firewall rules, security groups, and network access control lists must mirror the primary's intended policy exactly, generated from the same infrastructure-as-code source rather than maintained as a hand-copied second version, because a security policy that quietly drifted out of sync with the primary is either a security hole (DR is more permissive than intended) or a fresh outage cause of its own (DR is more restrictive than intended and blocks legitimate traffic) discovered at the worst possible moment, during an actual failover.
Cutover plan and pre-provisioning steps
- Continuous replication: data replicates from on-prem to the DR site on an ongoing basis (not started at incident time), with replication lag monitored and alerted well inside the 15-minute RPO budget, commonly alerting at a fraction of that budget so there is response time before the hard limit is breached.
- Warm, not cold, compute: DR compute capacity exists in a scaled-down but running state (or as immediately-launchable machine images and infrastructure-as-code that can stand up full capacity in minutes, not hours), since a fully cold DR environment (nothing provisioned until the incident) reliably consumes too much of a one-hour budget just on provisioning.
- Detection and decision: automated health checks detect the primary's failure; for this RTO, the decision to fail over should be automated for clearly-unambiguous failure signatures (the primary is entirely unreachable) with a human-in-the-loop override available for ambiguous cases, since a fully manual decision gate risks consuming a large, variable fraction of the one-hour budget on paging and judgment alone.
- Promotion and traffic cutover: promote the DR data store to primary, confirm replication lag was within budget at the moment of failure, then flip DNS (or the equivalent traffic-steering mechanism) to the DR endpoint.
- Validation: run an automated smoke-test suite against the now-live DR environment before declaring the incident resolved, confirming the application is not just reachable but functionally correct.
Validation and testing approach
A DR design meeting these targets needs the RTO's own budget broken into its component steps and each step measured on a recurring schedule, not assumed:
| Step | What it covers | Why it needs its own measured budget |
|---|---|---|
| Detection | Automated health check confirms primary is down | A slow or flapping health check delays everything downstream |
| Decision | Automated trigger, or human confirmation for ambiguous cases | The largest source of variance if left fully manual |
| Promotion | DR data store becomes primary, replication lag confirmed within RPO | Must be tested, not assumed instant |
| Traffic cutover | DNS or load-balancer change takes effect for real clients | TTL and client-side caching behavior determine the real-world delay, not just the technical trigger time |
| Validation | Automated smoke tests confirm functional correctness | Skipping this risks declaring recovery complete while the application is actually broken |
Each row should be timed in a scheduled DR drill (quarterly at minimum for a target this strict) with the sum compared against the one-hour budget, and any row consistently eating a disproportionate share of the budget is the one to re-engineer first.
Worked example: compute placement and session handling for a failover-safe application tier
A DR design that gets the network right can still fail the RTO in practice if the application tier itself cannot serve traffic correctly the instant it lands in the DR site. The specific risk is session state: if user sessions are held in local, in-memory storage on the primary's application servers, those sessions do not exist at all in the DR site after failover, every user is silently logged out or, worse, lands on a DR instance that treats them as unauthenticated mid-transaction. Two designs address this, with a real trade-off between them:
- Token-based sessions (a signed token, commonly a JSON Web Token, JWT, containing the session claims, verified statelessly by any server using a shared signing key rather than looked up from a session store): this is the stronger fit for a strict RTO, because any DR instance holding the same verification key can validate a session on first contact with zero session-state replication needed at all. The trade-off is that revoking a session before its token expires needs an explicit mechanism (a short token lifetime forcing frequent re-issuance, or a small, separately-replicated revocation list) since a bare stateless token, once issued, is valid until it expires no matter what.
- A replicated session store (a distributed cache, for example a Redis cluster, replicated from on-prem to the DR site): this keeps session data centrally managed and revocable, but now the session store itself is one more thing that must replicate within the 15-minute RPO and be included in the promotion step of the cutover plan, adding another moving part to the exact recovery path this design is trying to keep as tight as possible.
For a design specifically optimizing for a sub-one-hour RTO, token-based sessions with a short lifetime and a lightweight, separately-replicated revocation list is the recommended default: it removes session-store promotion from the critical failover path entirely, and the revocation list, being small and infrequently written, is far easier to keep within the RPO budget than a full session store carrying the constant read/write traffic of every active user's session.
Trade-offs & pitfalls
- Treating the DR network path as something that can be provisioned at incident time, rather than pre-provisioned and continuously used for real replication traffic, is the most common way a one-hour RTO target is missed by a wide margin.
- A fully manual failover decision gate is the single largest source of unpredictable delay against a strict RTO, automate detection and the trigger for unambiguous failures, and time-box the human step for ambiguous ones.
- Letting DR-site security policy drift out of sync with the primary, by maintaining it as a hand-copied second version instead of generating both from the same source, creates a hidden failure mode discovered only during an actual failover.
- Choosing a replicated session store without accounting for it as a full participant in the RPO and cutover plan (its own replication lag, its own promotion step) is a common way DR designs meet their database RPO target while quietly missing it for session state.
- Skipping the post-failover smoke-test validation step in the interest of declaring recovery complete quickly risks calling an incident resolved while the application is actually serving broken responses.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Network Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs