Amazon Network Engineer (Junior Level) - Comprehensive Interview Preparation Guide
Amazon's interview process for junior-level Network Engineer roles typically follows a structured approach combining technical assessment with behavioral evaluation. The process begins with a recruiter screening call, progresses through 1-2 technical phone screens focusing on networking fundamentals and troubleshooting skills, and culminates in 4-5 onsite interview rounds covering technical depth, network architecture design, real-world troubleshooting scenarios, and Amazon's Leadership Principles.
Interview Rounds
Recruiter Screening
What to Expect
Initial 30-45 minute conversation with Amazon recruiter to assess background, experience level, career goals, and fit for the Network Engineer role. Recruiter will review your resume, discuss your networking experience, explain the role and team structure, and assess cultural alignment with Amazon's Leadership Principles. This round determines if you advance to technical phone screens.
Tips & Advice
Have your resume accessible and be ready to discuss your hands-on networking experience. Prepare 2-3 concise examples of network challenges you've solved or infrastructure projects you've worked on. Research the team or organization you're interviewing for. Ask about team structure, current infrastructure challenges, and what success looks like in the first 90 days. Be clear about your experience level as a junior and your enthusiasm for learning. Mention any relevant certifications (CompTIA Network+, Cisco CCNA, etc.) if applicable.
Focus Topics
Relevant Certifications and Credentials
Discussion of networking certifications (CompTIA Network+, Cisco CCNA, Juniper JNCIA) or cloud certifications (AWS Certified Cloud Practitioner, AWS Certified Solutions Architect) that validate technical knowledge.
Practice Interview
Study Questions
Motivation for Network Engineering at Amazon
Clear reasoning for pursuing this specific role at Amazon, what appeals about the company, and what you hope to accomplish in the position.
Practice Interview
Study Questions
Amazon Leadership Principles Alignment
Ability to articulate how your work style and values align with Amazon principles such as Customer Obsession, Ownership, Invent and Simplify, and Earn Trust.
Practice Interview
Study Questions
Understanding of Network Engineer Role Scope
Knowledge of what Network Engineers do daily: infrastructure design, equipment configuration, troubleshooting, security implementation, monitoring, and capacity planning.
Practice Interview
Study Questions
Professional Background and Experience Summary
Clear articulation of your networking experience, previous roles, hands-on projects, and career progression to junior Network Engineer level.
Practice Interview
Study Questions
Technical Phone Screen 1: Network Fundamentals and Troubleshooting
What to Expect
90-minute technical phone screen focusing on networking fundamentals, common troubleshooting scenarios, and practical problem-solving. You'll be asked about OSI model layers, TCP/IP concepts, common network protocols, and guided through real-world troubleshooting scenarios like connectivity issues, DNS failures, or routing problems. Interviewer will assess your understanding of how networks operate, your troubleshooting methodology, and ability to use standard tools.
Tips & Advice
Start troubleshooting scenarios by asking clarifying questions (what exactly is broken, what have you already verified). Use a systematic approach: start with basics (IP connectivity, routing), then move to specific issues (DNS, firewall, application). Demonstrate familiarity with tools like ping, traceroute, netstat/ss, dig, nslookup, arp, and ip route. For each scenario, explain your thought process out loud. Be honest about gaps in knowledge but show eagerness to learn. Use proper terminology but avoid unnecessarily complex jargon. If stuck, ask clarifying questions or propose multiple hypotheses and how you'd test each one.
Focus Topics
DNS Resolution and Common DNS Issues
Understanding DNS hierarchy, recursive vs authoritative lookups, common DNS issues (server unreachable, configuration problems), and tools to diagnose DNS failures.
Practice Interview
Study Questions
IP Addressing, Subnetting, and Routing Basics
Practical IP addressing concepts including public vs private addresses, subnet masks, default gateways, routing tables, static vs dynamic routing, and how packets are routed between networks.
Practice Interview
Study Questions
Network Connectivity Troubleshooting Scenarios
Methodology for troubleshooting common issues: host unable to reach network, DNS failures, connectivity to specific subnet unreachable, port connectivity problems, and MTU-related packet fragmentation issues.
Practice Interview
Study Questions
OSI Model and Network Layers
Understanding of the seven layers of the OSI model (Physical, Data Link, Network, Transport, Session, Presentation, Application) and which protocols/issues occur at each layer.
Practice Interview
Study Questions
Network Diagnostic Tools and Commands
Practical use of tools: ping and traceroute for connectivity testing, netstat/ss for port and socket information, dig/nslookup for DNS queries, arp/ip neigh for ARP tables, ip route for routing tables, nc for port testing.
Practice Interview
Study Questions
TCP/IP Protocol Suite Fundamentals
Core understanding of IPv4 and IPv6 addressing, CIDR notation, TCP vs UDP, DNS resolution process, DHCP process, ARP, ICMP, and how these protocols interact.
Practice Interview
Study Questions
Technical Phone Screen 2: Infrastructure and Network Services
What to Expect
90-minute technical phone screen diving deeper into network infrastructure components, services, and advanced troubleshooting. Topics include VLAN configuration and inter-VLAN routing, firewall and NAT concepts, network security basics, DHCP server configuration and issues, load balancing fundamentals, and scenarios where multiple services interact. This round assesses ability to think about complete infrastructure rather than isolated issues.
Tips & Advice
For infrastructure questions, think end-to-end: how do services connect, what can go wrong at each layer, how do you verify configuration. For VLAN/firewall/NAT questions, draw diagrams if possible to explain your thinking. Be prepared to troubleshoot scenarios where the issue isn't obvious (e.g., clients get IP addresses but no internet access - think DHCP options, gateway, NAT, DNS). Demonstrate understanding of network security fundamentals without needing to be a security expert. If interviewer asks about technologies you haven't used, acknowledge the gap but explain how you'd approach learning them. Connect infrastructure concepts back to real-world scenarios.
Focus Topics
Maximum Transmission Unit (MTU) and Packet Fragmentation
MTU concept, typical MTU sizes, path MTU discovery, fragmentation issues in VPNs and tunnels, and troubleshooting scenarios where large packets fail to reach destination.
Practice Interview
Study Questions
Network Address Translation (NAT) and Port Forwarding
NAT concepts and mechanisms, source NAT vs destination NAT, port forwarding configuration, how NAT affects traffic flow, and troubleshooting scenarios where services are inaccessible externally despite working internally.
Practice Interview
Study Questions
Network Services and Integration Troubleshooting
Understanding how multiple network services integrate, diagnosing complex scenarios where DNS, routing, firewall, and NAT all play roles, and systematic approach to multi-component troubleshooting.
Practice Interview
Study Questions
DHCP Server Configuration and Troubleshooting
DHCP process (Discover, Offer, Request, Acknowledge), DHCP options (default gateway, DNS servers, subnet mask), DHCP pool exhaustion, scope configuration, and troubleshooting when clients receive IPs but lack connectivity.
Practice Interview
Study Questions
VLAN Configuration and Inter-VLAN Routing
Virtual LAN concepts, tagged vs untagged VLAN traffic, trunk ports vs access ports on switches, inter-VLAN routing configuration, routing between VLANs, and troubleshooting VLAN connectivity issues.
Practice Interview
Study Questions
Firewall Rules, ACLs, and Network Segmentation
Firewall concepts, Access Control Lists (ACLs), stateful vs stateless filtering, inbound vs outbound rules, network segmentation strategies, and how firewall rules impact connectivity.
Practice Interview
Study Questions
Onsite Interview 1: Deep Dive Network Troubleshooting and Analysis
What to Expect
60-minute technical interview conducted onsite (or via video) focusing on complex troubleshooting scenarios and network analysis. You'll work through realistic network problems that require systematic diagnosis, tool usage, and problem-solving. This round assesses depth of knowledge, troubleshooting methodology, and communication of technical concepts. Interview may include whiteboarding network diagrams, drawing packet flows, or discussing command output.
Tips & Advice
Approach problems methodically: state assumptions, ask clarifying questions, propose hypotheses, and describe how you'd test each one. Use whiteboarding to show your thinking visually. Explain what you're doing and why, not just providing answers. Show awareness of monitoring and alerting - how would you detect this issue proactively? Discuss trade-offs in your solutions. If you reach a blocking point, acknowledge it but propose alternative troubleshooting paths. Demonstrate knowledge of enterprise-scale networking (redundancy, failover, monitoring) appropriate for junior level.
Focus Topics
Performance Diagnosis and Optimization
Identifying when network problems are performance-related (latency, throughput, jitter) vs connectivity-related, using monitoring data, identifying bottlenecks, and proposing optimizations.
Practice Interview
Study Questions
Network Monitoring and Observability Concepts
Understanding monitoring strategies, what metrics to track (latency, packet loss, interface errors), alerting thresholds, and how proactive monitoring prevents issues.
Practice Interview
Study Questions
Technology Stack Integration and Multi-Component Issues
Diagnosing complex scenarios involving multiple infrastructure components (DNS + routing + firewall + application), identifying which component is failing, and systematic elimination.
Practice Interview
Study Questions
Connectivity Issue Diagnosis and Resolution
Handling scenarios where services are unreachable, packets aren't reaching destinations, or specific routes are broken. Includes understanding when issues are routing, firewall, NAT, or application-layer.
Practice Interview
Study Questions
Systematic Network Troubleshooting Methodology
Structured approach to network diagnostics: gathering information, forming hypotheses, testing systematically from lower layers upward, eliminating variables, and narrowing root cause.
Practice Interview
Study Questions
Onsite Interview 2: Network Architecture and Infrastructure Design
What to Expect
60-minute technical interview focusing on designing network infrastructure for specific scenarios or applications. You may be asked to design a network architecture for a given scenario (e.g., supporting a growing user base, connecting multiple offices, migrating to cloud). This round assesses understanding of network design principles, scalability, redundancy, security, and cost considerations. You'll discuss trade-offs between design options and justify architectural decisions. For junior level, focus is on understanding design principles rather than making architecture decisions independently.
Tips & Advice
Start by asking clarifying questions about requirements (users, scale, security needs, budget, growth expectations). Build your design step-by-step, explaining each component and why it's necessary. Consider scalability, redundancy, and security proactively. Draw diagrams showing network topology, data flows, and component interactions. Discuss what could go wrong and how you'd mitigate those risks. Ask about constraints and non-functional requirements. For junior-level design questions, focus on understanding WHY certain architectural patterns are used rather than prescribing complex designs. Propose simple, workable solutions and explain how they'd evolve as requirements grow.
Focus Topics
Cloud Networking and Hybrid Architectures
Understanding cloud networking concepts (VPC, subnets, security groups), hybrid on-premises and cloud setups, VPN and interconnect concepts, and implications for infrastructure design.
Practice Interview
Study Questions
Security by Design and Network Segmentation
Incorporating security into network design, DMZs, application segmentation, security zones, firewall positioning, and limiting blast radius of compromises.
Practice Interview
Study Questions
Scalability and Growth Planning
Designing networks that grow with business needs, capacity planning, bandwidth provisioning, and avoiding architectural decisions that create scaling bottlenecks.
Practice Interview
Study Questions
Redundancy, Failover, and High Availability
Designing networks without single points of failure, failover mechanisms, redundant paths, active-active vs active-passive configurations, and availability requirements.
Practice Interview
Study Questions
Network Architecture Design Fundamentals
Basic principles of network design: core, distribution, and access layers; redundancy and failover; performance and availability; security zones and segmentation.
Practice Interview
Study Questions
Onsite Interview 3: Behavioral and Leadership Principles
What to Expect
60-minute interview focused on Amazon Leadership Principles, behavioral competencies, teamwork, communication, and cultural fit. Interviewer will ask behavioral questions using the STAR method (Situation, Task, Action, Result) to assess how you handle challenges, collaborate with others, learn and grow, and approach problems. Questions typically target principles like Customer Obsession, Ownership, Learn and Be Curious, Insist on Highest Standards, Bias for Action, Frugality, and Earn Trust.
Tips & Advice
Prepare 5-7 concrete examples from your experience that demonstrate Amazon Leadership Principles. Use the STAR method: Situation (context), Task (your responsibility), Action (what you did), Result (outcomes). Quantify results when possible (time saved, incidents prevented, cost reduced). Be specific and honest - interviewers can tell when stories are generic or exaggerated. Show ownership, learning orientation, and collaboration. Discuss how you handle feedback and failure. Ask thoughtful questions about how the team embodies these principles. Connect your examples to the Network Engineer role and how you'd apply these principles in this position. Avoid making yourself sound like you know everything; instead show growth mindset and eagerness to learn.
Focus Topics
Problem-Solving Under Pressure and Handling Failure
Examples of high-pressure situations, how you stay calm during production incidents, learning from failures, and preventing similar issues in future.
Practice Interview
Study Questions
Teamwork and Communication
Examples of collaborating with team members, communicating technical information to non-technical stakeholders, handling disagreements professionally, and supporting colleagues.
Practice Interview
Study Questions
Amazon Leadership Principle: Customer Obsession
Understanding end-user needs, designing infrastructure with user experience in mind, considering business impact of reliability and performance, and balancing technical solutions with business goals.
Practice Interview
Study Questions
Amazon Leadership Principle: Learn and Be Curious
Showing growth mindset, examples of learning new technologies or expanding skills, curiosity about how things work, and continuous improvement in your capabilities.
Practice Interview
Study Questions
Amazon Leadership Principle: Ownership
Demonstrating accountability for outcomes, taking initiative without waiting for direction, following through on commitments, and thinking long-term about infrastructure you're responsible for.
Practice Interview
Study Questions
Frequently Asked Network Engineer Interview Questions
A regulated enterprise requires regulated traffic to be physically separated, inspected at the branch, fully captured for audits and strongly encrypted. How does that change your WAN overlay design compared with an ordinary rollout?
Sample Answer
Direct answer
The regulated traffic gets its own overlay: a separate edge device and dedicated physical circuits (two, from different carriers) at each regulated branch, a firewall that inspects locally on the regulated side, a capture tap (a point where a copy of every packet is sent to a recorder) placed before encryption so the recording is readable, and strong IPsec (the standard suite for encrypting IP traffic between sites) between the branches and the regulated hubs. It fails closed: if one regulated circuit is down, regulated traffic moves to the second dedicated circuit, and if both are down it stops; it never slides onto the general internet path (failing open would be the opposite, letting traffic through unprotected). That differs from an ordinary rollout, which shares transport, segments with VRFs (virtual routing tables), inspects centrally and samples logs.
What changes, requirement by requirement
| Requirement | Ordinary rollout | Regulated design |
|---|---|---|
| Physical separation | Logical segments on shared CPE (customer premises equipment, the branch router) and circuits | Separate edge device, switch ports and carrier circuit for the regulated enclave (the isolated part of the network that holds regulated systems) |
| Inspection at the branch | Central or cloud-delivered | Inline firewall on the regulated side in each branch |
| Full capture for audit | Sampled flows and logs | Packet capture of regulated traffic, stored and protected |
| Strong encryption | Default IPsec profile | IKEv2 (the protocol that authenticates the two ends and negotiates the keys) with certificates, AES-256-GCM (a cipher that encrypts and authenticates in one step), perfect forward secrecy (each rekey uses fresh secrets, so one stolen key cannot decrypt past traffic), own key hierarchy (a separate set of keys and signing authority for this overlay) and short rekey interval |
| Failure behaviour | Fall back to another transport | Fail closed, second dedicated circuit from a different carrier |
Assumptions for the numbers below: 60 regulated branches, 15 Mbps average and 60 Mbps peak regulated traffic per branch.
Traffic path at a branch
Regulated hosts connect to regulated switch ports. Their traffic reaches the branch firewall, which inspects it. A capture tap on the firewall's inside interface copies the plaintext to a local recorder. Only after that does the traffic enter the IPsec tunnel on the dedicated circuit. Capturing after encryption would record ciphertext that an auditor could not read, so the tap sits on the plaintext side.
The general-purpose network at the branch has no path into this enclave. There is no route leaking between the regulated and ordinary overlays, and they use separate keys and separate controller policy domains.
Capture sizing
15 Mbps×86,400 s/8=162 GB per day per branch
- Local recorder for 7 days: 162 x 7 = 1,134 GB, so a 2 TB drive pair per branch.
- Central archive for 90 days across 60 branches: 162 x 90 x 60 = 874.8 TB. Treat that as a storage budget item and confirm with the compliance owner whether the rule requires full payload for 90 days or session metadata plus selective capture. Do not assume it.
- Link load: copy the capture to the central archive over the same regulated overlay in a low-priority class. Peak 60 Mbps plus a capture copy that can also reach 60 Mbps is 120 Mbps, so a 200 Mbps regulated circuit runs at 60% in the worst case, not at 100%. That worst case is the surviving circuit carrying everything after the other fails.
The capture archive is itself regulated data: encrypt it at rest, hash files on write for chain of custody, and restrict who can read it.
Encryption and MTU
MTU is the largest packet a link carries, 1,500 bytes on ordinary Ethernet. Encryption wraps every packet in extra bytes, so the packet you can send inside the tunnel is smaller than 1,500. ESP (Encapsulating Security Payload) is the IPsec format that does the wrapping, and tunnel mode means the whole original packet is encrypted and given a new outer IP header. For AES-GCM on IPv4 the added bytes are:
| Part | Bytes | What it is |
|---|---|---|
| Outer IP header | 20 | New header that carries the packet between the two tunnel ends |
| SPI | 4 | Security Parameters Index, a number telling the receiver which tunnel and keys apply |
| Sequence number | 4 | Counter that lets the receiver reject replayed packets |
| IV | 8 | Initialisation vector, per-packet input to GCM so identical data encrypts differently |
| Pad length + next header | 2 | The two trailer bytes (how much padding, and what protocol is inside) |
| ICV | 16 | Integrity check value, the authentication tag that detects tampering |
Total: 20 + 4 + 4 + 8 + 2 + 16 = 54 bytes. On a 1,500-byte circuit the largest inner packet is therefore 1,500 - 54 = 1,446 bytes. ESP may also add 0 to 3 pad bytes so the encrypted part is a multiple of 4 bytes; at 1,446 bytes plus the 2 trailer bytes that is 1,448, already a multiple of 4, so no padding is needed (RFC 4303, RFC 4106). If NAT traversal (wrapping ESP in UDP so it can cross a device doing address translation) is needed, the extra 8-byte UDP header cuts it to 1,446 - 8 = 1,438. Set the tunnel MTU to 1,446 and clamp TCP MSS to 1,446 - 40 = 1,406; the 40 is a 20-byte IP header plus a 20-byte TCP header. MSS clamping means the edge device rewrites the maximum segment size that two hosts agree on during the TCP handshake, so they never build packets too big for the tunnel. With NAT traversal the same steps give an MTU of 1,438 and an MSS of 1,438 - 40 = 1,398.
Trade-offs and pitfalls
- Full inspection with TLS decryption lowers firewall throughput well below the data-sheet headline, so size the appliance from the inspection-enabled figure.
- Dedicated circuits add a circuit and an edge device per regulated branch, with their recurring fees. Reduce the scope by moving regulated work to fewer sites where the rules allow.
- Fail-closed needs a plan: two circuits from different carriers, and a documented procedure for staff when both are down.
- Test the separation: from the ordinary network, confirm there is no route, ARP reachability or management path to the enclave.
Two of your top customers want mutually exclusive behaviors from the product, and both say they will leave if they do not get theirs. How do you decide?
Sample Answer
Direct answer
I would not decide on revenue size or on who is louder. First I find out whether the two demands are truly exclusive or only the two requested solutions are, by getting to the goal behind each request. If the goals can both be met, I design for both. If they truly cannot, I pick the customer that fits the product we are building, tell both of them the decision and the reasoning before it ships, and plan for the loss we have chosen to risk.
Step 1: Test the ultimatum
"We will leave" is a claim, not data. Check the renewal date, how deeply the customer uses the product (active users, integrations they built), what alternatives they have and what switching would cost them. A customer two months from renewal with a competitor pilot is a different situation from one with eighteen months left on contract. Then ask each customer: what are you trying to get done when this behavior matters, and what happens if you do not get it?
Step 2: Move from the behavior up to the goal (illustrative example)
- Customer A, a finance firm, demands that edits to shared dashboards need manager approval before going live.
- Customer B, a media company, demands that analysts publish edits instantly.
- Asking why: A needs auditability (to show an auditor who changed what, and when). B needs speed during breaking news. Neither wants the other's world.
- Design that serves both: a workspace-level setting for publish mode (approval required or instant), with the edit history always on. A gets approvals plus the audit trail. B gets instant publishing, with after-the-fact review available.
- Cost: two modes to test and support and one more settings screen. Compare that with losing both accounts.
Step 3: If the behaviors are truly exclusive, decide on criteria
| Question | Why it matters |
|---|---|
| Which customer is closer to the ideal customer profile (ICP, the kind of customer the company is built to serve best)? | Strategy, not size, should steer the product |
| What do other customers in that segment want? | Two voices are not a market; check request counts and usage data |
| Which option is cheaper to build and easier to reverse? | Reversible choices can be made faster |
| Which keeps the product coherent? | A product that does two contradictory things tends to do neither well |
| Revenue and strategic value (reference customer, expansion) | A tiebreaker, not the decider |
My rule: choose on ICP fit and breadth of demand, and use revenue only to break ties.
Step 4: Communicate
Before any announcement, the product lead and the account's executive sponsor call each customer. Say what we heard, what we decided, why, and what we offer instead (a workaround, an API or integration, a date to revisit). Assume the two customers may compare notes, so both hear the same reasoning.
Pitfalls
- Splitting the difference into a halfway behavior that satisfies neither.
- Building two forks of the product.
- Letting one ultimatum teach every customer that ultimatums work.
What would flip my call
If the setting is cheap, build both. If one customer sits outside the ICP, accept losing them. If usage data shows the "we will leave" is a bargaining position, I treat it as a preference, not a deadline.
Design a migration plan to move a Kubernetes environment from a flat network (where production and non-production share a cluster) to a properly segmented one, using network policies and a service mesh. Include a rollback plan and how you'd continuously verify the new segmentation holds.
Sample Answer
Migrate in observe-then-enforce phases, never straight to enforcement, because the biggest risk in this kind of migration isn't writing the wrong policy, it's not knowing about a dependency until the policy that blocks it goes live.
Migration plan
flowchart LR
A[Discover: passive traffic mapping] --> B[Design target segmentation model]
B --> C[Roll out policies in audit/log mode]
C --> D[Cutover dev namespace]
D --> E[Cutover staging]
E --> F[Canary slice of production]
F --> G[Full production enforce]
G --> H[Continuous verification job]
C -.rollback.-> R[Revert via GitOps]
D -.rollback.-> R
E -.rollback.-> R
F -.rollback.-> R
G -.rollback.-> R
- Discover: run passive traffic observation (mesh telemetry, flow logs, or an eBPF-based tool) for a real baseline period before writing any policy. Never rely on documentation alone: undocumented calls are exactly what a migration like this breaks.
- Design the target model: production and non-production get separate namespaces (or clusters) with an explicit authorization matrix of which services in which environment may call which.
- Audit-mode rollout: deploy the new NetworkPolicies and mesh authorization policies in a permissive, log-only mode where supported, so you see what WOULD be denied without actually denying it yet.
- Staged cutover, lowest risk first: dev namespace, then staging, then a small canary slice of production, watching error rates at each step before moving on.
- Add mesh identity (mTLS) for east-west once network-level segmentation is stable, as a second, independent enforcement layer.
- Continuous verification: keep a scheduled synthetic test (a small client pod that attempts both allowed and disallowed calls on a recurring basis) running permanently, not just as a one-time launch gate.
Rollback plan
Manage policies through GitOps (a workflow where the desired state is a version-controlled manifest that a controller continuously reconciles against) so rollback is a revert of the last commit, applied automatically. Keep a pre-tested "break-glass" wide-open policy that operators can apply directly, with mandatory logging and required approval, for the rare case where waiting on GitOps reconciliation is too slow during an active incident; that path should be alarmed on and time-boxed so it can't quietly become the new normal.
Worked example
Suppose the discovery phase's two-week observation window turns up three legacy scheduled jobs in an unrelated namespace that call a production database service directly, with no clear owner and no entry in the architecture docs. That's the exact kind of finding this phase exists to catch: had the team skipped straight to a deny-all policy, those jobs would have failed silently at 2 a.m. with no immediate connection back to "we just rolled out network segmentation."
Trade-offs and pitfalls
The audit-first approach adds real calendar time before the "real" cutover happens, which is a legitimate cost against the safety it buys. Rollback is easy for the network-policy layer itself, but harder once application code has started depending on the new state, for example if client libraries quietly stop doing their own authentication because they now assume mTLS is always present; rolling back the mesh halfway can leave both layers of auth disabled at once, a regression that's easy to miss because nothing errors loudly. The most common pitfall is letting the log-only phase run indefinitely because "everything looks fine": set a hard date to force the decision to either enforce or explicitly re-scope, or the migration quietly stalls at zero actual enforcement.
Describe man-in-the-middle (MITM) attacks including passive interception and active manipulation techniques (e.g., ARP spoofing, TLS stripping). Explain which network and application-layer logs and telemetry you would examine to detect MITM activity and provide two immediate mitigations for a corporate network.
Sample Answer
A man-in-the-middle (MITM) attack is when an attacker inserts themselves between two communicating parties, either just observing the traffic (passive interception) or actively changing it (active manipulation), without either legitimate party realizing a third party is involved.
Passive vs active techniques
- Passive interception: the attacker taps traffic without altering it, for example by mirroring a switch port, running a rogue access point, or using ARP spoofing purely to redirect traffic through themselves for silent capture. This is a confidentiality violation only.
- Active manipulation: the attacker changes what each side sees.
- ARP spoofing (ARP poisoning): since the Address Resolution Protocol has no authentication, the attacker sends forged ARP replies mapping a real IP (often the default gateway) to the attacker's own MAC address, so victim traffic is routed through the attacker.
- TLS stripping: the attacker sits between the client and server and intercepts the client's initial plaintext HTTP request before it upgrades to HTTPS, keeping the client on cleartext HTTP while the attacker proxies the real connection to the server over HTTPS. The user's browser shows a working connection, but everything between the client and attacker is unencrypted.
Detection telemetry
- Network layer: ARP table and gateway-MAC monitoring (one MAC suddenly claiming an IP it never has, or the gateway's MAC changing), unsolicited/gratuitous ARP replies, rogue DHCP (Dynamic Host Configuration Protocol) servers, switch port-security violations, and traceroute or latency anomalies suggesting an extra hop.
- Application layer: certificate mismatches or downgrade attempts in browser/proxy logs, requests arriving over plain HTTP where HTTPS was expected, violations of HTTP Strict Transport Security (HSTS) policy, and session tokens being used from two different IP addresses or user agents within seconds of each other.
Worked example
Attacker on the same LAN as the victim broadcasts a gratuitous ARP reply: "10.0.0.1 (the gateway) is at AA:BB:CC:DD:EE:FF (attacker's NIC)." The victim's ARP cache updates, so all of the victim's gateway-bound traffic now flows through the attacker first. If the switches run Dynamic ARP Inspection (DAI) against a DHCP-snooped IP-to-MAC binding table, that same forged reply is compared to the trusted binding, found not to match, and dropped and logged instead of being accepted, which is exactly the detection signal described above.
Two immediate mitigations
- Enable Dynamic ARP Inspection plus DHCP snooping on access switches, so ARP replies are validated against a trusted binding table and spoofed ones are dropped at the switch.
- Enforce HSTS and disable any HTTP fallback on externally reachable services, so a stripping attempt has no plaintext path to downgrade to.
Trade-offs and pitfalls
Dynamic ARP Inspection only protects the Layer 2 segment it is deployed on, it does nothing for MITM happening outside your own switches (a public Wi-Fi hotspot, for example). HSTS only protects repeat visits unless the domain is preloaded in browsers, so there is a first-visit gap. Detection controls alone (without prevention) still leave a real exposure window between the attack starting and someone acting on the alert.
Walk through how TCP congestion control evolves during a long-lived connection: slow start, congestion avoidance, fast retransmit, and fast recovery. State which sender-side variable changes at each stage and what event triggers the transition to the next stage.
Sample Answer
Direct answer
Over the life of a connection, TCP's congestion window grows exponentially in slow start, switches to growing linearly in congestion avoidance once it approaches a known safe ceiling, and reacts to loss with fast retransmit and fast recovery rather than always restarting from scratch.
Structured elaboration
- Slow start: the connection begins with a small congestion window (historically 1 segment; modern stacks start higher, commonly around 10 segments per RFC 6928) and roughly DOUBLES the window every round trip, since each of the ACKs for the previous batch triggers sending two new segments. This continues until either loss occurs, or the window reaches a threshold called
ssthresh(slow start threshold), at which point the sender switches strategies. - Congestion avoidance: once at or above
ssthresh, growth switches from exponential to roughly linear (classically, additive increase of about one segment per round trip), a much more cautious probe for additional capacity. - Fast retransmit: if the sender sees three duplicate ACKs (the receiver repeatedly acknowledging the same byte, implying a specific segment is missing but LATER data did arrive), it retransmits the missing segment immediately, without waiting for the retransmission timer to expire, since three duplicate ACKs is strong, specific evidence of loss rather than simple reordering.
- Fast recovery: after a fast retransmit, rather than collapsing all the way back to slow start,
ssthreshis set to about half the current window, and the window itself is set near that halved value, so the sender doesn't have to re-earn all its previous progress from a window of one segment; it resumes near where it estimates the path can actually sustain.
Worked example
Picture a connection whose window has grown to 64 segments in flight when a single segment is lost and detected via three duplicate ACKs (not a full timeout). Fast retransmit resends the missing segment immediately. Fast recovery sets ssthresh to roughly 32 (half of 64) and the window to near that value, then resumes congestion avoidance's linear growth from there, rather than collapsing to slow start's small initial window and doubling all the way back up. Contrast this with a RETRANSMISSION TIMEOUT (no duplicate ACKs arrived at all, meaning the loss was severe enough that the whole flight of data went missing): that's a much stronger loss signal, and the sender resets ssthresh to half the current window but drops the actual window all the way back to slow start's minimum, since a timeout implies the path may be far more broken than a few duplicate ACKs would suggest.
Trade-offs & pitfalls
It's a common mistake to say TCP always halves its window on any loss and moves on; a full retransmission timeout is treated far more conservatively (full reset to slow start) than a fast-retransmit-detected loss (a much gentler recovery), because the ABSENCE of any duplicate ACKs at all is itself informative: it suggests either a much larger loss event or a badly congested/broken path, not just one unlucky dropped segment.
Tell me about a time internal or external pressure, such as a deadline, a client, or a business commitment, pushed you toward a decision that conflicted with a principle or value your company had explicitly committed to (for example privacy, security, or data quality). Walk through how you recognized the conflict, what you did about it, how you communicated your position to stakeholders, and what the final outcome was.
Sample Answer
Direct answer
When a deadline, a client, or a business ask pushes toward something that conflicts with a principle a company has committed to, such as privacy, security, or data quality, the strongest answers show three things: you noticed the conflict explicitly rather than complying without registering it, you raised it through the right channel rather than either silently complying or unilaterally blocking the work, and you drove toward a resolution rather than just splitting the difference.
Structured elaboration
- Notice: name the specific moment you recognized the tension, and what concrete detail made you pause.
- Raise it: describe how you raised it, ideally backed by data or a concrete risk rather than an appeal to principle alone. A values-based objection lands far better when it is backed by the actual risk it protects against.
- Navigate: what you actually did in the interim, whether you proposed a compromise or a phased approach, who you looped in, and how you kept the relationship functional even while disagreeing.
- Outcome: what actually happened. An honest outcome, including "I was overruled and here is what I did next," is often more credible than a suspiciously clean win.
Worked example
A team was under pressure to ship a change quickly, and the fastest path meant skipping a validation step that existed specifically to catch a known class of data-quality problem. Rather than quietly skipping it or unilaterally blocking the release, the response was to time-box a reduced version of the validation, checking the highest-risk subset in the time available, and to flag explicitly and in writing what wasn't covered and what the residual risk was, so the decision to accept that risk was made deliberately by the right people rather than by default. The release shipped on time, and the flagged gap was closed within the following two days as agreed, rather than being silently forgotten.
Trade-offs and pitfalls
A story where you unilaterally blocked the work and were later vindicated can read as inflexible if it doesn't also show you understood the business pressure; the strongest answers show empathy for that pressure while still holding the line. A story where you quietly went along with the shortcut is not really an example of this competency at all; the action needs to show you actively surfaced the tension, not merely noticed it internally. Vague appeals to "our values" without a concrete risk attached tend to land weaker than a specific technical or business risk, clearly stated.
Tell me about a time you had to communicate a project risk, delay, or scope change to stakeholders. How did you frame the message, what options did you present, and how did you protect trust?
Sample Answer
Situation: On a prior project, we uncovered a late dependency issue that would push a release by a few weeks.
Task: I needed to tell stakeholders early, explain the impact clearly, and keep trust intact.
Action: I didn’t wait until we had perfect data. I shared the risk as soon as the pattern was clear, framed it around business impact, and presented options rather than just the problem. I explained what was affected, what was still on track, and what we could do next: reduce scope, add temporary support, or adjust the release sequence. I also set a short update cadence so no one had to guess.
Result: The group made a quick decision on scope, leadership appreciated the early warning, and the conversation stayed focused on trade-offs instead of blame. The key was being direct, specific, and calm.
What I learned is that trust is protected by speed, honesty, and a recommendation. If I bring a risk with a clear path forward, stakeholders usually stay engaged instead of feeling surprised or managed around.
The firewall team and the network team keep using the word 'NAT' for different things. Walk me through the kinds of address translation you meet in an enterprise network, what each one rewrites, and how each affects logging, troubleshooting and stateful firewalls.
Sample Answer
Direct answer
NAT (network address translation) rewrites IP addresses, and sometimes ports, in packet headers as traffic crosses a device. The kinds differ in what they rewrite and in which direction: static NAT (one-to-one), dynamic NAT (pool), PAT (port address translation, many-to-one), destination NAT or port forwarding, twice NAT, and NAT64/NPTv6 for IPv6. When teams say "NAT" they usually mean PAT outbound (firewall team) or static/destination NAT inbound (network team).
Kinds and what each rewrites
| Type | Rewrites | Typical use |
|---|---|---|
| Static NAT | Source (outbound) or destination (inbound), fixed 1:1 | Publish a server on a dedicated public address |
| Dynamic NAT | Source, from a pool, 1:1 for the session's life | Rare today; limited by pool size |
| PAT / NAPT (overload) | Source address and source port | Everyone behind one public address |
| Destination NAT (port forwarding) | Destination address and port | Public 203.0.113.10:443 to internal 10.1.5.20:8443 |
| Twice NAT | Source and destination both | Overlapping address spaces after a merger |
| NAT64 (with DNS64) | IPv6 to IPv4 packets | IPv6-only clients reaching IPv4 servers |
| NPTv6 (RFC 6296) | IPv6 prefix only, stateless | Provider-independent prefix mapping |
PAT and destination NAT appear on nearly every enterprise firewall. Twice NAT rewrites both addresses of one packet, which is what makes two networks that use the same range reachable from each other. NAT64 lets an IPv6-only client reach IPv4 servers, with DNS64 (a DNS service that invents IPv6 answers for names that only have IPv4 addresses) pointing the client at the translator. NPTv6 swaps one IPv6 prefix for another and leaves the rest of the address and the ports alone, so it keeps no per-connection state. An ISP may add carrier-grade NAT (CGNAT) using 100.64.0.0/10 (RFC 6598), a second PAT layer outside your control.
Effect on logging
A log entry contains only one side of the translation. The firewall sees the inside address, an upstream server sees the translated one, so you correlate with the NAT table or NAT log (inside IP and port, outside IP and port, timestamps). With PAT many inside hosts share one outside address, so the source port and an exact timestamp (clocks synchronized with NTP) are required to name the host. Without NAT logging, an abuse complaint naming the public IP cannot be traced to a person.
Effect on troubleshooting
- Capture on both sides of the translating device and compare addresses.
- Inspect the live table:
show ip nat translationson Cisco IOS,conntrack -Lon Linux (conntrack is the Linux kernel's connection-tracking table, which stores each flow and its translation). - Check hairpin cases (also called NAT loopback: an inside client reaching an inside server by its public address, so the packet must be translated and sent back into the same network).
- PAT port exhaustion shows up as new connections failing while old ones work.
- Protocols that carry IP addresses inside the payload (active FTP, SIP) need an ALG (application-layer gateway, a helper in the firewall that rewrites those embedded addresses) or they fail. IPsec ESP does not survive PAT without NAT-T (NAT traversal, UDP 4500).
Effect on stateful firewalls
A stateful firewall tracks each flow in a connection table and permits replies automatically. The translation is stored with the flow, so return traffic is un-translated and matched. Which address a rule must name depends on the platform's order of operations. On Linux netfilter (the kernel packet-processing framework behind iptables and nftables), PREROUTING is the first hook a packet meets on arrival and POSTROUTING is the last before it leaves. Destination NAT happens in PREROUTING, before the filter, so a forward rule matches the already-translated internal destination, while source NAT occurs in POSTROUTING, after the filter, so the rule matches the original inside source. Concretely (illustrative addresses): an internet client sends a packet to 203.0.113.10:443, and a port forward rewrites the destination to 10.1.5.20:8443 before filtering. The allow rule therefore has to name 10.1.5.20 port 8443; a rule written for 203.0.113.10 port 443 never matches. For outbound PAT, the filter sees the packet while the source is still 10.1.5.20, so that rule names 10.1.5.20, not 203.0.113.10. Other vendors differ (some match pre-NAT addresses, some real addresses), so read the product's packet flow documentation before writing policies, because a rule written against the wrong side silently never matches.
Worked example
Inside host 10.1.5.20 opens a connection from port 51000 to 198.51.100.7:443 through PAT address 203.0.113.10. The firewall rewrites the source to 203.0.113.10:62001 and records the pair. The server's log shows 203.0.113.10:62001. The firewall's own table maps that back to 10.1.5.20:51000, so only the two logs together identify the host.
Pitfalls
- No NAT logging retained, so incidents cannot be attributed.
- Treating NAT as a security control. Security comes from the stateful policy, not from address hiding.
- Overlapping ranges with no twice NAT plan.
Explain how an Ethernet switch learns MAC addresses and decides where to send a frame. Cover what happens when the destination is unknown, and what that flooding means for a large broadcast domain.
Sample Answer
Direct answer
A switch learns by reading the source MAC address of every frame it receives and recording "this address lives behind this port, in this VLAN" in its MAC address table (also called the CAM table, after content-addressable memory, the fast lookup hardware that holds it). It forwards by looking up the destination address. A known unicast destination goes out one port only. An unknown unicast destination, a broadcast, or a multicast (unless the switch filters multicast, for example with IGMP snooping) is flooded out every other port in the same VLAN. A MAC address is the 48-bit hardware address of a network card.
How learning and forwarding work, in order
- A frame arrives on port Gi1/0/5 in VLAN 10 with source MAC
aaaa.aaaa.aaaa. The switch adds or refreshes the entry (VLAN 10,aaaa.aaaa.aaaa, Gi1/0/5) and resets that entry's age timer. - It looks up the destination MAC in the same VLAN's table. The table is effectively per VLAN: the same MAC may appear in two VLANs on two ports.
- Three outcomes:
- Known, on a different port: forward out that one port only.
- Known, on the port the frame arrived on: filter (drop). The destination is behind the same port, for example a hub or a downstream switch, and already saw the frame.
- Unknown: flood out all ports in that VLAN except the one it arrived on (the ingress port, meaning the port a frame enters by).
- Entries age out. The usual Cisco default is 300 seconds, set with
mac address-table aging-time <seconds>(check the platform's command reference for the permitted range and for what a value of 0 does). Static entries (mac address-table static) never age.
Per-entry fields and special destinations
| Item | What the table or switch holds | Why it matters |
|---|---|---|
| VLAN | The VLAN ID the address was learned in | Same MAC can exist in several VLANs |
| MAC address | 48-bit address from the frame source field | The lookup key |
| Port | Ingress port (or port-channel) | Where to send known unicast |
| Type | Dynamic (learned, ages) or static (configured, never ages) | Static pins a server or blocks moves |
| Age | Time since the entry was last refreshed by a frame from that source | Idle entries expire after the aging time |
- Broadcast (destination
ffff.ffff.ffff) is never learned as a source and is always flooded within the VLAN: every host must process it. - Multicast, where the switch does no multicast filtering, is flooded within the VLAN, because no multicast address ever appears as a source. A switch can be given a feature to restrict it to interested receivers (IGMP snooping, where the switch listens to hosts' group-join messages), which is outside the table described here.
What flooding means for a large broadcast domain
A broadcast domain is every port that receives a broadcast sent by any one host, which on a switch is one VLAN. Flooding turns each unknown unicast, broadcast and multicast frame into N-1 copies (if the VLAN has N active ports, one copy leaves every port except the one the frame arrived on, so a 24-port VLAN produces 23 copies of every flooded frame). Every host in the VLAN spends CPU on broadcasts (ARP, DHCP discovery, some discovery protocols), and every switch-to-switch link carries all of it. If the topology has a layer 2 loop that spanning tree has not blocked, flooded frames circulate forever and multiply (a broadcast storm), and the MAC table thrashes because the same source address keeps appearing on different ports.
Worked example: unknown unicast and aging
Host A (aaaa.aaaa.aaaa, port 1) sends to host B (bbbb.bbbb.bbbb, port 2) on a fresh 4-port switch in one VLAN.
- A's first frame is an ARP request to the broadcast address. The switch learns A on port 1, then floods to ports 2, 3 and 4.
- B replies, unicast to A. The switch learns B on port 2 and, because A is known, sends the reply out port 1 only. Ports 3 and 4 never see it.
- A now sends to B: both are known, so each frame goes out exactly one port.
- If B stays silent for 300 seconds (the default aging time) its entry expires. A's next frame to B is an unknown unicast and floods again until B transmits something.
Mitigating excessive flooding
- Shrink the flood scope: keep broadcast domains small. One VLAN per function, with a layer 3 boundary between them, bounds the flood scope. A /24 (254 usable addresses) per VLAN is a common ceiling in campus designs because broadcast load grows with host count.
- Prevent loops: keep spanning tree healthy so flooded frames cannot circulate. On user ports enable PortFast (the port goes straight to forwarding, skipping the listening and learning states) together with BPDU guard (if the port ever receives a spanning-tree BPDU, meaning a switch has been plugged in, the switch shuts the port into the err-disabled state). A stray switch then cannot change the topology.
- Stop table flooding attacks: cap MAC addresses per access port. Port security, the switch feature that limits how many and which MAC addresses a port may learn, has a default limit of one secure MAC address per port, raised with
switchport port-security maximum <value>. The default violation mode is shutdown (error-disabled), and restrict or protect modes drop offending frames instead. This also blunts a MAC flooding attack, where an attacker fills the table with fake sources so legitimate destinations become unknown and are flooded. - Avoid asymmetric paths that leave a destination unlearned. A switch learns only from frames it sees B send. Suppose A sends to B through switch sw-1, but B's replies travel back over a different path that never passes through sw-1. sw-1 never sees B as a source, so it never learns B's port and floods every frame addressed to B, for as long as the asymmetry lasts. Fix the path so replies cross sw-1, or add a static entry for B.
- Size the aging time to the traffic. Lengthening it above 300 seconds reduces re-flooding for quiet hosts but slows stale-entry cleanup after a host moves.
Trade-offs and pitfalls
- Learning is from the source only. A host that never transmits is never learned and always gets flooded traffic.
- A host that moves ports is relearned on its next frame. If the old entry has not aged, traffic misdirects only until that frame arrives.
- A MAC table has finite capacity. When a table is full, behaviour on new sources depends on the platform, so check the datasheet rather than assuming.
A link between two routers flaps intermittently. How would you detect and alert on flapping reliably, avoid paging for every bounce, and tell a flap from a genuine outage?
Sample Answer
Direct answer
Treat "link is down" and "link is flapping" as two different alerts built from different questions. The rules below are written for Prometheus (a monitoring system that stores metrics as time series and evaluates alert expressions) and routed by Alertmanager (its companion that groups, deduplicates and delivers the notifications). An outage is one continuous state: the interface has been down for a set time (3 minutes here). A flap is a pattern: six or more state changes inside 15 minutes (three down-and-up bounces), whatever state it is in now. Page a human for the outage. Open a ticket for the flap, hold it open through the brief up periods so it does not resolve and re-fire every bounce, and suppress it while a real outage is already paging for the same interface. Then find the cause by comparing both ends, the error counters and the optics, because a flap is usually physical.
Detecting reliably: three signals, each with a blind spot
| Signal | What it gives | Blind spot |
|---|---|---|
| Polled state (ifOperStatus from the interface MIB, 1 = up, 2 = down, polled every 30 s) | Easy to alert on and graph | A bounce shorter than the poll interval can leave the state "up" at both polls |
| ifLastChange (the sysUpTime value when the interface entered its current state, RFC 2863) | Moves on every transition, so a short bounce still shows | Several bounces between two polls look like one change. It resets when the device reinitializes, giving one false change |
| Event stream: linkDown and linkUp traps, syslog, or a gNMI (gRPC Network Management Interface) ON_CHANGE subscription (the device sends an update only when a value changes) | One event per transition, with the device's own timestamp | Traps and syslog are commonly sent over UDP and can be lost, so reconcile with polling |
Use polled state for alerting, and the event stream as the exact count and timeline when you investigate. Do not treat an absence of polled changes as proof of a stable link.
Alert rules
groups:
- name: link-health
rules:
- alert: LinkDown
expr: ifOperStatus{ifAlias=~"UPLINK:.*"} == 2 and on (instance, ifIndex) ifAdminStatus == 1
for: 3m
labels:
severity: page
annotations:
summary: "{{ $labels.instance }} {{ $labels.ifName }} down for 3 minutes"
- alert: LinkFlapping
expr: changes(ifOperStatus{ifAlias=~"UPLINK:.*"}[15m]) >= 6
keep_firing_for: 20m
labels:
severity: ticket
annotations:
summary: "{{ $labels.instance }} {{ $labels.ifName }} changed state {{ $value }} times in 15 minutes"
changes()is a Prometheus function that returns the number of times a series' value changed in the window. Each down and each up counts as one, so 6 means three bounces.- The
for: 3mclause keeps an alert pending until the condition has held across 3 minutes of evaluations. A flap whose down periods are shorter than 3 minutes therefore never pages as an outage. keep_firing_for: 20mkeeps the flap alert firing after its condition last held, so the up phases between bounces do not resolve it and send another notification.ifAdminStatus == 1(administratively enabled) excludes ports an operator shut on purpose: a port that is administratively down also shows operational state down.and on (instance, ifIndex) ifAdminStatus == 1joins two series: theon (...)list says to pair samples that have the same device and interface index, so each interface's operational status is compared with its own admin status.- The
ifAlias=~"UPLINK:.*"selector assumes a naming convention that marks monitored links in the interface description. Use whatever your source of truth makes reliable.
Worked check: a 30-second poll series 1 1 2 1 1 1 2 1 1 2 1 1 (1 up, 2 down) holds one sample per 30 s, at t = 0, 30, 60 and so on to 330 s. The values change at t = 60, 90, 180, 210, 270 and 300 s, so the sixth change is at t = 300 s, and the series has six changes by the 5-minute mark, so LinkFlapping fires at 5 minutes with value 6. A link that goes down once and stays down has one change and never reaches the flap threshold, but fires LinkDown 3 minutes after it fell. A link whose state is 2 but whose ifAdminStatus is 2 fires neither.
Not paging for every bounce
route:
receiver: noc
group_by: [instance]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- matchers: ['severity="page"']
receiver: pager
inhibit_rules:
- source_matchers: ['alertname="LinkDown"']
target_matchers: ['alertname="LinkFlapping"']
equal: [instance, ifName]
- Inhibition means a firing alert can silence a related one. Grouping by
instancemeans a line card failure that downs 24 links produces one notification for the device, not 24.group_wait(default 30 s) holds the first notification so related alerts arrive together,group_interval(default 5 minutes) spaces later updates for the same group, andrepeat_interval(default 4 hours) spaces reminders. - The inhibition rule suppresses
LinkFlappingfor an interface whileLinkDownis firing for the same instance and interface, because the outage page already covers it. - Only
severity="page"goes to the pager. The flap goes to the network operations center (NOC) queue as a ticket, so a link that bounces ten times overnight leaves one ticket, not ten pages.
Telling a flap from a genuine outage
| Observation | Flap | Outage |
|---|---|---|
| Operational state over 15 min | Repeated 1, 2, 1, 2 | One transition to 2 and staying there |
| Routing adjacencies (the neighbor sessions between routers) and traffic | Come and go, reconverge (recompute their routes) each time | Gone until fixed, traffic on the backup path |
| Error counters before the event | Often climbing (CRC, cyclic redundancy check, a checksum failure that signals a corrupted frame, and other input errors) | May be clean (a cut or power loss) |
| Alert | LinkFlapping, ticket | LinkDown, page |
Both can be true in sequence: a flapping link can end as an outage when the optic or cable finally fails, which is why the flap ticket matters.
Finding the cause, in order
- Compare both ends' logs. The same instant on both sides suggests the cable, patch panel (a rack of sockets where cables are cross-connected) or the optical path between them. One side logging link loss while the other stays up suggests that side's port or transceiver, or the far end's transmitter. The transceiver (also called the optic) is the pluggable module that turns the switch's signal into light or electrical signal on the cable, and a line card is a removable board holding a group of ports.
- Read the error counters. ifInErrors rising before each bounce points at signal quality (dirty fiber, damaged cable, bad optic).
- Read the optical power levels the transceiver reports (digital optical monitoring, DOM): transmit or receive power drifting near the limits explains flaps without any counter errors.
- Look at the timing. Intervals that repeat almost exactly (say every 40 seconds) suggest a protocol or timer, a negotiation loop or a power cycle. Random intervals suggest a physical fault.
- Change one thing. Swap the patch cable, then the optic, then move the port, and watch that the flap count over the next 15 minutes stays under the threshold.
Verify the fix
Run changes(ifOperStatus{...}[1h]) and expect 0 over a full hour, check the error counters are not rising, and let the keep_firing_for period elapse so the ticket clears on its own.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Network Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs