Amazon Network Engineer (Entry Level) - Comprehensive Interview Preparation Guide
Amazon's interview process for entry-level technical roles typically begins with a recruiter screening call, followed by a technical phone screen, and concludes with an onsite loop of 4-5 rounds. For entry-level network engineers, this includes technical interviews focusing on networking fundamentals and troubleshooting, one infrastructure/architecture-focused round, and dedicated behavioral rounds assessing alignment with Amazon's Leadership Principles. All behavioral assessments use the STAR method (Situation, Task, Action, Result) and reference Amazon's 16 Leadership Principles.
Interview Rounds
Recruiter Screening
What to Expect
Initial 20-30 minute call with an Amazon recruiter to assess basic fit, background, and role interest. The recruiter will ask about your experience with networks, your motivation for the role, relocation willingness, and salary expectations. This is a preliminary fit check before moving to technical evaluation. Success here depends on clear communication, demonstrated interest in networking, and alignment with Amazon's culture. No technical depth is expected at this stage.
Tips & Advice
Be clear and concise about your background and why you're interested in networking. Mention any relevant coursework, certifications (CompTIA Network+, Cisco basics), or projects you've completed. Research Amazon's cloud infrastructure and express genuine interest in supporting large-scale systems. Prepare 2-3 questions about the role, team structure, and growth opportunities. Maintain enthusiasm and professionalism. Have your resume and any relevant links ready. If you lack formal networking experience, emphasize your learning ability and foundational technical knowledge.
Focus Topics
Relevant Certifications and Foundations
Mention any relevant certifications (CompTIA Network+, Cisco CCNA basics, AWS Fundamentals), online courses, or technical projects that demonstrate foundational networking knowledge.
Practice Interview
Study Questions
Amazon Culture and Leadership Principles Awareness
Demonstrate familiarity with Amazon's Leadership Principles (especially Ownership, Bias for Action, and Learn and Be Curious) and explain how they resonate with your professional values.
Practice Interview
Study Questions
Professional Background and Motivation
Clearly communicate your educational background, any networking experience (academic, internships, personal labs), and genuine reasons for pursuing a network engineer role at Amazon.
Practice Interview
Study Questions
Technical Phone Screen
What to Expect
A 45-60 minute technical screening with a senior engineer or hiring manager. This round tests foundational networking knowledge through a mix of conceptual questions, troubleshooting scenarios, and potentially a simple hands-on problem. Expect questions covering OSI model layers, TCP/IP fundamentals, common network tools, and basic troubleshooting workflows. You may be asked to walk through how you'd diagnose a connectivity issue or explain how specific networking components work. This is not a deep-dive but validates you have entry-level competency and can think methodically about network problems.
Tips & Advice
Structure your answers methodically, explaining your reasoning step-by-step. When presented with a troubleshooting scenario, start by clarifying the symptoms, then systematically eliminate possibilities (e.g., 'First I'd verify if the device has connectivity at all, then check DNS separately from routing'). Use correct terminology but don't overcomplicate—clarity matters more than jargon. For any tool-related question, briefly explain what common tools do (ping, traceroute, netstat, ss, dig, arp). If you're unsure about an answer, say so and explain how you'd find the answer. Interviewers value honest uncertainty over incorrect confidence. Practice explaining networking concepts out loud before the call to build fluency.
Focus Topics
DNS Fundamentals and Resolution Troubleshooting
How DNS works, common record types (A, AAAA, CNAME, MX), the DNS resolution process, tools for diagnosing DNS issues (dig, nslookup, host), and recognizing when DNS is the problem vs. general connectivity.
Practice Interview
Study Questions
Firewalls, NAT, and Port Forwarding Basics
Basic understanding of how firewalls block traffic, Network Address Translation (NAT) and when it's used, port forwarding concepts, and recognizing when firewall rules might be blocking connectivity.
Practice Interview
Study Questions
IP Addressing, Subnetting, and CIDR Notation
Fluency with IPv4 and IPv6 addressing, subnet masks, CIDR notation, address classes, private vs. public ranges, and ability to calculate subnets and determine if two IPs are on the same subnet.
Practice Interview
Study Questions
Network Troubleshooting Tools and Commands
Practical knowledge of tools used in the job description: ping, traceroute, netstat, ss, dig (DNS), arp, ip, ifconfig, curl, nc (netcat). Understanding what each tool shows and when to use it.
Practice Interview
Study Questions
Routing and Network Path Troubleshooting
Understanding how packets are routed via routing tables, the role of default gateways, how traceroute works, recognizing when routing is the problem vs. other issues, and interpreting routing table output.
Practice Interview
Study Questions
OSI Model and TCP/IP Stack Fundamentals
Solid understanding of the seven OSI layers, how TCP/IP maps to these layers, and which protocols operate at which layers (e.g., DNS at application layer, TCP at transport, IP at network, Ethernet at data link).
Practice Interview
Study Questions
Onsite Technical Round 1 - Networking Fundamentals and Protocol Knowledge
What to Expect
First of four onsite rounds, this 55-60 minute session with a senior network engineer dives deeper into networking protocols, concepts, and fundamental troubleshooting. Expect a mix of conceptual questions ('explain how ARP works'), scenario-based problems ('if a host can reach the internet but not a specific subnet, what's broken?'), and some hands-on troubleshooting discussion. This round assesses your depth of foundational knowledge and logical thinking about network problems. You may be asked to whiteboard or sketch a simple network scenario and explain how traffic flows through it.
Tips & Advice
Be prepared to explain networking concepts at multiple levels of abstraction—from 'here's the big picture' to 'here's how the bits move.' When discussing protocols or troubleshooting scenarios, draw or describe diagrams if possible; visualizing helps. For troubleshooting scenarios, explicitly state your assumptions and walk through your diagnostic steps logically. If a scenario mentions 'clients receive IPs but have no internet,' break it into distinct problems (DHCP works, but routing or NAT might be broken) and explain how to test each. Interviewers value methodical thinking over immediately jumping to answers. If you make a mistake, acknowledge it and correct yourself.
Focus Topics
VLANs and Inter-VLAN Routing
VLAN basics, trunk vs. access ports, how traffic is separated by VLAN, how inter-VLAN routing is configured, and common misconfigurations that break VLAN communication.
Practice Interview
Study Questions
DHCP and IP Address Assignment Process
How DHCP works (DORA: Discover, Offer, Request, Acknowledge), DHCP options and what they do (default gateway, DNS servers, etc.), common DHCP issues, and how to verify DHCP functionality.
Practice Interview
Study Questions
ICMP, Ping, and Traceroute Mechanics
How ICMP works, what ping actually tests (reachability, not necessarily application connectivity), how traceroute traces the path packets take, and interpreting traceroute output.
Practice Interview
Study Questions
Ethernet, ARP, and Data Link Layer Operations
Deep understanding of how Ethernet frames work, MAC addresses, ARP protocol and how devices discover each other on a local network, VLAN basics, and data link layer switching.
Practice Interview
Study Questions
TCP and UDP Protocols - Differences and Use Cases
Understanding the differences between TCP (connection-oriented, reliable) and UDP (connectionless, fast), when each is used, how TCP establishes connections (three-way handshake), and port concepts.
Practice Interview
Study Questions
Common Network Issues and Troubleshooting Methodologies
Recognizing common scenarios: host on wrong subnet, MTU mismatches, duplicate IPs, misconfigured default gateway, DHCP not working, DNS resolution failing, firewall blocking ports. Systematic approaches to isolating each type of issue.
Practice Interview
Study Questions
Onsite Technical Round 2 - Network Configuration, Design, and Hands-On Scenarios
What to Expect
A 55-60 minute round focused on practical configuration knowledge and network design at an entry-level scope. This round tests your understanding of how to configure routers and switches (conceptually and with configuration syntax), design simple network architectures, and solve multi-layered troubleshooting problems. You may be given a scenario like 'design a network for an office with 50 employees that connects to the internet via a single ISP' and asked to explain your architecture, including addressing scheme, routing strategy, and security considerations. Configuration syntax questions may reference industry-standard formats (Cisco IOS or similar) but focus on concepts over memorization.
Tips & Advice
When designing a network, start by clarifying requirements (number of hosts, performance needs, security requirements) before diving into solutions. Explain your design decisions clearly ('I chose this addressing scheme because...', 'I'd segment this network this way to isolate critical systems...'). For configuration questions, focus on understanding what the configuration does, not memorizing exact syntax—explain the purpose of each component. If you don't recall exact commands, explain the logical steps and what you'd need to verify the configuration. For troubleshooting scenarios with multiple layers (DHCP issue, then routing issue, then firewall issue), tackle them systematically. Draw network diagrams to illustrate your thinking. Show awareness of security (firewalls, access controls) as mentioned in the job description.
Focus Topics
Capacity Planning and Scalability Basics
Understanding how to estimate bandwidth needs, recognize when network capacity is approaching limits, and plan for growth. Recognizing when infrastructure needs upgrades.
Practice Interview
Study Questions
Network Monitoring and Performance Metrics
Understanding what to monitor (bandwidth usage, latency, packet loss, interface errors, device availability), common monitoring tools, and how to interpret performance data to identify problems.
Practice Interview
Study Questions
Network Security Implementation Basics
Firewall rule concepts, access control lists (ACLs), network segmentation strategies, basic principles of least privilege, and recognizing when security controls are necessary (e.g., isolating sensitive systems).
Practice Interview
Study Questions
Router and Switch Configuration Concepts
Understanding basic configuration of routing and switching devices: interface configuration, IP addressing on interfaces, static routing, VLAN configuration, access control lists (ACLs), and basic device management.
Practice Interview
Study Questions
Multi-Layer Troubleshooting Scenarios
Complex scenarios requiring diagnosis at multiple layers: 'Clients can't reach a service; it might be routing, firewall, NAT, DNS, or the service itself.' Systematic approaches to narrowing down the root cause.
Practice Interview
Study Questions
Network Architecture Design for Entry-Level Scenarios
Designing small-to-medium networks (office, branch office, small campus). Choosing addressing schemes, planning subnetting, deciding on routing strategies (static vs. dynamic), segmenting networks with VLANs, and incorporating basic redundancy.
Practice Interview
Study Questions
Onsite Technical Round 3 - Infrastructure, Operations, and Tools
What to Expect
A 55-60 minute round with an infrastructure or operations engineer covering practical tools, operational considerations, and real-world challenges. This round tests familiarity with network management tools, monitoring platforms, documentation practices, and operational workflows. You may be asked about tools for network analysis (Wireshark, netflow, SNMP), configuration management, change control processes, or how you'd approach documenting network changes. Questions focus on the 'daily activities' mentioned in the job description: monitoring, configuration, troubleshooting, security updates, and documentation.
Tips & Advice
Discuss your familiarity with network management and monitoring concepts. If you've used tools like Wireshark, SNMP, or netflow (or cloud equivalents like CloudWatch), mention them. If not, explain how you'd learn them quickly. Emphasize the importance of documentation and change control—demonstrating operational maturity. Talk about your experience with Linux/Unix command-line tools since the job involves network monitoring and configuration. Discuss how you'd approach a security update deployment with minimal downtime. Show awareness of operational best practices like testing changes in non-production environments, maintaining change logs, and communicating with stakeholders. If you lack production experience, discuss lab setups you've built and how you'd apply those practices.
Focus Topics
Logging, Alerting, and Incident Response Basics
Understanding how to set up alerts for network issues, interpreting logs from network devices, structured incident response workflows, and post-incident documentation.
Practice Interview
Study Questions
Cloud Network Concepts (AWS/Azure/GCP Basics)
Basic familiarity with cloud networking: virtual networks, security groups, network ACLs, VPC concepts, cloud DNS, and how cloud networks differ from on-premises. AWS-specific if applicable.
Practice Interview
Study Questions
Security Updates and Patch Management
Understanding how to apply security updates to network equipment, planning patch deployment with minimal downtime, testing patches before production deployment, and recognizing critical vs. non-critical updates.
Practice Interview
Study Questions
Network Configuration Management and Documentation
Best practices for maintaining network documentation, version controlling configurations, change management processes, implementing changes safely, testing procedures, and rollback strategies.
Practice Interview
Study Questions
Linux/Unix Command-Line Fundamentals for Network Operations
Proficiency with command-line tools: ip, ifconfig, netstat, ss, route, traceroute, dig, tcpdump, iptables/firewall commands, basic shell scripting for network tasks, and SSH for remote management.
Practice Interview
Study Questions
Network Monitoring and Analysis Tools
Familiarity with tools for network monitoring: packet analyzers (Wireshark), flow data (netflow/sflow), SNMP, and system monitoring tools. Understanding what data each tool provides and when to use each.
Practice Interview
Study Questions
Onsite Behavioral Round - Amazon Leadership Principles
What to Expect
A 55-60 minute round, often with a 'Bar Raiser' (senior interviewer ensuring high hiring standards), dedicated entirely to behavioral assessment. This round evaluates your alignment with Amazon's 16 Leadership Principles through structured questions about past experiences. Expect 5-7 questions, each asking for specific examples using the STAR method (Situation, Task, Action, Result). Common topics include times you owned a problem, delivered results despite obstacles, learned from failure, disagreed with a manager, prioritized customer/user needs, or collaborated across teams. For a network engineer role, examples might involve troubleshooting a critical outage, implementing a complex configuration change, or improving network reliability.
Tips & Advice
Prepare 5-7 concrete examples from your background (academic projects, internships, personal projects, volunteer work) that align with different Leadership Principles. Each example should be a real situation where you took action and achieved a result. Quantify outcomes when possible ('reduced downtime by 30%', 'trained 5 junior colleagues'). Practice telling these stories concisely (2-3 minutes each) and be prepared to go deeper if asked. For entry-level candidates, examples don't need to show large-scale impact—focus on demonstrating learning, ownership of small tasks, and collaboration. Explicitly mention the Leadership Principle you're demonstrating ('This is an example of ownership because...'). Listen carefully to each question; sometimes a single story can address multiple principles. Avoid generic answers—specific details make stories credible. Prepare examples that connect to the job description, e.g., 'I owned the troubleshooting of a network issue' demonstrates Ownership and Bias for Action.
Focus Topics
Amazon Leadership Principle: Frugality
Accomplishing more with less, eliminating waste, optimizing costs, and finding creative solutions within constraints. Example: 'I solved a network bottleneck by reconfiguring existing equipment rather than purchasing new hardware.'
Practice Interview
Study Questions
Amazon Leadership Principle: Earn Trust
Being honest and direct, following through on commitments, admitting mistakes, listening to others, and earning respect through integrity. Example: 'When I made a configuration error, I immediately disclosed it and helped fix it rather than trying to hide it.'
Practice Interview
Study Questions
Amazon Leadership Principle: Learn and Be Curious
Continuously seeking knowledge, learning new technologies, staying updated on industry trends, seeking feedback, and teaching others. Example: 'I took an online networking course to deepen my knowledge beyond what my job required.'
Practice Interview
Study Questions
Amazon Leadership Principle: Ownership
Taking responsibility for outcomes, going beyond job requirements, being accountable for mistakes, and driving solutions without waiting for permission. Example: 'I identified and fixed a network configuration issue rather than just reporting it.'
Practice Interview
Study Questions
Amazon Leadership Principle: Bias for Action
Making reasonable decisions quickly with available information, taking calculated risks, moving forward rather than over-analyzing, and iterating based on results. Example: 'I implemented a temporary fix immediately while planning a permanent solution.'
Practice Interview
Study Questions
Amazon Leadership Principle: Deliver Results
Committing to ambitious goals, prioritizing work to deliver on commitments, and measuring success by business impact. Example: 'I worked extra hours to ensure a network maintenance didn't disrupt business operations.'
Practice Interview
Study Questions
Frequently Asked Network Engineer Interview Questions
Describe how to deploy a voice VLAN for IP phones on access ports. Include how the phone and PC behind the phone typically tag traffic, how phones learn the voice VLAN (CDP/LLDP-MED/DHCP), QoS considerations (CoS/DSCP trust), and the access-port configuration best practices to prioritize voice.
Sample Answer
Direct answer
A voice VLAN gives IP phones their own VLAN (and its own subnet) on the same access port as the PC plugged into the phone. The port stays an access port for the PC's data VLAN, and a separate voice VLAN is attached with switchport voice vlan. The switch tells the phone which VLAN to use through CDP (Cisco Discovery Protocol); the phone then tags its own traffic with that VLAN and priority, while the PC behind it stays untagged. Tagging means inserting an 802.1Q label that names the VLAN into the Ethernet frame; an untagged frame carries no such label, and the switch assigns it to the port's access VLAN. QoS (quality of service) must be set so the switch trusts the phone's markings and not the PC's.
How traffic is tagged
- Phone: with
switchport voice vlan <id>, the phone sends voice in the voice VLAN with 802.1Q priority 5 (Cisco's description of the vlan-id option). - PC behind the phone: sends untagged frames, and the phone's built-in switch passes them through unchanged. Cisco states untagged data from devices attached to the phone passes through regardless of the port's trust state.
- Result: the single link carries tagged voice and untagged data, which behaves like a mini trunk (a trunk is a link that carries several VLANs, tagged) without being configured as a trunk. Cisco notes voice VLAN is only supported on access ports, not on trunk ports.
How phones learn the voice VLAN
| Method | How |
|---|---|
| CDP | Per Cisco's guide, the switch sends CDP packets that instruct the phone how to send voice traffic. CDP is enabled by default on interfaces, and it must be enabled on the port. |
| LLDP-MED (Link Layer Discovery Protocol for Media Endpoint Devices) | The network-policy TLV (type-length-value field) lets the switch advertise VLAN, CoS (class of service), DSCP (differentiated services code point) and tagging mode for voice to the phone. Whether a network-policy profile and switchport voice vlan can coexist on one interface, and in which order they must be applied, is platform and release specific, so check the platform's LLDP-MED documentation. |
| DHCP (Dynamic Host Configuration Protocol) | DHCP gives the phone its IP address, not its VLAN. CDP and LLDP-MED run at layer 2, so the phone gets the VLAN before it has any IP address, then sends its DHCP request tagged in that VLAN, which means the voice VLAN's subnet needs its own DHCP scope. Some phone vendors also deliver extra settings in DHCP options; that is vendor-specific and not part of Cisco's switch-side voice VLAN configuration, so confirm it in the phone vendor's documentation. |
The switchport voice vlan keywords per Cisco: vlan-id (voice tagged in that VLAN, priority 5), dot1p (voice tagged with VLAN ID 0, 802.1p priority), untagged (voice sent untagged), none (phone uses its own configuration).
Access-port configuration
vlan 40
name DATA
vlan 140
name VOICE
interface GigabitEthernet1/0/15
description IP phone with PC
switchport mode access
switchport access vlan 40
trust device cisco-phone
switchport voice vlan 140
switchport priority extend cos 0
spanning-tree portfast
trust device cisco-phonemakes the port trust QoS markings only when a Cisco phone is detected. Cisco lists it as a prerequisite before enabling voice VLAN, and its configuration steps apply it beforeswitchport voice vlan, which is the order shown above.switchport priority extend cos 0makes the phone rewrite the PC's CoS to 0 so a PC cannot claim voice priority. (switchport priority extend trustinstead trusts the PC's priority, which you would not want.)- PortFast makes the port skip the spanning-tree waiting states and forward immediately, so a phone can boot and get an address without a delay.
- Cisco states PortFast is enabled automatically when the voice VLAN is configured and is not removed when voice VLAN is removed.
- QoS syntax differs between switch platforms, so check the platform's QoS guide for the trust command.
QoS considerations: CoS and DSCP trust
Layer 2 priority is the 3-bit CoS (class of service) in the 802.1Q tag, a number from 0 to 7 where the phone sends voice at 5, as above. Layer 3 priority is the 6-bit DSCP (differentiated services code point) value in the IP header, which survives routing where the 802.1Q tag does not. Voice media is conventionally marked DSCP 46, called EF (expedited forwarding), which is 101110 in binary; RFC 4594 recommends EF for the telephony class and CS5 (40) for signaling. Its top three bits 101 equal 5, matching the IP precedence of 5 that Cisco's guide gives as the default for voice traffic (and 3 for voice control traffic). The trust boundary is the point in the network where you stop believing the priority marks that devices put on their own traffic and start enforcing your own. Trust boundary design: trust the phone (conditionally, via trust device cisco-phone), do not trust the PC, and keep end-to-end queuing for voice on every uplink so the priority marking is acted on.
Verification
show interfaces GigabitEthernet1/0/15 switchport reports the access VLAN and the voice VLAN as separate fields, so you should read 40 in the first and 140 in the second; if the voice field says none, the switchport voice vlan line is missing or was overridden. show running-config interface GigabitEthernet1/0/15 confirms the lines. Confirm the phone registers in the voice VLAN and the PC gets a data-VLAN address.
Trade-offs and pitfalls
- Keep the voice VLAN in the uplink trunk's allowed list.
- Do not put a switch or hub that you do not control behind the phone.
- Treat the voice VLAN as an enforcement point: filter with ACLs so PCs cannot reach voice servers they should not.
- Prefer
trust device cisco-phone(conditional) to unconditional trust, which lets a PC spoof priority.
You have four contiguous IPv4 networks: 10.0.0.0/24, 10.0.1.0/24, 10.0.2.0/24 and 10.0.3.0/24. What single prefix covers them, and how do you work it out bit by bit? Then explain when summarizing a set of subnets like this would not be possible or would be risky.
Sample Answer
Direct answer
10.0.0.0/22 covers all four networks exactly: it spans 10.0.0.0 to 10.0.3.255, which is the same 1,024 addresses as four /24s (4 x 256). Summarization is risky or impossible when the networks are not contiguous and aligned, or when the summary would claim addresses you do not own or route elsewhere.
Bit-by-bit derivation
Write the differing octet (the third) in binary for each network. The first two octets (10.0) are identical, which is 16 common bits.
| Network | Third octet |
|---|---|
| 10.0.0.0/24 | 00000000 |
| 10.0.1.0/24 | 00000001 |
| 10.0.2.0/24 | 00000010 |
| 10.0.3.0/24 | 00000011 |
The first six bits of the third octet (000000) are the same in all four; only the last two vary. Common bits: 16 + 6 = 22. Mask for /22 is 255.255.252.0 (252 = 11111100). Check: the block size is 256 - 252 = 4 in the third octet. The last two bits of that octet are free in a /22, so the first network must have those two bits at 00, which means its third octet is a multiple of 4. 10.0.0.0 is, so four /24s fit exactly.
When summarizing fails or is risky
- Not aligned. A set must start on a multiple of its block size and be a power-of-two count. 10.0.1.0/24 to 10.0.4.0/24 are also four consecutive networks, but their third octets in binary are 00000001, 00000010, 00000011 and 00000100, which already disagree at the sixth bit, so they share only 16 + 5 = 21 bits, and 10.0.1.0 is not a multiple of 4. They collapse to 10.0.1.0/24 + 10.0.2.0/23 + 10.0.4.0/24. One covering prefix would have to be 10.0.0.0/21 (10.0.0.0 to 10.0.7.255), eight /24s, four of which are not yours.
- Gaps. Three networks 10.0.0.0/24 to 10.0.2.0/24 summarize only as 10.0.0.0/23 plus 10.0.2.0/24. A single /22 would cover 10.0.3.0/24, which is not part of the set.
- Too coarse a summary. It advertises addresses you do not own or that live elsewhere. If 10.0.3.0/24 is at another site and that site also advertises its /24, longest-prefix match saves you. Longest-prefix match means a router forwards each packet using the matching route with the longest prefix, the most specific one: a packet for 10.0.3.5 matches both 10.0.0.0/22 and 10.0.3.0/24, and the /24 wins. If the other site does not advertise it, traffic for that /24 is pulled to you and dropped or looped (a blackhole, where packets silently disappear).
- Hidden failures. A summary stays up while one component subnet is down, so the rest of the network keeps sending traffic toward a router that cannot deliver it. For example, if 10.0.2.0/24 loses its link, the /22 is still advertised, so distant routers keep sending 10.0.2.x packets to you, where they are dropped, instead of failing fast or switching to a backup path. The specific /24 withdrawal is hidden inside the summary.
- Non-contiguous design. Subnets of the same block split across two sites cannot be summarized at either site.
Why aggregation is worth doing
One /22 route replaces four /24 routes, so routing tables and updates shrink and flaps (a route repeatedly going down and up) of one /24 stop propagating beyond the summarizing router. At large scale that cuts memory and CPU use and makes convergence (all routers agreeing on the new topology after a change) faster. The price is precision, which is why summaries should be planned in the addressing scheme: give each site or region an aligned block and summarize at its edge.
Practice
The summarizing router needs a discard (null) route for the summary so traffic to unused parts of the block is dropped there rather than looped back out a default route. Some protocols install it for you, as described at the end of this section; when you originate the summary yourself, for example with a static route that is then redistributed, you add it by hand. Without it, a packet for an unused address such as 10.0.3.77 reaches the router because of the advertised summary, finds no specific route, follows the default route back toward the upstream router that sent it, and bounces until its TTL (hop limit) expires. With the discard route, the /22 matches (it is more specific than the default route, which matches everything) and the packet is dropped at once. On Cisco IOS this is ip route 10.0.0.0 255.255.252.0 Null0, where Null0 is a virtual interface that discards everything sent to it. The advertisement itself is configured separately: in OSPF, an area range command on the router at the area edge; in BGP, an aggregate-address command. Both of those install the discard route automatically: Cisco documents that an OSPF ABR (area border router) or ASBR installs a discard route to Null0 for a summary, with a discard-route setting that can turn this off, and that BGP aggregate-address inserts a route toward Null0 into the routing table to prevent forwarding loops. So the manual Null0 line is for summaries you create by static configuration or by other means. Do not disable the automatic discard route unless you have another route that makes unused space in the summary drop safely.
Tell me about a time internal or external pressure, such as a deadline, a client, or a business commitment, pushed you toward a decision that conflicted with a principle or value your company had explicitly committed to (for example privacy, security, or data quality). Walk through how you recognized the conflict, what you did about it, how you communicated your position to stakeholders, and what the final outcome was.
Sample Answer
Direct answer
When a deadline, a client, or a business ask pushes toward something that conflicts with a principle a company has committed to, such as privacy, security, or data quality, the strongest answers show three things: you noticed the conflict explicitly rather than complying without registering it, you raised it through the right channel rather than either silently complying or unilaterally blocking the work, and you drove toward a resolution rather than just splitting the difference.
Structured elaboration
- Notice: name the specific moment you recognized the tension, and what concrete detail made you pause.
- Raise it: describe how you raised it, ideally backed by data or a concrete risk rather than an appeal to principle alone. A values-based objection lands far better when it is backed by the actual risk it protects against.
- Navigate: what you actually did in the interim, whether you proposed a compromise or a phased approach, who you looped in, and how you kept the relationship functional even while disagreeing.
- Outcome: what actually happened. An honest outcome, including "I was overruled and here is what I did next," is often more credible than a suspiciously clean win.
Worked example
A team was under pressure to ship a change quickly, and the fastest path meant skipping a validation step that existed specifically to catch a known class of data-quality problem. Rather than quietly skipping it or unilaterally blocking the release, the response was to time-box a reduced version of the validation, checking the highest-risk subset in the time available, and to flag explicitly and in writing what wasn't covered and what the residual risk was, so the decision to accept that risk was made deliberately by the right people rather than by default. The release shipped on time, and the flagged gap was closed within the following two days as agreed, rather than being silently forgotten.
Trade-offs and pitfalls
A story where you unilaterally blocked the work and were later vindicated can read as inflexible if it doesn't also show you understood the business pressure; the strongest answers show empathy for that pressure while still holding the line. A story where you quietly went along with the shortcut is not really an example of this competency at all; the action needs to show you actively surfaced the tension, not merely noticed it internally. Vague appeals to "our values" without a concrete risk attached tend to land weaker than a specific technical or business risk, clearly stated.
You're the on-call network engineer and must choose between applying broad traffic shaping that will intentionally degrade a lower-priority business function to stabilize the network, or allowing congestion to continue risking larger customer impact. Describe your decision-making process: how you assess business impact, who you involve, required approvals, communication to stakeholders/customers, rollback plan, and documentation for post-incident review.
Sample Answer
Direct answer
Choosing between deliberately degrading a lower-priority business function and letting congestion continue unmanaged is a business-impact decision dressed as a technical one, and treating it that way, gathering the SAME evidence and approvals you would for any consequential business trade-off, rather than making the call alone as a purely technical judgment, is what actually protects both the business and yourself.
Structured elaboration
- Assess business impact concretely, not just technically: quantify, as precisely as time allows, what "lower-priority business function" actually means in revenue, customer, or SLA terms, and compare that against the projected impact of continued, unmanaged congestion (which customers, which SLAs, how much revenue) so the decision is grounded in comparable, concrete terms rather than a vague technical instinct about which is "worse."
- Identify who actually needs to be involved in this decision: this is not purely a network-engineering call, since it deliberately trades one business function's degradation for another's protection; the OWNERS of the function being degraded need to be consulted or at minimum informed before the decision is executed, not after, given they're the ones who can best assess whether that specific trade-off is acceptable from their side.
- Determine what approvals are actually required, and get them BEFORE acting where the timeline allows: for a decision with this level of business consequence, a single network engineer unilaterally choosing which business function to sacrifice, without any sign-off from someone with the authority and context to weigh that trade-off, is a governance gap; if the situation is severe enough that waiting for approval risks worse outcomes, have a PRE-AGREED escalation path for exactly this kind of emergency decision, rather than improvising the approval process live for the first time during the incident.
- Communicate proactively to affected stakeholders and customers, not just internally, framing what's being deprioritized and why, ideally BEFORE they notice the degradation themselves rather than only reactively once they report it.
- Build in an explicit rollback plan and trigger condition: define in advance what condition would cause you to REVERSE the traffic-shaping decision (the primary congestion resolving, or the "lower-priority" function's degradation turning out to be worse than expected), rather than treating the shaping decision as a one-way, unmonitored action.
- Document for post-incident review: capture the reasoning, who was consulted, what alternatives were considered, and what the actual measured impact was on both sides of the trade-off, since a decision this consequential deserves a thorough after-the-fact review regardless of how it turns out.
Worked example
Applying traffic shaping specifically to a batch-reporting feature (assessed as lower-priority, with a known, currently-acceptable SLA for delay) to protect the primary customer-facing transactional service from congestion-driven failures: before applying, I'd confirm the batch-reporting feature's OWNER agrees this specific SLA hit is acceptable given the alternative, get a quick sign-off from an on-call incident commander or equivalent authority given the time pressure, notify affected stakeholders proactively with a specific expected duration, and set an explicit trigger (once transactional congestion drops below a defined threshold) to lift the shaping.
Trade-offs & pitfalls
Making this kind of decision unilaterally, purely as a technical judgment call, without involving the people who actually own the function being sacrificed, is a common failure mode under time pressure; even a brief, fast consultation is usually better than none, and a pre-agreed emergency-escalation path removes the need to improvise that process for the first time during a live incident. Skipping proactive communication in favor of just acting and explaining later, if at all, damages trust with the very stakeholders whose function you just deprioritized on their behalf.
Explain how NAT, hairpinning (NAT loopback), and service discovery interact when traffic traverses multiple firewalls or VPCs. Walk through how you would diagnose a resulting failure (for example, a callback or health check that stops working) and propose concrete mitigations and troubleshooting steps.
Sample Answer
Network Address Translation (NAT), hairpinning (also called NAT loopback), and service discovery can interact in a way that works perfectly from outside the network but silently fails for a caller inside it, because the caller and an external client can end up taking genuinely different paths to reach what looks like the same service.
How the interaction breaks things
Hairpinning is what has to happen when a client behind the same NAT gateway as a service tries to reach that service via its public address: traffic has to leave the gateway and come back in through it, rather than staying purely internal, and not every NAT or firewall device supports this correctly. If service discovery or DNS resolves callers to the service's externally-published address regardless of where the caller sits (a very common default when there's a single public DNS name), an internal caller ends up depending on hairpin support it may not actually have.
Worked example: a health check that stops working
An internal health-check process resolves the target service's public DNS name (because that's the only name service discovery publishes) and connects to its public IP through the same firewall/NAT path an external client would use. If that gateway doesn't support hairpin NAT, the SYN either gets silently dropped or the return traffic takes an asymmetric path back through a different device that never saw the outbound leg and drops it as an invalid, untracked connection. External clients keep working fine (they were never depending on hairpinning), so the failure looks host- or path-specific and is easy to misdiagnose as an application bug rather than a network topology issue.
Diagnosing the failure
- Capture from the internal caller's own interface to see whether the SYN leaves at all, and whether any SYN-ACK returns, and if it does, whether it arrives via the interface you'd expect.
- Check exactly what address service discovery handed the caller: public or internal.
- Check whether the outbound and expected-return paths traverse the same stateful device; a stateful firewall that sees only one leg of a connection will drop the other leg as invalid, which looks identical to a routing problem from the outside.
- Confirm whether the specific NAT/firewall device in the path actually supports hairpin/loopback NAT, since support varies significantly by vendor and platform.
Mitigations
Use split-horizon DNS (or equivalent internal service discovery) so internal callers resolve to an internal address or private virtual IP entirely, bypassing the external NAT and load-balancer path rather than depending on hairpinning to work at all. Where hairpinning genuinely must be relied on, explicitly verify and enable loopback support on the specific gateway rather than assuming it. Ensure routing symmetry, the same stateful device sees both directions of a given flow, since asymmetric routing across multiple firewalls is a common root cause even when hairpinning itself is supported. Finally, test service-to-service calls from inside the network as part of normal validation, rather than only testing from outside, since that's exactly the blind spot this failure mode hides in.
For a B2B SaaS provider needing low-latency private connectivity into customer AWS accounts, compare exposing services via AWS PrivateLink (VPC Endpoint Services) versus exposing internet-facing ALBs protected by WAF and CDN. Discuss security posture, performance/latency, operational complexity for onboarding customers, cost, and which approach you'd recommend for private B2B integrations.
Sample Answer
Direct answer
For a B2B (business-to-business) SaaS provider whose customers are enterprises with their own AWS accounts, AWS PrivateLink is the right default for the core integration path: it keeps traffic off the public internet entirely and gives each customer a private, predictable connection into your service, at the cost of a heavier one-time setup per customer. Keep an internet-facing Application Load Balancer (ALB) behind a Web Application Firewall (WAF) and a content delivery network (CDN) as the path for anything that doesn't need that guarantee (public marketing APIs, self-serve trial signups, webhooks). Route customers to PrivateLink specifically when they demand it and to the public path otherwise.
Structured elaboration
PrivateLink works by having you (the provider) put a Network Load Balancer (NLB) in front of your service and register it as a VPC Endpoint Service. Each customer creates an Interface VPC Endpoint (an ENI, or elastic network interface, placed directly in their own VPC) that resolves to a private IP address inside their network and forwards traffic to your NLB over AWS's internal backbone. It never touches the public internet, and you can require explicit acceptance of each customer's endpoint connection request (an allowlist) rather than auto-accepting everyone.
| Dimension | PrivateLink (Endpoint Service) | Internet-facing ALB + WAF + CDN |
|---|---|---|
| Security posture | No public exposure of the service at all; attack surface is limited to whoever you accept as an endpoint consumer. Still needs security groups and IAM (Identity and Access Management, the system controlling who can do what) on your side, but removes an entire class of internet-facing risk. | Public DNS name is inherently probe-able. WAF blocks common web exploits and can rate-limit, but the origin is discoverable and reachable by anyone unless you layer IP allowlisting or mutual TLS (mTLS, where both sides present certificates) on top. |
| Performance and latency | Single hop over the AWS backbone, low and consistent variance since it never leaves AWS's network; the main added latency is the NLB itself, which is minimal. | CDN edge points of presence terminate TLS close to the customer, which can look fast for the initial connection, but B2B API traffic is rarely cacheable, so the data plane still crosses the public internet for the actual request, with more variable latency. |
| Operational complexity for onboarding | Higher per customer: the customer must create the interface endpoint, accept a private DNS name (which requires proving domain ownership via a TXT record if you want name-based routing instead of the AWS-generated endpoint DNS name), and you must accept the connection request. This is a real integration task, not a config flag. | Trivial: hand the customer a URL and an API key. No infrastructure change on their side. |
| Cost | You pay a per-AZ (Availability Zone) endpoint-service hourly charge plus a per-GB data-processing charge, multiplied by however many customers connect; this scales roughly linearly with customer count regardless of traffic volume. | Cost is centralized in the ALB, WAF rule evaluations, and CDN request/data-transfer pricing, and scales with request volume rather than customer count, which is cheaper for a large number of low-traffic customers. |
Worked example
A SaaS provider with 40 enterprise customers, each sending a moderate, steady volume of API calls, decides on PrivateLink for the primary integration and keeps the ALB/WAF/CDN path only for the public signup and marketing site. Rough reasoning, not a live pricing quote: 40 customers each maintaining one interface endpoint across two AZs is 80 endpoint-hours per hour of runtime, small in absolute dollars compared to the value of the deals it unblocks, because for this buyer segment (regulated enterprises procuring a vendor integration), "our data never leaves our VPC" is frequently a hard security-review requirement, not a nice-to-have. If this same provider instead had 5,000 small business customers each sending a handful of API calls a day, the same math flips: 5,000 endpoints, most nearly idle, is now a real fixed cost with little security benefit for a segment that doesn't have a security-review gate in the first place, and the ALB/WAF/CDN path plus mTLS to a handful of larger accounts on request is the better default.
Trade-offs and pitfalls
The most common mistake is treating this as a binary product decision instead of a per-customer routing decision: offer PrivateLink as a tier-gated option, not the only path, or the onboarding overhead becomes a tax on every customer regardless of whether they need it. A second pitfall is assuming PrivateLink means "no security work required" on your side; you still need to scope the endpoint service's allowed-principal list, apply security groups to the NLB's targets, and monitor connection acceptance, since PrivateLink narrows the attack surface but doesn't eliminate the need for defense in depth. What would flip the recommendation: if the customer base shifts toward a long tail of small, low-touch accounts rather than a smaller number of large enterprise accounts, or if per-customer support cost for PrivateLink onboarding starts exceeding the deal value it protects.
Design an observability dashboard for on-call engineers that helps detect imminent capacity exhaustion before SLOs are violated. List the panels, derived metrics, short and long term alerts, and escalation rules you would include. Explain how you would set thresholds to minimize alert fatigue while ensuring timely action.
Sample Answer
Panels, ordered by how directly they connect to user impact:
- SLO (service level objective, the internal reliability target)/error-budget burn panel: current burn rate against budget remaining, the single "are we about to break our promise" panel, placed at the top.
- Saturation panels per tier (CPU/GPU, memory, connection pool, queue depth), each showing headroom to limit, not just the current value.
- A time-to-exhaustion panel: a derived metric projecting, from a short-window trend, how many minutes or hours remain before a saturation metric crosses its limit at the current rate of change.
- Autoscaler activity: recent scale events plus current versus max replica count, so on-call can see at a glance whether the automatic system is already responding or already maxed out.
- Per-environment differences for a hybrid cloud/on-prem estate: on-prem panels show physical link and rack-level power/cooling headroom, a ceiling that simply doesn't exist in the cloud, while cloud panels show account-level quota utilization, since a cloud capacity ceiling is often an API rate limit or instance quota, not a hardware one. A combined "worst headroom across all providers" tile sits at the top so on-call doesn't have to hunt for which environment is actually the constraint.
Derived metrics
- Headroom percentage, (limit minus current) divided by limit, shown instead of a raw value, since 72% utilization means something very different on a 1 Gbps link versus a 100 Gbps link.
- Trend slope over a rolling window, feeding the time-to-exhaustion projection.
- Burn-rate multiplier for the SLO panel: how many times faster than sustainable the resource is being consumed.
Short-term vs long-term alerts and escalation
Short-term: a fast-approach alert, for example time-to-exhaustion under 30 minutes at the current trend, pages on-call immediately. Long-term: a slow-trend alert, growth that would exhaust capacity within 3-7 days, files a ticket for the next capacity-planning review rather than paging, since there's time to act deliberately. A short-term alert unacknowledged within 10 minutes escalates to secondary on-call, then to the team lead. A fast-burn alert that coincides with an already-open incident groups under that incident instead of paging a second time for the same root cause.
Minimizing alert fatigue while staying timely
Distinguish a genuine sustained trend from a single spike before paging: require the time-to-exhaustion projection to still hold on a second consecutive evaluation, five minutes apart, rather than firing on one noisy sample. A spike that resolves within one window never pages at all. Once a fast-burn alert has fired and paged, suppress duplicate pages for the same resource until it either resolves or a fixed re-notify interval passes, for example 30 minutes, so a sustained breach doesn't page every single evaluation cycle. Group correlated alerts, like a saturation alert, a queue-depth alert, and an autoscaler-at-max alert all firing together on the same service, into a single incident notification instead of three separate pages, since they share one root cause.
Describe how you would run capacity planning for an upcoming period across a group with varying availability (vacations, hiring timelines, growing demand). Explain how you'd convert available capacity into a realistic delivery forecast, what buffers you'd build in, and how you'd get budget or headcount approved and executed without disrupting ongoing work.
Sample Answer
Direct answer
Converting uneven team capacity into a believable delivery forecast means starting from real, not nominal, person-time (accounting for vacations, ramp-up, and hiring timelines), applying a deliberate buffer for the work that always eats capacity but never appears on a roadmap, and taking the resulting gap, not a guess, to leadership as the basis for a headcount or budget ask, before the plan is committed publicly.
Structured elaboration
- Start from nominal capacity, then subtract what's real. Multiply headcount by the planning period to get a nominal ceiling, then subtract planned time off and add back new hires only at their actual ramp rate, since a new hire's first weeks are rarely full productive capacity.
- Apply a buffer. Reserve a deliberate percentage of the remaining time for on-call, meetings, and unplanned work, since a forecast built on 100 percent utilization is fiction and will slip on contact with the first unplanned incident.
- Compare against demand. Convert the committed backlog into the same person-time unit as capacity, and be explicit about the resulting gap or surplus, rather than quietly hoping it works out.
- Get budget or headcount approved without disrupting ongoing work. Bring the specific gap, in the same units used for the forecast, to leadership before committing to the plan, with named options to close it (added headcount, a contractor, or an explicit scope cut), so the ask is a concrete trade-off decision rather than an open-ended request.
Worked example
A team of 5 engineers is planning a 12-week quarter, a nominal 60 person-weeks of capacity.
- One planned parental leave removes 6 person-weeks: 60 minus 6 equals 54.
- A new hire joins in week 7, ramping at 50 percent for their first 4 weeks (2 effective person-weeks) and full speed for the final 2 weeks of the quarter (2 effective person-weeks), adding 4 effective person-weeks rather than the 6 nominal weeks a fully ramped person would contribute in that span: 54 plus 4 equals 58 raw available person-weeks.
- Applying a 20 percent buffer for on-call, meetings, and unplanned work: 58 times 0.8 equals about 46 committed person-weeks.
- The backlog the team has been asked to deliver this quarter is estimated at 55 person-weeks.
- Gap: 55 minus 46 equals 9 person-weeks short, about 16 percent of demand.
Taken to leadership before the quarter's plan is finalized: either fund a contractor for roughly 9 person-weeks of the quarter, or agree explicitly on which 9 person-weeks of backlog move to next quarter. Either way, the decision is made and documented before the roadmap is communicated externally, not discovered mid-quarter.
The same convert-to-forecast-plus-buffer discipline applies when the constrained resource is infrastructure rather than headcount. For a network capacity plan forecasting 40 percent annual traffic growth, the same shape holds: convert the growth forecast into a specific utilization threshold (for example, an upgrade trigger at 70 percent link utilization on a given segment), stage the capital request into phased upgrades tied to when each segment is projected to cross that threshold instead of one lump-sum ask, and get each phase approved on its own timeline so operations are never disrupted by an all-at-once cutover.
Trade-offs and pitfalls
The most common mistake is planning off nominal headcount times weeks, without subtracting real ramp-up and time-off effects, which produces a forecast that looks fine on a spreadsheet and fails the first month. A second mistake is presenting the gap as a vague concern rather than a specific number with named options, which makes it easy for leadership to defer a decision until the shortfall is already a missed commitment. Building in too large a buffer is its own failure mode: an overly conservative estimate under-promises capacity the team actually has and can erode trust in future forecasts just as much as an overly optimistic one.
What is the difference between traffic shaping and policing, and how do common queuing approaches decide which packets go first? Where in an enterprise would you apply markings?
Sample Answer
Direct answer
Both policing and shaping hold traffic to a configured rate using a token bucket (a counter that refills at the allowed rate and is spent as packets pass). They differ in what happens to a packet that arrives with no tokens: a policer drops it (or re-marks it), a shaper queues it and sends it later. Queuing is the separate decision of which waiting packet leaves the interface next. Markings (DSCP, Differentiated Services Code Point, a 6-bit value in the IP header) are written once at the edge you trust and read by every queue downstream. On the public internet they are often reset to zero, so they matter only inside networks you or your carrier control.
Policing versus shaping
Cisco's documentation puts it plainly: a policer typically drops traffic, though it can instead change a packet's marking, and a shaper typically delays excess traffic in a buffer. Shaping applies to traffic leaving an interface, while policing can be applied in either direction. The terms in the token bucket:
- CIR (committed information rate): the average rate allowed.
- Bc (committed burst): how many bytes can pass at once, which is the bucket depth.
- Tc (time interval): mean rate = burst size / time interval.
With numbers: CIR = 10 Mbps and Bc = 15,000 bytes (120,000 bits) give Tc = Bc / CIR = 120,000 / 10,000,000 = 0.012 s = 12 ms. The bucket refills completely every 12 ms, and 120,000 bits per 12 ms is the 10 Mbps mean rate. The 15,000 bytes are the same 10 packets of 1,500 bytes used in the example below.
| Policing | Shaping | |
|---|---|---|
| Excess traffic | Dropped or re-marked | Queued, sent later |
| Latency added | None | Up to the queue length |
| Effect on TCP | Loss, so the sender backs off with a saw-tooth (TCP's rate climbs until a packet is lost, then halves, so a plot of its rate looks like saw teeth) | Smoother, with longer round trips |
| Where | Edge, in or out (customer rate limit, control-plane protection) | Egress, to match a slower link or a contracted rate |
Use shaping when you must stay under a carrier's committed rate without losing packets, and policing when you must enforce a limit on traffic you do not want to buffer, such as untrusted or bulk traffic.
Worked example: 20 Mbps offered to a 10 Mbps limit
167 packets of 1,500 bytes arrive 0.6 ms apart (20 Mbps for about 100 ms). The policer has a 15,000-byte bucket (10 packets); the shaper has a 32-packet queue and drains at 10 Mbps (1.2 ms per packet).
PKT = 1500 # bytes per packet
BITS = PKT * 8
RATE = 10e6 # 10 Mbps policer and shaper rate
BURST = 10 * PKT # bucket depth: 15,000 bytes
arrivals = [i * BITS / 20e6 for i in range(167)] # 20 Mbps offered for about 100 ms
# Policer: tokens refill at RATE, a packet that finds too few tokens is dropped
tokens, last, passed, dropped = BURST, 0.0, 0, 0
first_drop = None
for i, t in enumerate(arrivals):
tokens = min(BURST, tokens + (t - last) * RATE / 8)
last = t
if tokens >= PKT:
tokens -= PKT
passed += 1
else:
dropped += 1
first_drop = i if first_drop is None else first_drop
print(f"policer: {passed} passed, {dropped} dropped (first drop at packet {first_drop}), no packet delayed")
# Shaper: packets wait in a 32-packet queue and leave at exactly RATE
QUEUE = 32
serial = BITS / RATE # 1.2 ms per packet
free_at, queued_until, sent, dropped, worst = 0.0, [], 0, 0, 0.0
for t in arrivals:
queued_until = [d for d in queued_until if d > t]
if len(queued_until) >= QUEUE:
dropped += 1
continue
depart = max(t, free_at) + serial
free_at = depart
queued_until.append(depart)
sent += 1
worst = max(worst, depart - t - serial)
print(f"shaper: {sent} sent, {dropped} dropped, worst queueing delay {worst * 1000:.1f} ms, last packet leaves at {free_at * 1000:.1f} ms")
Output:
policer: 92 passed, 75 dropped (first drop at packet 19), no packet delayed
shaper: 114 sent, 53 dropped, worst queueing delay 37.2 ms, last packet leaves at 136.8 ms
The policer passes the 10-packet burst plus what the refill allows (92 packets, about 11 Mbps over this window) and drops 75, with no added delay. The shaper delays instead: the queue builds until the worst packet waits behind 31 others, 37.2 ms of delay for the worst packet, and it still drops 53 once the 32-packet queue fills, because the offered load stays at twice the rate for the whole window. A shaper absorbs bursts shorter than its queue and drops when overload outlasts it.
Tracing the first packets by hand. Packets arrive every 0.6 ms (1,500 x 8 bits / 20 Mbps), and the policer's bucket refills 10,000,000 / 8 x 0.0006 = 750 bytes in each gap. Packet 0 finds 15,000 bytes and spends 1,500, leaving 13,500. Packet 1 gets 750 back and spends 1,500, leaving 12,750. Each packet nets minus 750, so the bucket empties after packet 18. Packet 19 finds only 750 bytes and is dropped; packet 20 finds 750 + 750 = 1,500 and passes; from there the policer alternates pass and drop, which is half of 20 Mbps, the 10 Mbps limit. The shaper instead holds packets: packet 0 leaves at 1.2 ms; packet 1 arrives at 0.6 ms, finds the line busy until 1.2 ms, leaves at 2.4 ms and so waited 0.6 ms; every later packet waits 0.6 ms longer than the one before, so packet 62 waits 37.2 ms, and packet 63 finds all 32 queue slots full and is dropped.
How queues decide which packet goes first
Each row below fixes the weakness of the row above it.
| Method | Rule | Weakness |
|---|---|---|
| FIFO (first in, first out) | Arrival order | A bulk transfer delays voice |
| Strict priority | Serve the high queue until empty | Can starve everything below if unlimited |
| Weighted fair queuing (WFQ, class-based as CBWFQ) | Each class gets a share by weight | No strict latency guarantee |
| Low-latency queuing (LLQ) | Strict priority queue that is capped, plus CBWFQ for the rest | Needs the priority class sized correctly |
Weights in practice: classes weighted 50, 30 and 20 on a congested 10 Mbps link get 5, 3 and 2 Mbps. If the 20 class goes idle, the other two share its capacity in proportion and get 6.25 and 3.75 Mbps. LLQ is the usual enterprise answer: voice in the capped priority queue, other classes by weight. A queue also needs a drop policy for when it fills; plain tail drop (a full queue discards the packet that just arrived) discards new arrivals, while weighted random early detection (WRED) starts dropping packets at random before the queue is full, and does so earlier for lower-priority traffic.
Where to apply markings
First, classification (deciding which class a packet belongs to, by port, address, application signature or existing mark) comes before marking (writing the value). Mark as close to the source as you trust the device, then let every later hop act on the mark.
| Place | What to do |
|---|---|
| Access switch port | The trust boundary (the first device whose arriving marks you stop believing and rewrite yourself). Trust the mark from known devices such as desk phones; re-mark everything from PCs and servers |
| Wireless | Map between the Wi-Fi priority and DSCP (Wi-Fi frames carry a User Priority from 0 to 7 that selects one of four access categories: voice, video, best effort and background; RFC 8325 maps EF to User Priority 6, the voice access category), and do not pass through marks from unauthenticated devices |
| Data center edge | Mark by source and destination, not by what the application sets |
| WAN edge | Re-mark to the carrier's agreed class menu; re-classify inbound traffic |
| Layer 2 and MPLS | A VLAN tag carries a 3-bit priority (CoS, class of service), and MPLS carries a 3-bit Traffic Class, so only 8 classes survive there; by default the MPLS value is the top 3 bits of the DSCP |
The same value appears in different notations: EF is 46 in decimal, 101110 in binary, and 0xb8 in the full type-of-service byte (46 x 4 = 184 = 0xb8), because the DSCP occupies the top 6 bits of that byte.
Why markings are often bleached or ignored on the public internet
The public internet has no end-to-end service contract. RFC 8100 notes that many networks re-mark unknown or unexpected DSCPs to zero when traffic enters, so a mark that is valid in your network can arrive as best effort. One reason is that a carrier cannot let customers who mark everything as high priority claim its premium classes. Only agreed interconnections preserve marks, and even then only for the agreed classes. The practical conclusion is that a mark guarantees nothing across the internet: your benefit comes from the queues at your own egress, where the mark is honored.
Pitfalls
- Shaping above the real link rate, so the queue forms in the carrier's device instead of yours.
- Marking from untrusted hosts, which lets any application take priority.
- An unbounded priority queue that starves other classes.
- Policing TCP bulk traffic too tightly: the resulting loss collapses throughput.
How would you measure latency, jitter, packet loss and throughput between two points in a cloud environment? Compare the tools you would use at different layers and what each can miss.
Sample Answer
Direct answer
Measure each property with the probe that sees it, from both ends, over enough time to be meaningful, and keep passive counters running next to the active tests. Latency is the round-trip time (RTT) from ping or a timed HTTP request. Jitter (variation in delay between packets) and packet loss come from a UDP test stream or a long run of probes. Throughput comes from a TCP test with iperf3 (a free, widely used throughput tester). No single tool is complete: each one sees one layer and misses something the next layer would show.
What each tool measures and what it misses
| Property | Tool and layer | What it can miss |
|---|---|---|
| Latency | ping (ICMP echo, layer 3, the network layer where IP addresses and routing live) | ICMP can be rate limited, deprioritized or blocked by security groups (the cloud's virtual firewall rules), and under equal-cost multipath (ECMP, where routers spread flows across parallel paths by hashing the 5-tuple of addresses, protocol and ports) a ping may take a different path than your application's flows |
| Latency | curl -w timings (TCP handshake plus first byte, layer 7, the application layer where HTTP lives) | Mixes network and server time: a slow application looks like a slow network |
| Jitter | iperf3 -u (UDP) receiver jitter, ping mdev (despite the name, the standard deviation of the RTTs, so a measure of their spread) | Averages hide bursts. Two formal definitions exist, with a numeric example under the table. RFC 3550 defines a smoothed estimator, J = J + ( |
| Loss | ping -c, iperf3 -u lost/total datagrams, mtr per hop | Few probes give a noisy figure (worked example below). mtr documents that routers may give ICMP echo lower priority, so loss shown at a middle hop that does not continue to the last hop is usually rate limiting, not real loss |
| Throughput | iperf3 TCP, then -P for parallel streams and -R for the reverse direction | One flow is bounded by window divided by RTT, so a single-stream test on a long path under-reports what many flows achieve. Providers also cap bandwidth per instance and per flow: AWS documents single-flow traffic limited to 5 Gbps unless instances share a cluster placement group (an option that packs instances physically close together; 10 Gbps), and burst credits (extra bandwidth smaller instance types may use only while a saved-up allowance lasts) that run out |
| Passive (all four) | ss -ti on the host (socket statistics: per-TCP-connection RTT, retransmit count and window size), provider counters such as AWS Elastic Network Adapter (ENA, the virtual network card) ethtool -S eth0 fields bw_in_allowance_exceeded, bw_out_allowance_exceeded, pps_allowance_exceeded | Passive data covers only real traffic that happened. Virtual-network drops caused by instance allowances are invisible to a ping inside the guest but show in those counters |
Worked jitter example. Four packets sent at 20 ms spacing arrive with one-way transit times of 40, 46, 41 and 45 ms. D, the change in transit between consecutive packets, is 6, 5 and 4 ms. The RFC 3550 estimator starts at J = 0 and goes J = 0 + (6 - 0)/16 = 0.375, then 0.375 + (5 - 0.375)/16 = 0.664, then 0.664 + (4 - 0.664)/16 = 0.873 ms. A de-jitter buffer (the receiver's short wait that smooths out uneven arrival for voice or video) needs the RFC 5481 view instead: delay minus the minimum delay gives 0, 6, 1 and 5 ms, so a buffer of about 6 ms absorbs all four packets.
Method rules that apply in any cloud: test both directions (paths and limits can differ), run from the same zone and across zones/regions so you can separate a local fault from a path fault, run long enough to include a busy period, and repeat probes on a schedule so you have a baseline rather than one number. One-way delay needs synchronized clocks on both ends, so use RTT unless you control time sync.
Worked example you can reproduce
This runs in any Linux container started with the NET_ADMIN capability (the permission to change network settings; iperf3, iproute2, iputils-ping and curl installed). tc netem (the Linux network emulator) delays packets on the loopback interface by 20 ms plus or minus 5 ms and drops 1 percent. Because every packet crosses loopback in each direction, expect an RTT near 40 ms and about 1.99 percent round-trip ping loss (1 - 0.99 x 0.99).
ip link set dev lo mtu 1500
tc qdisc add dev lo root netem delay 20ms 5ms loss 1%
iperf3 -s -D
python3 -m http.server 8080 --bind 127.0.0.1 &
ping -c 200 -i 0.2 127.0.0.1 | tail -3
iperf3 -c 127.0.0.1 -u -b 10M -t 20 | grep receiver
curl -s -o /dev/null -w 'connect=%{time_connect}s ttfb=%{time_starttransfer}s total=%{time_total}s\n' http://127.0.0.1:8080/
Output from one run (netem's randomness is not seeded, so your numbers will differ slightly):
200 packets transmitted, 194 received, 3% packet loss, time 40517ms
rtt min/avg/max/mdev = 33.127/44.944/65.456/7.373 ms
[ 5] 0.00-20.04 sec 23.6 MBytes 9.89 Mbits/sec 4.293 ms 150/17267 (0.87%) receiver
connect=0.043648s ttfb=0.093469s total=0.093511s
How to read it. The receiver's Lost/Total Datagrams column reads 150/17267: of the 17,267 datagrams it accounted for, 150 never arrived, 150 / 17,267 = 0.87 percent one way, with 4.293 ms of jitter. That is close to the 1 percent configured: 1 percent of 17,267 would be 173 lost, and the natural spread for that many packets is about 13 (sqrt(17,267 x 0.01 x 0.99)), so 150 is within two standard deviations of 173. The grep receiver keeps only the receiver's summary line, because the receiver is the end that counts what arrived. The ping shows 3 percent loss against an expected 1.99 percent: with only 200 probes the standard error is sqrt(0.0199 x 0.9801 / 200), about 0.99 percentage points, so 3 percent is one standard error high and is not evidence of a second problem. That is why a loss figure needs thousands of probes or a long window before you act on it. The curl connect time of 43.6 ms is one RTT (the TCP handshake needs a round trip), agreeing with the ping average of 44.9 ms, and the 93 ms first byte adds a second round trip for the request.
Trade-offs and pitfalls
- Reporting averages. Report percentiles (95th, 99th) for latency and the maximum for loss bursts, since users feel the tail.
- Active tests consume bandwidth and, across regions, can incur data-transfer charges. Run heavy iperf3 tests off-peak and short, keep light probes continuous.
- A clean ICMP result does not clear the application: test with the same protocol and port the application uses (
mtr -T -P 443sends TCP SYN probes to a chosen port). - Treat a throughput number as a ceiling for that test, not a promise: a TCP result tells you about the host's windows and the path together.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Network Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs