Network Monitoring and Performance Questions
Network telemetry and performance operations: SNMP polling and traps (including counter wraparound and SNMPv3 access), NetFlow, sFlow and IPFIX flow export, sampling and its accuracy, streaming telemetry (gNMI), and eBPF or packet-capture telemetry; interface-level metrics (utilization, errors, discards, queue depth, microbursts), active synthetic probing alongside passive counters, link-flap detection, baselining and anomaly detection on network signals including elephant-flow spotting, network SLIs and SLOs, alerting, alert-storm suppression and NOC dashboards, telemetry pipeline design, storage, retention, downsampling and cardinality for network data (including securing the collection path and handling bursty remote sites), BGP and link-state monitoring including prefix hijack and route-leak detection, and network capacity monitoring, percentile utilization and bandwidth headroom planning. Also covers measuring and tuning network-level latency, jitter, packet loss and throughput (bufferbloat, queueing, TCP tuning for long paths). Excludes the generic metrics, logs and traces stack and alert design, the layered fault-isolation method and packet-capture troubleshooting, TCP and protocol fundamentals, application and CDN latency engineering, cloud VPC design and security detection.
How does SNMP work for network monitoring? Walk through how a poller gets interface and device health data, how polling differs from traps, and what changes between the older and newer protocol versions.
Sample Answer
Direct answer
SNMP (Simple Network Management Protocol) is a request/response protocol over UDP. A poller (the "manager") periodically asks each device's SNMP agent for the current value of named variables, such as an interface's byte counter or operational state, and stores the answers. Traps and informs are the reverse direction: the device pushes an unsolicited event message when something happens. The three protocol versions differ mainly in counter size, bulk retrieval and, above all, security: v1 and v2c authenticate with a plain-text community string, while v3 adds real authentication, encryption and per-user access control.
How a poller gets interface and device health data
- The data model. Every readable variable is an object in a MIB (Management Information Base, a schema file), identified by an OID (object identifier, a dotted number such as
1.3.6.1.2.1.2.2.1.10). Interface data lives in the IF-MIB (RFC 2863): one row per interface, indexed byifIndex. Device health (CPU, temperature, power) usually comes from vendor MIBs or the standard ENTITY-SENSOR MIB (RFC 3433). - The transport. The agent on the device listens on UDP port 161. RFC 3417 suggests 161 for the agent and UDP 162 for the receiver of notifications. The message carries a version, a credential (community or v3 user) and a PDU (protocol data unit, one request or response).
- The operations. GetRequest fetches named OIDs. GetNextRequest returns the next OID in order, which is how a table is "walked". GetBulkRequest (v2c and v3) returns many rows per round trip, so walking a 48-port switch takes a few packets instead of hundreds. SetRequest writes a value and is rarely enabled for monitoring.
- The poller loop. Discover interfaces once (walk
ifDescr,ifHighSpeed,ifOperStatus). Then every interval read the counters, subtract the previous reading and divide by elapsed seconds to get a rate. Counters are cumulative totals, so a single reading means nothing on its own.
Executed in a Linux container against a net-snmp agent (the manager commands are real; the interface values are whatever that container had):
$ snmpget -v2c -c public -On 127.0.0.1 1.3.6.1.2.1.31.1.1.1.6.1 1.3.6.1.2.1.2.2.1.8.1 1.3.6.1.2.1.31.1.1.1.15.1
.1.3.6.1.2.1.31.1.1.1.6.1 = Counter64: 0
.1.3.6.1.2.1.2.2.1.8.1 = INTEGER: 1
.1.3.6.1.2.1.31.1.1.1.15.1 = Gauge32: 10
$ snmpbulkwalk -v2c -c public -On -Cr5 127.0.0.1 1.3.6.1.2.1.2.2.1.10
.1.3.6.1.2.1.2.2.1.10.1 = Counter32: 0
.1.3.6.1.2.1.2.2.1.10.2 = Counter32: 0
...
Reading it: -On prints OIDs as plain numbers (net-snmp: displays the OID numerically). In the first OID, the trailing .1 after the object's number is the ifIndex, so all three values describe interface 1. The type before each value says how to treat it: Counter64 and Counter32 are cumulative totals that only increase and wrap to zero at their maximum, Gauge32 is a level that can go up and down, and INTEGER is a plain number or enumerated state. The counters read 0 because the test interface had carried no traffic yet; the worked example below shows what a changing counter turns into. The walk output is cut after two rows here, and a real walk prints one line per interface, so the second OID suffix (.1, .2) is how you tell interfaces apart.
The three objects in the first command are ifHCInOctets (64-bit received bytes), ifOperStatus (1 means up; RFC 2863 lists up, down, testing, unknown, dormant, notPresent, lowerLayerDown) and ifHighSpeed (interface speed in units of 1,000,000 bit/s). -Cr5 sets max-repetitions, the number of rows the agent returns per bulk response (net-snmp's default is 10).
Polling versus traps
| Polling (GET / GETBULK) | Traps and informs | |
|---|---|---|
| Who starts it | The manager, on a timer | The device, when an event occurs |
| Good for | Trends, utilisation, capacity, health history | Immediate events: link down, restart, authentication failure |
| Delivery | Request and response, so a lost answer shows as a timeout | Trap: no confirmation at all. Inform: the receiver must acknowledge, though RFC 3416 gives no delivery guarantee |
| Failure mode | Cost scales with devices x interfaces x frequency | A lost trap is silent, so a missed "link down" goes unnoticed |
Run both. Polling is the source of truth and also detects a device that stopped talking; traps (or syslog) shorten the time to notice an event from one poll interval to seconds. A trap is a hint to poll that device now, not a replacement for the poll.
What changes between versions
| v1 (RFC 1157) | v2c (RFC 1901 and RFC 3416 operations) | v3 (RFC 3411 to 3418 family) | |
|---|---|---|---|
| Credential | Community string in clear text | Community string in clear text | Named user, with engine ID and timeliness checks |
| Bulk retrieval | No (GetNext only) | GetBulk | GetBulk |
| Counters | 32-bit only | 64-bit counters (Counter64) available | 64-bit available |
| Notifications | Trap-PDU with its own fields (enterprise, agent-addr, generic-trap, specific-trap, time-stamp) | SNMPv2-Trap (unconfirmed) and Inform (confirmed) | Same PDUs as v2c, protected by USM |
| Security | Effectively none | Effectively none (RFC 1901 is an Experimental-status document) | USM (user-based security model): integrity, origin authentication, confidentiality, replay window of 150 seconds; VACM (view-based access control) limits which objects each user may read or write |
Counter64 is not a cosmetic change. In the same container, asking for the 64-bit counter with -v1 fails:
$ snmpget -v1 -c public -On 127.0.0.1 1.3.6.1.2.1.31.1.1.1.6.1
Error in packet
Reason: (noSuchName) There is no such variable name in this MIB.
Failed object: .1.3.6.1.2.1.31.1.1.1.6.1
RFC 2578 allows Counter64 only where a 32-bit counter would wrap in under an hour. RFC 2863 gives the minimum wrap time of a 32-bit byte counter as 57 minutes at 10 Mbit/s and 34 seconds at 1 Gbit/s, which is why every modern interface should be polled via the ifHC* objects.
For a beginner the takeaway is short: use SNMPv3 at the authPriv level with SHA-based authentication and AES encryption where the device offers them. The algorithm names below are the history behind that choice and what you will meet in device documentation. v3 security levels (RFC 3411): noAuthNoPriv (no authentication, no encryption), authNoPriv (authenticated), authPriv (authenticated and encrypted). Authentication choices have grown from HMAC-MD5-96 and HMAC-SHA-96 (RFC 3414) to the SHA-2 family (RFC 7860); encryption from CBC-DES (RFC 3414) to AES-128 in CFB mode (RFC 3826). Which algorithms a given device accepts varies by vendor and software release, so check before standardising.
Worked example
A switch uplink shows ifHCInOctets = 10,000,000,000 at 10:00:00 and 10,450,000,000 at 10:01:00. The delta is 450,000,000 octets (bytes), so 450,000,000 x 8 / 60 = 60,000,000 bit/s = 60 Mbit/s, which is 6.0% of a 1 Gbit/s link whose ifHighSpeed reads 1000.
Trade-offs and pitfalls
- Community strings sent over UDP can be read by anyone on the path, and a default such as
publicis guessed first by scanners. Use v3 authPriv where devices support it; where only v2c exists, restrict by source address and make it read-only. - Poll interval is a trade: 60 seconds is a common default; faster polling raises device CPU and collector load, slower polling smooths away short congestion.
- Counters reset when a device restarts. Compare
sysUpTime(time since the agent started) between polls and discard the rate when it went backwards. - An unanswered poll is a data point: alert on "no data" separately from "value is bad".
Which interface-level metrics would you monitor on routers and switches to catch congestion, errors and degradation early? For each, explain what it tells you and how to read it.
Sample Answer
Direct answer
Monitor eight things per interface: state, traffic, speed, errors, discards, packet-type counts, optics and queue drops. The first six come from the standard IF-MIB (RFC 2863, the standard table of interface counters a device exposes over SNMP); optics and per-class queue drops are usually vendor-specific. If you can only start with a few, metrics 1 to 5 (state, octets, speed, errors and discards) already catch most congestion and cable trouble; the other three (packet-type counts, optics and queue drops) sharpen the diagnosis. Read every counter as a rate (the change between two polls divided by elapsed seconds) and every error counter as a ratio against packets, never as a raw total. Traffic tells you load, errors tell you physical trouble, discards tell you queue pressure or policy.
The eight metrics
| # | Metric (object) | What it tells you | How to read it |
|---|---|---|---|
| 1 | Operational and administrative state (ifOperStatus, ifAdminStatus), plus ifLastChange | Whether the interface is up. Admin up with oper down is a fault; admin down is deliberate. ifLastChange is the sysUpTime when it entered its current state | Alert on admin-up and oper-down for important links. Frequent changes in ifLastChange indicate a flapping link |
| 2 | Octet counters (ifHCInOctets, ifHCOutOctets) | Traffic volume, hence load | Rate = (delta x 8) / seconds, compared per direction. 64-bit counters because a 32-bit counter wraps in 34 s at 1 Gbit/s |
| 3 | Speed (ifHighSpeed) | The denominator for utilisation, in units of 1,000,000 bit/s | Utilisation = rate / (ifHighSpeed x 1,000,000). The older ifSpeed tops out at 4,294,967,295, so fast interfaces report a wrong value |
| 4 | Errors (ifInErrors, ifOutErrors) | Inbound packets with errors that prevented delivery; outbound packets that could not be sent because of errors | Ratio to packets. A clean link should show none; sustained growth typically means a bad cable, dirty or failing optic, or a speed or duplex mismatch (duplex mismatch: one end sends and receives at the same time, the other end takes turns, so they collide) |
| 5 | Discards (ifInDiscards, ifOutDiscards) | Packets dropped although nothing was wrong with them | Output discards typically mean the egress queue (the buffer where packets wait to leave the port) overflowed (congestion); input discards can mean policy or resource limits. Ratio to packets |
| 6 | Packet-type counters (ifHCInUcastPkts, ifHCInMulticastPkts, ifHCInBroadcastPkts, and the Out equivalents) | The mix of traffic. A sudden jump in broadcast or multicast share suggests a loop or a misbehaving host | Share of total packets and its change over time |
| 7 | Optical and environmental readings (ENTITY-SENSOR, RFC 3433, the standard table for sensor readings such as temperature; per-transceiver power from vendor MIBs) | A transceiver degrading before errors appear | Trend against the module's own thresholds, rather than a fixed value |
| 8 | Queue and buffer drops per class (vendor-specific MIBs, or streaming telemetry, where the device pushes counters every few seconds instead of waiting to be polled) | Which traffic class is being dropped, which ifOutDiscards cannot say | Compare the class that matters (voice, storage) with its configured share, the slice of the link that quality of service (QoS) rules reserve for it |
Reading the numbers
- Always a rate from two readings. Counters are cumulative; a poller subtracts the earlier reading from the later one and divides by the real elapsed time (use the poll timestamps, not the nominal interval).
- Handle wrap and reset. Per RFC 2863 a manager must discard a difference when
ifCounterDiscontinuityTimechanged between the two polls, in addition to checkingsysUpTimefor agent restart. - Direction matters. On a full-duplex link each direction has its own capacity. A link at 20% in and 95% out is congested outbound.
- Errors versus discards diagnose different layers. Errors point at layer 1 (the physical layer: cable, optic, signal quality); discards point at layers 2 to 3 (the frames and IP packets that the device itself receives and forwards, so queues and policy). If both rise on one port, check the physical layer first, because a faulty link also causes retransmissions that add load.
Worked example
An uplink with ifHighSpeed = 1000 over one 60 s interval: output octets +3,750,000,000; output packets +4,500,000 with +45,000 ifOutDiscards; input packets +3,000,000 with +120 ifInErrors.
- Output utilisation: 3,750,000,000 x 8 / 60 = 500 Mbit/s, 50% of 1,000 Mbit/s.
- Discards: 45,000 / 4,500,000 = 1.0%. Healthy utilisation with 1% discards means bursts overflow the queue, so look at queue drops (metric 8) and consider buffer or QoS changes.
- Errors: 120 / 3,000,000 = 0.004%. Small, but it matters if it keeps climbing.
- For contrast (illustrative numbers), the same link in a healthy minute carrying the same 4,500,000 output packets shows 0 discards, which is 0%, and 0 errors. The unhealthy signs are the discards ratio moving off zero and the error count growing from one poll to the next.
Pitfalls
- Averages hide bursts: a 60 s average of 50% can contain 100 ms bursts at line rate. Discards are the cheap signal that those bursts happened.
- Using 32-bit counters on fast links or
ifSpeedfor utilisation silently produces wrong charts. - Port channels (several physical links grouped into one logical link, also called a bundle): monitor member links as well as the bundle, since the bundle's average can hide one member that is saturated or erroring.
- Counters for a given
ifIndex(the integer a device uses to number each interface) can change after a reboot on some devices; map by interface name, not index alone (RFC 2863 notes ifIndex values may be reassigned whensysUpTimeresets).
You want to detect congestion, rising packet loss, jitter and device failure before users complain, in a medium enterprise network. What would you collect, how often, and where would you set initial alert thresholds?
Sample Answer
Direct answer
Collect five things: interface counters (traffic, errors, discards), device health (CPU, memory, temperature, power, fans), reachability, active probe results between sites (latency, jitter, loss), and syslog. Poll counters every 60 seconds, run probes every 10 seconds, and start with duration-based thresholds: warn at 70% link utilisation and alert at 85%, both sustained for 10 minutes; flag any rising error counter; flag discards above 0.1% of packets. These are starting values to tune against 30 days of your own baseline, not standards.
What to collect, by failure type
| Problem | Signal | Source | Interval | Initial threshold |
|---|---|---|---|---|
| Congestion | Utilisation = (delta octets x 8) / (seconds x interface speed) per direction | SNMP ifHCInOctets, ifHCOutOctets, ifHighSpeed | 60 s | Warn 70%, critical 85%, for 10 consecutive minutes |
| Congestion (queue overflow) | Outbound discards as a share of outbound packets (packets dropped because the outbound queue was full) | SNMP ifOutDiscards and packet counters | 60 s | Above 0.1% for 10 minutes |
| Rising packet loss on a link | Input and output errors as a share of packets | SNMP ifInErrors, ifOutErrors | 60 s | Warn on any increase for 3 consecutive polls; critical above 0.01% of packets |
| Loss anywhere on a path | Probe loss percentage | Active probes (ICMP and UDP) between sites | 10 s probes, evaluated over 5 minutes | Warn above 1%, critical above 5% |
| Jitter | Probe delay variation | Active UDP probes | 10 s | Above 2x the 30-day baseline for 5 minutes |
| Device failure | Reachability, restart, CPU, memory, temperature, power and fan state | ICMP, sysUpTime, vendor or ENTITY-SENSOR objects (RFC 3433) | 10 s ping, 60 to 300 s health | 3 missed probes; uptime reset; CPU above 80% for 10 minutes; any fan or power fault |
Why each choice:
- Utilisation must be read per direction. A full-duplex 1 Gbit/s link can carry 1 Gbit/s in each direction, so an input average of 40% and an output average of 90% means the output is the problem.
- Discards and errors differ. RFC 2863 defines errors as packets that contained errors preventing delivery, and discards as packets dropped even though no error was detected (typically a full queue). Errors suggest a physical fault; discards suggest congestion or policy.
- Interface counters cannot see loss inside a provider network. Only probes between your endpoints see it.
- Jitter needs a baseline, not a universal number. A path that is naturally variable (for example a satellite or a congested internet VPN) would alert constantly on a fixed value, so compare to its own history.
- A duration requirement stops flapping. One 60 s burst to 95% is normal; ten minutes is a trend.
Worked example
A 1 Gbit/s uplink shows these deltas over one 60 s poll: 6,000,000,000 output octets and 6,000,000 output packets, of which 9,000 were discarded; on the input side 4,000,000 packets and 600 errors.
- Output utilisation: 6,000,000,000 x 8 / 60 = 800,000,000 bit/s = 80%. Above the 70% warning level, but the rule needs 10 consecutive minutes, so one poll does not alert.
- Output discards: 9,000 / 6,000,000 = 0.15%, above the 0.1% threshold. If it persists for 10 minutes it alerts, and it is the stronger signal that the queue is overflowing.
- Input errors: 600 / 4,000,000 = 0.015%, above the 0.01% critical level, so this points to a physical fault on the receive side: the cable, the optic (the pluggable laser module in the port), or the transmitter of the device at the other end of the cable.
Trade-offs and pitfalls
- Alert on rates and ratios computed from two successive counter readings, never on the raw counter. A 32-bit octet counter wraps in about 34 seconds at 1 Gbit/s; use the 64-bit
ifHCobjects. - Probe loss and interface errors do not always agree: run both and look at where they diverge.
- After a month, replace fixed numbers with per-link baselines (for example hour-of-week percentiles: for each link, keep the readings taken in the same hour of the same weekday over the past several weeks, such as every Tuesday 09:00 to 10:00, and alert when the current value is above, say, the 95th percentile of those readings, meaning higher than 95 out of every 100 of them), so a branch that always runs hot does not alert and a quiet link that suddenly doubles does.
- Do not alert on every port. Server-facing ports get status history; infrastructure links and uplinks get thresholds.
A cross-continental replication link has 300 ms RTT and 1 Gbps of capacity, but transfers reach only a fraction of it. Work out what limits throughput, what you would change on the hosts, and how you would validate each change.
Sample Answer
Direct answer
The limit is almost certainly the TCP window, not the link. A sender can only have one window of unacknowledged data in flight per round trip, so throughput is at most window divided by round-trip time (RTT). To fill 1 Gbit/s at 300 ms RTT you need the bandwidth-delay product (BDP) in flight, 37.5 MB, and a default host configuration, or any loss on the path, keeps far less than that in the air. On the hosts I would confirm window scaling is negotiated, raise the receive and send buffer ceilings on both ends to cover the BDP, then check for loss and for a policer on the path, validating each step with a long, single-variable iperf3 run.
The arithmetic
BDP = capacity x RTT = 1,000,000,000 bit/s x 0.300 s / 8 = 37,500,000 bytes (35.8 MiB)
Throughput <= window / RTT
64 KiB window (no scaling): 65,535 x 8 / 0.3 = 1.75 Mbit/s
6 MiB window: 167.8 Mbit/s
16 MiB window: 447.4 Mbit/s
32 MiB window: 894.8 Mbit/s
64 MiB window: 1000 Mbit/s (link-limited)
Window scaling (RFC 7323) multiplies the 16-bit window field by a shift of at most 14, giving a maximum window of 1 GiB, which would allow 28.6 Gbit/s at this RTT, so the protocol is not the ceiling. The option is only sent in the initial SYN and SYN-ACK (the first two packets of the TCP handshake), so a middlebox (a firewall, NAT or proxy sitting on the path) that strips it or a host with net.ipv4.tcp_window_scaling set to 0 (default on Linux is 1) leaves you at the 64 KiB figure. The Linux kernel documentation gives the receive-buffer ceiling (third value of net.ipv4.tcp_rmem) as between 131072 bytes and 32 MB depending on RAM, and the send-buffer ceiling (third value of net.ipv4.tcp_wmem) as between 64K and 4 MB depending on RAM. The sender must keep every unacknowledged byte buffered, so a small send ceiling limits you as much as a small receive window.
Loss caps you too
Even with big buffers, loss resets the window: TCP's congestion control (the algorithm that decides how fast the sender may go) cuts its congestion window (cwnd, the sender's own limit on unacknowledged data) when it sees a loss and then regrows it slowly, and on a 300 ms path the regrowth takes many round trips. RFC 5348 gives a throughput model for standard TCP, in bytes per second:
X = s / (R x (sqrt(2p/3) + 12 x sqrt(3p/8) x p x (1 + 32 x p^2)))
Here s is the segment size in bytes, R the RTT in seconds and p the loss event rate (roughly, the fraction of packets that begin a loss). The RFC takes one packet per acknowledgement and a retransmission timeout of 4R, which is where the 12 comes from. The second term is the penalty for timeouts and is tiny when p is small. At s = 1448 bytes and R = 0.3 s, with p = 1e-6 as the worked case: sqrt(2 x 1e-6 / 3) = 0.000816, the timeout term is 0.0000000073 and can be ignored, so R x (0.000816) = 0.000245 s, and 1448 x 8 = 11,584 bits divided by 0.000245 s is 47.3 Mbit/s. The other rows use the same steps. I computed:
p = 1e-6 -> 47 Mbit/s
p = 1e-8 -> 473 Mbit/s
p for 1 Gbit/s -> about 2.2e-9 (roughly one loss event per 450 million packets)
The model describes classic AIMD (additive increase, multiplicative decrease) behaviour, and CUBIC, the congestion-control algorithm Linux uses by default, does somewhat better on long paths, so treat these as order of magnitude. The conclusion holds: at 300 ms RTT a clean path matters, so a replication link that shows even 0.001 percent loss (p = 1e-5, roughly 15 Mbit/s by the model) needs a loss fix before any tuning pays off.
What I would change on the hosts (in this order)
- Confirm window scaling is on at both ends and survives the path (check with
ss -tiduring a transfer, thewscale:field is present when negotiated). - Raise the ceilings on sender and receiver, for example
sysctl -w net.ipv4.tcp_rmem="4096 131072 67108864" net.ipv4.tcp_wmem="4096 16384 67108864"(min, default, max in bytes; 64 MiB is about 1.8 times the BDP). Only connections that need the room use it, but many such connections multiply worst-case memory, so size the maximum to the BDP of your slowest-to-fill links, not to every connection. - Fix loss: look at interface errors and discards on the path, a policer (a rate limiter that drops traffic above a committed rate) or shallow buffer on the provider circuit, and a path maximum transmission unit (MTU) problem (
ping -M do -s 1472 <peer>fails if a 1500-byte packet cannot pass with the do-not-fragment bit set). - If the application can, use several parallel streams. Each stream is limited by its own window, so with the roughly 84 Mbit/s per-stream figure measured below you need about 12 streams (1000 / 84).
- Only then test another congestion-control algorithm (
net.ipv4.tcp_congestion_control; iperf3 selects one per test with-C).
Executed lab and how to validate each change
I reproduced the path in a Linux container with NET_ADMIN: ip link set dev lo mtu 1500 then tc qdisc add dev lo root netem delay 150ms rate 1gbit limit 60000 (150 ms each way, so ping showed an RTT of about 300 ms), iperf3 -s -D, and iperf3 -c 127.0.0.1 -t 30 -O 10 (omit the first 10 s of ramp-up from the average). Reading the setup: netem is the Linux network emulator, a queueing discipline that damages traffic on purpose. delay 150ms holds every packet 150 ms before sending, and because loopback sends both directions through that one qdisc the round trip is 300 ms. rate 1gbit makes packets take the time a 1 Gbit/s link would take. limit 60000 is the most packets the emulator may hold while delaying them, set large enough for a full window in flight (37,500,000 / 1448 is about 26,000 packets). iperf3 is a throughput test tool: -s -D starts its server in the background and -c runs the client. Results, single stream:
| Configuration | iperf3 result |
|---|---|
| Container defaults (tcp_rmem max 32 MiB, tcp_wmem max 4 MiB) | 84.8 Mbit/s in one 20 s run, 139 Mbit/s in a 30 s run |
-w 64K (sets a 64 KiB socket buffer) | 2.19 Mbit/s |
Both ceilings raised to 64 MiB (docker run --sysctl) | 702 Mbit/s over a 40 s run (still ramping) |
The 64 KiB case matches the 1.75 Mbit/s calculation in order of magnitude. The default case shows buffers and ramp-up at work: ss -ti (which prints per-connection TCP internals) about 20 s into a default run showed cwnd:2296 and rwnd_limited:1895ms(10.1%). Reading those fields: cwnd:2296 is the congestion window in segments, 2296 x 1448 = 3,324,608 bytes (about 3.3 MB in flight, close to the 4 MiB send-buffer ceiling and under a tenth of the 37.5 MB BDP). wscale: (not quoted from this run) would show the window-scale shifts when negotiated. rwnd_limited is the time the sender had to hold back because the receiver's advertised window (rwnd, the free buffer space the receiver announces) was full. The percentage is a share of the connection's busy time, the time it was actively sending data (the busy: field in the same line), not of the wall-clock time since it started: 1895 ms / 0.101 is about 18.8 s of busy time, so dividing by a flat 20 s gives 9.5 percent and the two figures are consistent. In words, the receiver's window held the sender back for about a tenth of its busy time. Run-to-run variation is large (84 versus 139), so repeat each test at least three times and compare medians.
| Change | Check before and after | Pass condition |
|---|---|---|
| Window scaling | ss -ti shows wscale: on the established connection | present, nonzero shift |
| Buffer ceilings | iperf3 -c <peer> -t 40 -O 10, then -R | at least 800 Mbit/s averaged after the ramp, both directions (my chosen bar; set yours from the application need) |
| Loss fixed | iperf3 TCP Retr column (segments the sender had to resend), plus iperf3 -u -b 500M lost/total | retransmits near zero, UDP loss near 0 percent |
| Path MTU | ping -M do -s 1472 | replies, no "message too long" |
| Parallel streams | iperf3 -P 12 aggregate | roughly the sum, confirms per-flow limit |
Change one thing per run, start a fresh connection for each test, and keep a one-line log of settings and results.
Trade-offs and pitfalls
A 37.5 MB window can queue that much data at the bottleneck. If the provider's buffer is smaller, the burst is dropped and the loss model takes over, so ramp tests gently and watch discards. Also check whether the 1 Gbit/s is a policed commitment: an iperf3 -u run above the committed rate showing loss at one fixed ceiling points to a policer, not a window problem. Finally, remember tuning is per host pair: the receiver's settings matter as much as the sender's.
What is the difference between active health checks and passive telemetry for network monitoring? Give examples of data from each and say when you rely on which.
Sample Answer
Direct answer
Passive telemetry is data the network and its devices already produce about real traffic and their own state: interface counters, flow records, syslog, traps, streaming telemetry, packet captures. Active health checks are traffic you create on purpose and measure: a ping, a TCP connect, an HTTP request, a DNS query. Passive tells you what is happening to traffic that exists and why a device behaves as it does. Active tells you whether a path works and how fast right now, even when no user is using it. Rely on active checks to detect that users are hurt, and on passive data to find out where and why.
Side by side
| Active (synthetic) | Passive | |
|---|---|---|
| Source | Probes you schedule | Counters, flow records, logs, traps, telemetry, captures |
| Examples | ICMP echo (ping) round-trip time and loss; TCP connect time to a service port; HTTP request status and phase timings; DNS query success and time; TWAMP (Two-Way Active Measurement Protocol, RFC 5357) round-trip IP performance (a standard that measures delay and loss by sending test packets between two cooperating endpoints) | ifHCInOctets, ifInErrors and ifInDiscards counters read with SNMP (Simple Network Management Protocol) from the interface MIB (Management Information Base, the standard data tables); flow records; linkDown and linkUp traps (unsolicited SNMP notifications from the device); syslog; gNMI (gRPC Network Management Interface) subscriptions, a stream where the device pushes values to the collector |
| Adds load | Yes, a small amount you control | None to the data path (counter reads add device CPU work) |
| Sees real user traffic | No, a proxy for it | Yes, the actual traffic |
| Detects a broken path with no traffic on it | Yes | No: counters are flat and look healthy |
| Explains cause | Weak: says that, not why | Strong: errors, drops, utilization, who is talking |
| Blind spot | Probe path may differ from the user path; probe load may be treated differently from user traffic | Interface up and counters clean while a dependency (DNS, a firewall policy, an upstream provider) is failing |
Worked example
A branch complains that the internal app is down. Passive view: the uplink is up (ifOperStatus is up), utilization is 38 percent and error counters are zero, so nothing looks wrong. Active view: the branch probe host's TCP connect to the app on port 443 fails and the DNS probe for the app name times out, while ping to the hub router succeeds. Read as a ladder (the layer ladder: reachability, then name resolution, then connect, then application request), the first rung that fails marks the layer: reachability to the hub passes and name resolution fails. The TCP and HTTP probes look the application up by name, so their failures follow from the DNS failure and add no separate evidence. The failure is in name resolution, not the link. Passive data alone would have said "network fine" and active data located the layer in one look. Next, passive data comes back into play: flow records show no DNS queries leaving the branch to the resolver, pointing at the branch resolver or its firewall rule.
Detection time differs by method. A 60-second SNMP poll sees a link-down within 60 seconds in the worst case, and a linkDown trap (sent on transition into the down state, RFC 2863) arrives within moments. A probe every 10 seconds that must fail 3 times in a row detects an outage in 3 x 10 = 30 seconds, and it also catches failures that never change an interface state (a blackholed route, where packets are silently discarded because a route points nowhere useful, dead DNS, broken policy).
When I rely on which
| Situation | Primary | Why |
|---|---|---|
| Alert that users cannot reach a service | Active | It measures the symptom directly and covers failures with no interface event |
| Find which hop or device is at fault | Passive | Counters, errors and flows localize it |
| Capacity planning and trend | Passive | It reflects real load over months |
| Baseline of path latency when idle or at night | Active | No traffic needed |
| Security or "who sent that" questions | Passive (flow, logs) | Records of real conversations |
| Verifying a change or failover worked | Both | Probes confirm service, counters confirm traffic moved |
Pitfalls
- A probe to a router's own interface can be rate-limited or deprioritized by the router. For user-path measurements, probe through the router to a host.
- A single probe type is blind: ICMP success says nothing about a failed TCP service. Use the layer ladder (reachability, name resolution, connect, application request).
- Alert on symptoms, not every counter: paging on utilization over 80 percent for a link nobody is complaining about trains people to ignore pages.
That is every published Network Monitoring and Performance question for Systems Administrator so far. Browse the other topics in this category, or practice this one interactively.