Network Monitoring and Performance Questions
Network telemetry and performance operations: SNMP polling and traps (including counter wraparound and SNMPv3 access), NetFlow, sFlow and IPFIX flow export, sampling and its accuracy, streaming telemetry (gNMI), and eBPF or packet-capture telemetry; interface-level metrics (utilization, errors, discards, queue depth, microbursts), active synthetic probing alongside passive counters, link-flap detection, baselining and anomaly detection on network signals including elephant-flow spotting, network SLIs and SLOs, alerting, alert-storm suppression and NOC dashboards, telemetry pipeline design, storage, retention, downsampling and cardinality for network data (including securing the collection path and handling bursty remote sites), BGP and link-state monitoring including prefix hijack and route-leak detection, and network capacity monitoring, percentile utilization and bandwidth headroom planning. Also covers measuring and tuning network-level latency, jitter, packet loss and throughput (bufferbloat, queueing, TCP tuning for long paths). Excludes the generic metrics, logs and traces stack and alert design, the layered fault-isolation method and packet-capture troubleshooting, TCP and protocol fundamentals, application and CDN latency engineering, cloud VPC design and security detection.
Here is a small enterprise: two data centers with core, aggregation and about 50 leaf switches, internet edge routers, and WAN links to five branch offices. Design a pragmatic monitoring plan: what you monitor and how often, what you alert on, and what dashboards the operations team needs.
Sample Answer
Direct answer
Use SNMP v3 polling for health and counters, syslog and traps for events, sampled flow at the internet edge and WAN, and active probes from each branch to both data centers. Poll interface counters every 60 seconds, test reachability of the 21 core, aggregation, edge and branch devices every 10 seconds, and page only on conditions that need a human now (a tier-1 device down, an internet-edge BGP session lost, a WAN link down). Everything else becomes a ticket or a trend. Five dashboards cover the operations team.
Assumed inventory (so the numbers are checkable)
Per data center: 2 core, 4 aggregation, 25 leaf switches and 2 internet-edge routers, so 50 leaf, 4 edge, 4 core and 8 aggregation devices across the two sites. In addition there are 5 branch routers in total, one per branch office (not one per branch per data center). A leaf switch is the top-of-rack switch that servers plug into; core and aggregation are the layers above it.
| Tier | Devices | Count |
|---|---|---|
| 1: core, aggregation, edge, branch routers | 2x2 core + 2x4 aggregation + 4 edge + 5 branch | 21 |
| 2: leaf | 2 x 25 | 50 |
| Total | 71 |
Interface counts are assumptions I state so the volume check can be redone. A leaf has 48 ports: 44 face servers and 4 are uplinks, so 50 x 44 = 2,200 server ports and 50 x 4 = 200 leaf uplinks. I assume the other 21 devices carry about 24 interfaces each on average (downlinks, interconnects, edge and WAN ports), roughly 500, so 200 + 500 = 700 infrastructure interfaces.
What to monitor and how often
| Signal | Source | Interval | Scope |
|---|---|---|---|
| Reachability (is it up) | ICMP echo plus an SNMP read of sysUpTime (time since the device last restarted) | 10 s tier 1, 30 s leaf | all 71 devices |
Interface state and traffic: ifOperStatus, ifHCInOctets and ifHCOutOctets, ifInErrors, ifOutErrors, ifInDiscards, ifOutDiscards | SNMP v3 authPriv, GetBulk | 60 s | infrastructure links; server ports get state only |
| Device health: CPU, memory, temperature, fans, power supplies | Vendor MIBs or ENTITY-SENSOR (RFC 3433) | 60 to 300 s | all devices |
| Routing state: BGP peer state, OSPF or IS-IS neighbours | BGP4-MIB bgpPeerState (RFC 4273), traps such as bgpBackwardTransNotification, syslog | trap on event plus 60 s poll | edge, core, WAN |
| Events and config changes | Syslog to a central collector | continuous | all devices |
| Path quality: latency, loss, jitter | Active probes (ICMP and UDP) from each branch router or a probe host to a host in each data center | every 10 s | 5 branches x 2 data centers = 10 paths |
| Who is using the bandwidth | Sampled IPFIX/NetFlow or sFlow | continuous | internet edge and WAN only |
Volume check: the interface row polls 7 objects per infrastructure interface (ifOperStatus plus 6 counters: in and out octets, in and out errors, in and out discards), so 700 x 7 = 4,900 values per cycle. Server ports get state only: 2,200 values. That is 7,100 values every 60 s, about 118 values per second, 10,224,000 per day. This is the planning basis for sizing the time-series database (TSDB, a database built to store timestamped measurements); validate it with a load test before purchase.
What to alert on
Thresholds below are starting points I would choose, to be replaced by baselined values after about 30 days of data; they are not standards.
| Level | Condition | Why |
|---|---|---|
| Page now | Tier-1 device fails 3 consecutive 10 s probes (30 s) | Large blast radius (many users or sites lose service when it fails) |
| Page now | Internet-edge BGP session leaves Established (BGP, the routing protocol with the ISP, reports Established only while the session is healthy), or a WAN link goes oper-down (the interface's operational status, ifOperStatus, reports down) | User-visible outage |
| Page now | All four uplinks of a leaf are down | A rack is isolated |
| Ticket, same day | Link above 85% utilisation (warning at 70%) for 10 consecutive minutes | Congestion building |
| Ticket, same day | Error counters rising for 3 consecutive polls, or discards above 0.1% of packets for 10 minutes | Physical fault or queue overflow |
| Ticket, same day | Probe loss above 1% over 5 minutes, or latency or jitter above 2x baseline | Degraded path |
| Ticket, next day | Fan or power-supply fault with redundancy intact; CPU above 80% for 10 minutes | Fix before it becomes an outage |
| Ticket, same day | A device's sysUpTime went backwards | Unplanned restart |
Make alerts depend on topology: if an aggregation switch is down, suppress alerts for the leaves behind it and report one root cause. Send "no data received" as its own alert, since a dead poller looks like a quiet network.
Dashboards (five)
- Operations overview: green/amber/red per site, active alerts by severity, tier-1 reachability.
- WAN and branches: per-branch link utilisation, probe latency, loss and jitter, BGP or tunnel state, top talkers from flow.
- Data center fabric: leaf-uplink utilisation shown as a heat map per rack (a grid coloured by value, so the hot rack stands out), errors and discards, imbalance between a leaf's uplinks.
- Internet edge: BGP sessions, inbound and outbound traffic against the contracted rate, top talkers and destinations.
- Capacity: 95th percentile utilisation per link (sort the month's samples, discard the top 5%, read the highest remaining) and a 90-day trend, so upgrades are planned before alerts fire.
Trade-offs and pitfalls
- Per-port alerts on 2,200 server ports would bury the team; server ports get state history on a dashboard, not pages.
- Flow export everywhere multiplies collector cost and double-counts transit; edge and WAN already answer the "who is using the link" question.
- Streaming telemetry (gNMI) is the next step for the fabric where leaf platforms support it; it is a phase-2 upgrade because SNMP covers the stated needs.
- Rollout order: week 1 reachability and interface state, week 2 alerts with the starting thresholds, week 3 probes and syslog, week 4 flow and dashboards.
A poller reads a 32-bit octet counter on a 1 Gbps interface every 60 seconds. Two successive readings are 4,100,000,000 and 215,000,000. What was the average utilization over that interval, what happened between the readings, and how should a poller handle this in general? Would anything change at a 5-minute interval or on a faster link?
Sample Answer
Direct answer
The 32-bit counter wrapped. A 32-bit octet counter counts to 2^32 - 1 = 4,294,967,295 and then restarts at 0. The first reading (4,100,000,000) was 194,967,296 short of the limit, so the true increase is (4,294,967,296 - 4,100,000,000) + 215,000,000 = 409,967,296 octets. Over 60 s that is 409,967,296 x 8 / 60 = 54.66 Mbit/s, or 5.47% of 1 Gbit/s. But the two readings alone cannot prove a single wrap: two wraps would mean 4,704,934,592 octets, which is 627.3 Mbit/s (62.7%), and that is also below the line rate. The correct poller fix is to poll the 64-bit ifHCInOctets counter, and a 5-minute interval or a faster link makes the 32-bit counter strictly worse.
Why the 60 s reading is ambiguous
At line rate, 1 Gbit/s moves 125,000,000 octets per second, so 60 s can carry at most 7,500,000,000 octets, which is 1.75 wraps of a 32-bit counter. Every candidate must be non-negative and not exceed the line rate:
| Wraps in the interval | Octets | Mbit/s over 60 s | Utilisation | Possible? |
|---|---|---|---|---|
| 1 | 409,967,296 | 54.7 | 5.5% | yes |
| 2 | 4,704,934,592 | 627.3 | 62.7% | yes |
| 3 | 8,999,901,888 | 1,200.0 | 120% | no, exceeds link speed |
Each row adds one more 4,294,967,296-octet wrap to the 409,967,296 base, so the 3-wrap candidate is 409,967,296 + 2 x 4,294,967,296 = 8,999,901,888 octets.
So the answer is: 5.47% if the counter wrapped once, which is the standard assumption, but the data cannot exclude 62.7%. At 1 Gbit/s a 32-bit counter wraps in 2^32 x 8 / 10^9 = 34.4 s (RFC 2863: "at 1Gbs, the minimum is 34 seconds"), so a 60 s poll is already too slow to be trusted.
What a poller should do in general
- Use 64-bit counters (
ifHCInOctets,ifHCOutOctets) for anything at or above 10 Mbit/s. A 64-bit counter at 1 Gbit/s wraps in about 4,700 years; at 100 Gbit/s about 47 years. - Treat a decrease as a wrap only for 32-bit counters, add 2^32, and reject the sample if the implied rate exceeds the interface speed (
ifHighSpeedx 1,000,000 bit/s). That rejects reboots and ambiguous cases instead of charting spikes. - Detect resets.
sysUpTimeis the agent's time since it last started, counted in TimeTicks (hundredths of a second); if it went backwards the agent restarted.ifCounterDiscontinuityTimeis thesysUpTimevalue at the last moment any of the interface's counters was disturbed (for example reset), and per RFC 2863 a manager must discard a difference when it differs between the two polls. Neither is a wrap. - Poll faster than the wrap time if you must use 32 bits: at 1 Gbit/s under 34 s, at 10 Gbit/s under 3.4 s, at 100 Gbit/s under 0.34 s. At 10 Gbit/s and above, 32-bit counters are not usable at all.
- Use real timestamps for the elapsed time rather than the scheduled interval.
Code (Python 3.12), executed in a container:
WRAP32 = 2**32
def octet_rate(prev, curr, seconds, speed_bps, bits=32):
"""Return (bits_per_second, note) or (None, reason) for two counter readings."""
wrap = 2**bits
delta = curr - prev
note = "ok"
if delta < 0:
if bits != 32:
return None, "64-bit counter decreased: reset, discard sample"
delta += wrap
note = "wrapped once"
bps = delta * 8 / seconds
if bps > speed_bps:
return None, "rate above link speed: reset, multiple wraps or bad speed"
return bps, note
prev, curr = 4_100_000_000, 215_000_000
bps, note = octet_rate(prev, curr, 60, 1_000_000_000)
print(f"32-bit, one wrap assumed: {bps/1e6:.2f} Mbit/s, {bps/1e9*100:.2f}% ({note})")
alt = (curr + WRAP32 - prev + WRAP32) * 8 / 60
print(f"32-bit, two wraps: {alt/1e6:.2f} Mbit/s, {alt/1e9*100:.2f}%")
print("wrap time at 1 Gbit/s: %.1f s, at 10: %.2f s, at 100: %.2f s" % tuple(WRAP32 * 8 / s for s in (1e9, 10e9, 100e9)))
Output:
32-bit, one wrap assumed: 54.66 Mbit/s, 5.47% (wrapped once)
32-bit, two wraps: 627.32 Mbit/s, 62.73%
wrap time at 1 Gbit/s: 34.4 s, at 10: 3.44 s, at 100: 0.34 s
Would a 5-minute interval or a faster link change it?
- 5 minutes (300 s) at 1 Gbit/s: the same two readings now imply 409,967,296 x 8 / 300 = 10.93 Mbit/s if one wrap occurred. But 300 s can carry up to 37,500,000,000 octets, which is 8.73 wraps, and each extra wrap adds 4,294,967,296 x 8 / 300 = 114.5 Mbit/s. The candidates are 10.9, 125.5, 240.0, 354.5, 469.1, 583.6, 698.1, 812.7 and 927.2 Mbit/s: nine possible answers (0 to 8 extra wraps; a ninth extra wrap would give 1,041.7 Mbit/s, above line rate). The data cannot tell you which is true, and the usual "assume one wrap" answer underestimates by up to about 85 times.
- A faster link: at 10 Gbit/s the counter wraps every 3.44 s, so even a 10 s poll is hopeless; use 64-bit counters. RFC 2863 calls for 64-bit octet counters on interfaces that can run faster than 20 Mbit/s (and 64-bit packet counters from 650 Mbit/s up), which is why the
ifHCobjects exist.
The same arithmetic in a Go poller that exposes a metric
package main
import "fmt"
type Sample struct {
Octets uint64 // ifInOctets (32-bit) or ifHCInOctets (64-bit) reading
Uptime uint64 // sysUpTime in TimeTicks (hundredths of a second)
}
func rateBps(prev, curr Sample, seconds float64, bits uint, speedBps float64) (float64, error) {
if curr.Uptime < prev.Uptime {
return 0, fmt.Errorf("sysUpTime went backwards: agent restarted, discard sample")
}
var delta uint64
if curr.Octets >= prev.Octets {
delta = curr.Octets - prev.Octets
} else if bits == 32 {
delta = (1 << 32) - prev.Octets + curr.Octets
} else {
return 0, fmt.Errorf("64-bit counter decreased: reset, discard sample")
}
bps := float64(delta) * 8 / seconds
if bps > speedBps {
return 0, fmt.Errorf("%.0f bit/s exceeds link speed: discard sample", bps)
}
return bps, nil
}
func main() {
prev := Sample{Octets: 4_100_000_000, Uptime: 100_000}
curr := Sample{Octets: 215_000_000, Uptime: 106_000}
bps, err := rateBps(prev, curr, 60, 32, 1e9)
if err != nil {
fmt.Println(err)
return
}
fmt.Println("# TYPE interface_in_bits_per_second gauge")
fmt.Printf("interface_in_bits_per_second{device=\"sw1\",ifindex=\"1\"} %.0f\n", bps)
fmt.Printf("utilisation_percent=%.2f\n", bps/1e9*100)
rebooted := Sample{Octets: 50_000, Uptime: 500}
if _, err := rateBps(curr, rebooted, 60, 32, 1e9); err != nil {
fmt.Println("after reboot:", err)
}
}
Called with the two readings (sysUpTime 100,000 and 106,000 ticks, 60 s apart) and printed in the Prometheus text format, it produced the output below. In that format a line is name{label="value"} value, and a # TYPE name gauge line before it declares that the value can go up or down. The label values sw1 and ifindex="1" are illustrative names for the device and interface.
# TYPE interface_in_bits_per_second gauge
interface_in_bits_per_second{device="sw1",ifindex="1"} 54662306
utilisation_percent=5.47
after reboot: sysUpTime went backwards: agent restarted, discard sample
Reading the Go code step by step. Sample holds one reading: uint64 is an unsigned 64-bit whole number, wide enough to hold either counter and to do the wrap arithmetic without overflowing, and Uptime is sysUpTime. The bits parameter says whether the counter is 32 or 64 bits wide. The function works through four cases in order and then applies a rate check. First, if the uptime went backwards (not the case here: 106,000 is above 100,000, and 6,000 ticks is 60 s), the agent restarted, so the sample is discarded. Second, if the counter did not go down, the increase is a plain subtraction. Third, if it went down and the counter is 32-bit, the increase is (1 << 32) - prev + curr, where 1 << 32 means shift the number 1 left by 32 binary places, which is 2^32 = 4,294,967,296, the same wrap size used in the first section. Fourth, if a 64-bit counter went down, a real wrap is not credible, so it is a reset and the sample is discarded. After those four cases, the octets become bits (x 8) and are divided by the seconds, and the sample is rejected if the rate is above the link speed.
A synthetic day of counters: per-interface average and peak over 24 hours
Apply the same function to each consecutive pair of readings per interface and take the mean and maximum of the valid rates. The generator below makes synthetic 32-bit counter readings, so the run is reproducible: each interface has a typical rate (illustrative values of 40 and 300 Mbit/s), each minute's rate is that typical rate scaled by a random factor between 0.5 and 1.8, and the counter is advanced by the octets that rate carries in 60 s, wrapping at 2^32. The octet_rate function is the one defined above.
import random
WRAP32 = 2**32
LINK = 1_000_000_000
random.seed(7)
interfaces = {1: 40e6, 2: 300e6} # ifindex -> typical rate in bit/s
for ifindex, typical in interfaces.items():
counter = random.randrange(WRAP32) # arbitrary starting value
readings = [counter]
for _ in range(1440): # 1,440 one-minute intervals = 24 h
rate = typical * random.uniform(0.5, 1.8) # this minute's average, bit/s
counter = (counter + int(rate * 60 / 8)) % WRAP32
readings.append(counter)
rates, skipped = [], 0
for prev, curr in zip(readings, readings[1:]):
bps, note = octet_rate(prev, curr, 60, LINK)
if bps is None:
skipped += 1
else:
rates.append(bps)
print(f"ifindex {ifindex}: readings={len(readings)} intervals={len(rates)} skipped={skipped} "
f"avg={sum(rates)/len(rates)/1e6:.1f} Mbit/s peak={max(rates)/1e6:.1f} Mbit/s")
Output (Python 3.12, executed in a container):
ifindex 1: readings=1441 intervals=1440 skipped=0 avg=45.3 Mbit/s peak=71.9 Mbit/s
ifindex 2: readings=1441 intervals=1440 skipped=0 avg=342.5 Mbit/s peak=539.3 Mbit/s
1,441 readings make 1,440 intervals, and none is skipped because every interval's true rate stays below the 572.6 Mbit/s at which a 60 s interval would carry a full 2^32 octets and become ambiguous. The average sits near 1.15 x the typical rate because the random factor averages (0.5 + 1.8) / 2 = 1.15. Interface 2 wraps its counter roughly every 100 s, so most intervals are wrapped ones and the function still recovers them.
Report the peak as the peak of 60 s averages: a shorter burst is invisible, and the 5-minute average is lower still.
Pitfalls
- Summing deltas across a reboot produces a huge false spike; the
sysUpTimeandifCounterDiscontinuityTimechecks exist for this. - Rejecting an impossible rate is better than guessing: a gap in a chart is honest, a spike to 120% is not.
- Do not average percentages across intervals of different lengths: sum octets and divide by total seconds.
How would you measure real bandwidth utilization and remaining headroom on a link? Compare what counters, flow data and active tests each tell you, and how you decide when to upgrade.
Sample Answer
Direct answer
Measure utilization from interface counters (the device's own byte totals, polled by SNMP, the Simple Network Management Protocol), explain it with flow data (records of who talked to whom, exported by the device), and test the headroom actively only when you need ground truth. Headroom is capacity minus the busy-period load, not minus the average. I would upgrade when the 95th percentile (p95) of 5-minute utilization, in the busier direction, reaches 70 percent and stays there for a sustained period, or when the link shows discards, or when the forecast says it will cross 70 percent within the provider's lead time.
What each method tells you
| Method | Tells you | Cannot tell you |
|---|---|---|
| Interface counters | Bits per second in and out, discards and errors per interface, over time and cheaply. Use the 64-bit counters (ifHCInOctets, ifHCOutOctets; an octet is a byte) and ifHighSpeed (the interface speed in units of 1,000,000 bit/s) for capacity, because RFC 2863 requires 64-bit octet counters on any interface faster than 20 Mbit/s (and 64-bit packet counters too from 650 Mbit/s up), and a 32-bit octet counter at 1 Gbit/s wraps, rolling over to zero, in about 34 seconds | Who is using the link or why. Polled every 5 minutes it hides short bursts |
| Flow data (NetFlow, IPFIX) | Top talkers, applications, source and destination, who caused a rise. RFC 7011 defines a flow as packets passing an observation point in an interval that share common properties | Exact totals when the export is sampled, and it does not show drops or queue depth. It explains load, it does not measure capacity |
| Active tests (iperf3 between two ends, or a provider speed test) | The achievable throughput right now, including limits no counter shows such as a policer (a rate limiter that drops traffic above a contracted rate) or a bad path | Normal behaviour: it is a point-in-time, intrusive test that competes with production traffic |
Two counters matter beyond bytes. ifOutDiscards counts packets dropped on output although no error was detected, typically because the queue was full; that is a discard, a deliberate drop that signals congestion. ifInErrors counts inbound packets that arrived damaged (for example failing their checksum) and could not be delivered; that is an error, which signals a bad cable, optic or duplex problem rather than congestion. A link at 60 percent average with rising output discards is already out of headroom.
From counters to a rate
Rate is the change in the octet counter over the poll interval, times 8, handling wrap. The script below is the whole method and runs as is:
import math
def rate_bps(octets_prev, octets_now, seconds, bits=64):
delta = octets_now - octets_prev
if delta < 0: # the counter wrapped (or the device reset it)
delta += 2**bits
return delta * 8 / seconds
print(rate_bps(10_000_000_000, 21_250_000_000, 300) / 1e6, "Mbit/s")
print(round(rate_bps(4_000_000_000, 2_000_000_000, 300, bits=32) / 1e6, 1), "Mbit/s")
print(round((10 * 1000 + 290 * 270) / 300, 1), "Mbit/s average")
s = sorted([210, 250, 240, 300, 310, 280, 650, 700, 720, 330,
290, 260, 270, 305, 295, 285, 275, 265, 255, 245])
k = math.ceil(0.95 * len(s))
print("p95 =", s[k - 1], "Mbit/s (rank", k, "of", len(s), ")", "max =", s[-1])
300.0 Mbit/s
61.2 Mbit/s
294.3 Mbit/s average
p95 = 700 Mbit/s (rank 19 of 20 ) max = 720
Line by line:
- First line: the counter rose by 21,250,000,000 - 10,000,000,000 = 11,250,000,000 bytes (11.25 GB) in 300 s, which is 11,250,000,000 x 8 / 300 = 300 Mbit/s.
- Second line: the counter went from 4,000,000,000 down to 2,000,000,000, so it wrapped between polls. Adding 2^32 to the negative difference recovers 2,294,967,296 bytes, or 61.2 Mbit/s. That is right only if it wrapped once. At 1 Gbit/s a 32-bit counter can wrap several times in 5 minutes and the formula cannot see the extra wraps, which is why you poll 64-bit counters.
- Third line: 10 s at 1,000 Mbit/s plus 290 s at 270 Mbit/s averages 294.3 Mbit/s over the 5 minutes, even though 10 seconds ran at line rate, so averages must not decide headroom.
- Last line: the nearest-rank 95th percentile sorts the twenty samples, then takes the one at position ceil(0.95 x 20) = 19, which is 700 Mbit/s, far above the typical 250 to 300 Mbit/s. This is why you size on percentiles and not on the mean.
Deciding when to upgrade
- Compute utilization per direction (links are full duplex, meaning they send and receive at the same time, so inbound and outbound are separate budgets) and take the busier one, as a percentage of
ifHighSpeed(in Mbit/s). On the 20-sample example 700 / 1000 = 70 percent. - Alert when p95 reaches 70 percent and holds over, say, a week of business days, or on any sustained output discards.
- Explain the peak with flow data: if one backup job causes it, reschedule it before buying bandwidth.
- Forecast: fit the last 90 days of weekly p95 and extrapolate. Order the upgrade when the forecast reaches 70 percent within the carrier or procurement lead time. For example (illustrative numbers), thirteen weekly p95 values of 44, 45, 47, 48, 50, 51, 52, 54, 55, 57, 58, 59 and 61 percent fit a straight line rising about 1.4 points per week. That line reaches 70 percent about 6.5 weeks after the last sample, so with an 8-week lead time you order now.
- Confirm with a short active test off-peak if counters and users disagree.
- For a redundant pair, keep each link under 50 percent at peak so one can carry both after a failure.
Pitfalls
Counters reset on reboot, so look for discontinuities (RFC 2863's ifCounterDiscontinuityTime). The speed field can be wrong on bundles or sub-rate circuits, so compare to the contracted rate. Flow sampling means totals are estimates. Active tests on a production link need a change record and a bandwidth cap.
Which interface-level metrics would you monitor on routers and switches to catch congestion, errors and degradation early? For each, explain what it tells you and how to read it.
Sample Answer
Direct answer
Monitor eight things per interface: state, traffic, speed, errors, discards, packet-type counts, optics and queue drops. The first six come from the standard IF-MIB (RFC 2863, the standard table of interface counters a device exposes over SNMP); optics and per-class queue drops are usually vendor-specific. If you can only start with a few, metrics 1 to 5 (state, octets, speed, errors and discards) already catch most congestion and cable trouble; the other three (packet-type counts, optics and queue drops) sharpen the diagnosis. Read every counter as a rate (the change between two polls divided by elapsed seconds) and every error counter as a ratio against packets, never as a raw total. Traffic tells you load, errors tell you physical trouble, discards tell you queue pressure or policy.
The eight metrics
| # | Metric (object) | What it tells you | How to read it |
|---|---|---|---|
| 1 | Operational and administrative state (ifOperStatus, ifAdminStatus), plus ifLastChange | Whether the interface is up. Admin up with oper down is a fault; admin down is deliberate. ifLastChange is the sysUpTime when it entered its current state | Alert on admin-up and oper-down for important links. Frequent changes in ifLastChange indicate a flapping link |
| 2 | Octet counters (ifHCInOctets, ifHCOutOctets) | Traffic volume, hence load | Rate = (delta x 8) / seconds, compared per direction. 64-bit counters because a 32-bit counter wraps in 34 s at 1 Gbit/s |
| 3 | Speed (ifHighSpeed) | The denominator for utilisation, in units of 1,000,000 bit/s | Utilisation = rate / (ifHighSpeed x 1,000,000). The older ifSpeed tops out at 4,294,967,295, so fast interfaces report a wrong value |
| 4 | Errors (ifInErrors, ifOutErrors) | Inbound packets with errors that prevented delivery; outbound packets that could not be sent because of errors | Ratio to packets. A clean link should show none; sustained growth typically means a bad cable, dirty or failing optic, or a speed or duplex mismatch (duplex mismatch: one end sends and receives at the same time, the other end takes turns, so they collide) |
| 5 | Discards (ifInDiscards, ifOutDiscards) | Packets dropped although nothing was wrong with them | Output discards typically mean the egress queue (the buffer where packets wait to leave the port) overflowed (congestion); input discards can mean policy or resource limits. Ratio to packets |
| 6 | Packet-type counters (ifHCInUcastPkts, ifHCInMulticastPkts, ifHCInBroadcastPkts, and the Out equivalents) | The mix of traffic. A sudden jump in broadcast or multicast share suggests a loop or a misbehaving host | Share of total packets and its change over time |
| 7 | Optical and environmental readings (ENTITY-SENSOR, RFC 3433, the standard table for sensor readings such as temperature; per-transceiver power from vendor MIBs) | A transceiver degrading before errors appear | Trend against the module's own thresholds, rather than a fixed value |
| 8 | Queue and buffer drops per class (vendor-specific MIBs, or streaming telemetry, where the device pushes counters every few seconds instead of waiting to be polled) | Which traffic class is being dropped, which ifOutDiscards cannot say | Compare the class that matters (voice, storage) with its configured share, the slice of the link that quality of service (QoS) rules reserve for it |
Reading the numbers
- Always a rate from two readings. Counters are cumulative; a poller subtracts the earlier reading from the later one and divides by the real elapsed time (use the poll timestamps, not the nominal interval).
- Handle wrap and reset. Per RFC 2863 a manager must discard a difference when
ifCounterDiscontinuityTimechanged between the two polls, in addition to checkingsysUpTimefor agent restart. - Direction matters. On a full-duplex link each direction has its own capacity. A link at 20% in and 95% out is congested outbound.
- Errors versus discards diagnose different layers. Errors point at layer 1 (the physical layer: cable, optic, signal quality); discards point at layers 2 to 3 (the frames and IP packets that the device itself receives and forwards, so queues and policy). If both rise on one port, check the physical layer first, because a faulty link also causes retransmissions that add load.
Worked example
An uplink with ifHighSpeed = 1000 over one 60 s interval: output octets +3,750,000,000; output packets +4,500,000 with +45,000 ifOutDiscards; input packets +3,000,000 with +120 ifInErrors.
- Output utilisation: 3,750,000,000 x 8 / 60 = 500 Mbit/s, 50% of 1,000 Mbit/s.
- Discards: 45,000 / 4,500,000 = 1.0%. Healthy utilisation with 1% discards means bursts overflow the queue, so look at queue drops (metric 8) and consider buffer or QoS changes.
- Errors: 120 / 3,000,000 = 0.004%. Small, but it matters if it keeps climbing.
- For contrast (illustrative numbers), the same link in a healthy minute carrying the same 4,500,000 output packets shows 0 discards, which is 0%, and 0 errors. The unhealthy signs are the discards ratio moving off zero and the error count growing from one poll to the next.
Pitfalls
- Averages hide bursts: a 60 s average of 50% can contain 100 ms bursts at line rate. Discards are the cheap signal that those bursts happened.
- Using 32-bit counters on fast links or
ifSpeedfor utilisation silently produces wrong charts. - Port channels (several physical links grouped into one logical link, also called a bundle): monitor member links as well as the bundle, since the bundle's average can hide one member that is saturated or erroring.
- Counters for a given
ifIndex(the integer a device uses to number each interface) can change after a reboot on some devices; map by interface name, not index alone (RFC 2863 notes ifIndex values may be reassigned whensysUpTimeresets).
Tell me about a time you diagnosed and resolved a network performance or latency problem. What did the monitoring show, who did you work with, and what changed afterward?
Sample Answer
Direct answer
A strong answer is one concrete incident told in four moves: the symptom and why it mattered, what the monitoring data showed (and what it could not show), the specific actions you took and who else you pulled in, and what changed permanently afterward, followed by a short reflection on what you would do differently. Keep it first-person and name the metrics you looked at. Below is the shape, then a worked example you can replace with your own facts.
Structure to follow
- Situation (2 sentences). Who was affected, what they saw, how it was reported (users, an alert, a customer ticket) and the business stake.
- What the monitoring showed. Name the signals: interface utilization, output drops and errors, queue depth, round-trip latency and jitter (how much the delay varies from packet to packet) from a synthetic probe (a scripted test that sends fake traffic on a schedule), flow records (per-conversation summaries exported by routers: who talked to whom and how many bytes). Say honestly what the monitoring missed at first; the interesting part of most stories is a blind spot.
- Your actions, and who you worked with. The order of your checks, the hypothesis you ruled out, and the people you needed (an application owner, the carrier, a server or virtualization team, a security colleague) and what you asked each of them for.
- Result and what changed. The fix, how you proved it worked (the same measurement before and after), and the lasting change: a new alert, a shorter polling interval on that link, a runbook, a design change.
- Reflection. One thing you would do differently, stated as a behaviour ("I would have asked for per-second counters on day one"), not a virtue.
Example story (replace with your own facts)
Situation. A branch office reported that video calls froze for a few seconds every morning around the same time. The branch had a 1 Gbit/s WAN (wide-area network) link and the dashboard, which polls interface counters every 5 minutes, showed the link comfortably under half full.
What the monitoring showed. The 5-minute utilization graph looked healthy, but the interface's output-discard counter (packets the router dropped because its queue was full) was increasing. That contradiction was the clue: a 5-minute average of 400 Mbit/s is equally consistent with a steady 400 Mbit/s and with 20 seconds at 1,000 Mbit/s plus 280 seconds at about 357 Mbit/s (400 x 300 = 120,000 megabits in the window; the burst accounts for 20 x 1,000 = 20,000, so 120,000 - 20,000 = 100,000 remain; 100,000 / 280 s = 357). Drops with a calm average mean the link was saturating in short bursts the polling interval averaged away.
Actions and collaboration. I temporarily polled the counters every few seconds on that one link, which confirmed full-rate bursts of tens of seconds. I read the flow records for those minutes and found a scheduled file-sync job from the branch file server dominating the bursts, so I worked with the server owner to confirm the job's schedule, and with the application team on which traffic class the video calls used. The queueing policy on the router (its rules for which packets wait, and which are dropped, when the link is full) gave video and file-sync the same default treatment, so video packets were dropped alongside bulk traffic.
Result. We moved the sync job off the morning peak, put video in a priority class (a queue the router serves first, with a guaranteed share of the link) with a bandwidth guarantee and bulk sync in a lower class, and confirmed the fix with the same measurements that exposed the problem: the output-discard counter stopped increasing during the sync window and the synthetic call probe's jitter stayed flat. Afterward I added an alert on output discards (not only on average utilization) for all WAN links and a one-minute polling interval for links with a history of bursts.
What I would do differently. I would have asked the branch to open a ticket with timestamps on day one; matching their timestamps to counters would have shortened the investigation.
Pitfalls
- Do not present a story with no data in it ("I looked at the monitoring and found the problem"). Name the counter or probe and what its value told you.
- Do not claim sole credit. A network performance problem almost always crosses a team boundary, and the interviewer is listening for how you worked across it.
- Do not invent precise before-and-after numbers you cannot defend. A qualitative outcome you can prove ("drops stopped, jitter flat for two weeks") beats a fabricated percentage.
Unlock Full Question Bank
Get access to all 21 Network Monitoring and Performance interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.