Network Monitoring and Performance Questions

Network telemetry and performance operations: SNMP polling and traps (including counter wraparound and SNMPv3 access), NetFlow, sFlow and IPFIX flow export, sampling and its accuracy, streaming telemetry (gNMI), and eBPF or packet-capture telemetry; interface-level metrics (utilization, errors, discards, queue depth, microbursts), active synthetic probing alongside passive counters, link-flap detection, baselining and anomaly detection on network signals including elephant-flow spotting, network SLIs and SLOs, alerting, alert-storm suppression and NOC dashboards, telemetry pipeline design, storage, retention, downsampling and cardinality for network data (including securing the collection path and handling bursty remote sites), BGP and link-state monitoring including prefix hijack and route-leak detection, and network capacity monitoring, percentile utilization and bandwidth headroom planning. Also covers measuring and tuning network-level latency, jitter, packet loss and throughput (bufferbloat, queueing, TCP tuning for long paths). Excludes the generic metrics, logs and traces stack and alert design, the layered fault-isolation method and packet-capture troubleshooting, TCP and protocol fundamentals, application and CDN latency engineering, cloud VPC design and security detection.

EasyTechnical
37 practiced

How does SNMP work for network monitoring? Walk through how a poller gets interface and device health data, how polling differs from traps, and what changes between the older and newer protocol versions.

EasyTechnical
31 practiced

Which interface-level metrics would you monitor on routers and switches to catch congestion, errors and degradation early? For each, explain what it tells you and how to read it.

EasyTechnical
29 practiced

You want to detect congestion, rising packet loss, jitter and device failure before users complain, in a medium enterprise network. What would you collect, how often, and where would you set initial alert thresholds?

MediumTechnical
46 practiced

A cross-continental replication link has 300 ms RTT and 1 Gbps of capacity, but transfers reach only a fraction of it. Work out what limits throughput, what you would change on the hosts, and how you would validate each change.

EasyTechnical
34 practiced

What is the difference between active health checks and passive telemetry for network monitoring? Give examples of data from each and say when you rely on which.

That is every published Network Monitoring and Performance question for Systems Administrator so far. Browse the other topics in this category, or practice this one interactively.