Network Monitoring and Performance Questions

Network telemetry and performance operations: SNMP polling and traps (including counter wraparound and SNMPv3 access), NetFlow, sFlow and IPFIX flow export, sampling and its accuracy, streaming telemetry (gNMI), and eBPF or packet-capture telemetry; interface-level metrics (utilization, errors, discards, queue depth, microbursts), active synthetic probing alongside passive counters, link-flap detection, baselining and anomaly detection on network signals including elephant-flow spotting, network SLIs and SLOs, alerting, alert-storm suppression and NOC dashboards, telemetry pipeline design, storage, retention, downsampling and cardinality for network data (including securing the collection path and handling bursty remote sites), BGP and link-state monitoring including prefix hijack and route-leak detection, and network capacity monitoring, percentile utilization and bandwidth headroom planning. Also covers measuring and tuning network-level latency, jitter, packet loss and throughput (bufferbloat, queueing, TCP tuning for long paths). Excludes the generic metrics, logs and traces stack and alert design, the layered fault-isolation method and packet-capture troubleshooting, TCP and protocol fundamentals, application and CDN latency engineering, cloud VPC design and security detection.

MediumSystem Design
54 practiced

Design a pipeline to collect, enrich and analyze network flow records arriving at hundreds of thousands per second, including flow logs from cloud environments, with real-time dashboards and hourly rollups. Cover collectors, buffering, processing, storage and failure behavior.

MediumTechnical
47 practiced

Pull-based metrics collection is popular, but many network devices do not speak it and sit behind NAT or firewalls. How would you monitor them with a pull-based system, and when would pull be the wrong fit?

MediumTechnical
50 practiced

Design the metric names and labels for interface metrics across 1,000 network devices so operators can query by device role, region and interface type without exploding the number of series. What would you refuse to label by?

MediumTechnical
29 practiced

How would you monitor network connectivity for a microservices platform on Kubernetes? Which signals would you collect at pod and node level, and how would you detect policy or datapath problems?

MediumTechnical
47 practiced

Write a Python script using only the standard library that measures TCP connect latency to a list of endpoints concurrently and reports p50, p95 and p99 of successful connects. Handle timeouts and failures sensibly.

Unlock Full Question Bank

Get access to all 8 Network Monitoring and Performance interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.