Network Monitoring and Performance Questions
Network telemetry and performance operations: SNMP polling and traps (including counter wraparound and SNMPv3 access), NetFlow, sFlow and IPFIX flow export, sampling and its accuracy, streaming telemetry (gNMI), and eBPF or packet-capture telemetry; interface-level metrics (utilization, errors, discards, queue depth, microbursts), active synthetic probing alongside passive counters, link-flap detection, baselining and anomaly detection on network signals including elephant-flow spotting, network SLIs and SLOs, alerting, alert-storm suppression and NOC dashboards, telemetry pipeline design, storage, retention, downsampling and cardinality for network data (including securing the collection path and handling bursty remote sites), BGP and link-state monitoring including prefix hijack and route-leak detection, and network capacity monitoring, percentile utilization and bandwidth headroom planning. Also covers measuring and tuning network-level latency, jitter, packet loss and throughput (bufferbloat, queueing, TCP tuning for long paths). Excludes the generic metrics, logs and traces stack and alert design, the layered fault-isolation method and packet-capture troubleshooting, TCP and protocol fundamentals, application and CDN latency engineering, cloud VPC design and security detection.
How would you monitor network connectivity for a microservices platform on Kubernetes? Which signals would you collect at pod and node level, and how would you detect policy or datapath problems?
Sample Answer
Direct answer
Monitor in layers and let each layer name a different failure: pod and node signals (is the datapath on this host healthy), cluster services (DNS and service routing), and active probes (does pod A reach service B across nodes, now). Detect policy problems by looking at the verdict the enforcing network plugin gave each flow, not by guessing from application errors. The Kubernetes NetworkPolicy resource is only enforced by the network plugin (the CNI, Container Network Interface), so the plugin's flow and drop visibility is your best source.
Signals to collect
| Layer | Signal | Source | What it tells you |
|---|---|---|---|
| Pod | Flow verdicts (forwarded vs dropped) and drop reasons | Hubble (Cilium's flow-observability component) metrics hubble_flows_processed_total (label verdict) and hubble_drop_total (labels reason, protocol), when Cilium is the plugin | Policy denials versus other drops |
| Pod | TCP flag counts, for example SYN (the first packet of a TCP handshake) with no matching reply | hubble_tcp_flags_total | Connections that start and never complete |
| Pod | DNS queries and responses by response code | hubble_dns_responses_total (label rcode) | Lookups failing inside the cluster |
| Cluster | DNS request duration and response codes | CoreDNS (the cluster's DNS server) coredns_dns_request_duration_seconds and coredns_dns_responses_total{rcode}, served on port 9153 at /metrics | Slow or failing resolution (a rising SERVFAIL rate, the DNS server failed to answer, or NXDOMAIN rate, the name does not exist) |
| Node | Interface counters, TCP statistics, connection-tracking (conntrack, the kernel's per-connection state table) stats | node_exporter collectors netdev, netstat, conntrack (enabled by default) | NIC errors and drops, retransmits, conntrack pressure |
| Path | Synthetic probes between pods, across nodes and to external dependencies | blackbox_exporter (probes HTTP, HTTPS, DNS, TCP, ICMP, gRPC) | End-to-end reachability and latency now, independent of real traffic |
| Service | Request success rate and latency from the application or mesh | App or proxy metrics | Whether users are affected |
Read the table top to bottom as the order of suspicion in the next section: pod verdicts and DNS first, then node counters, then path probes. The Hubble metrics are disabled by default and are switched on in the Helm values with hubble.metrics.enabled, for example {drop,flow,tcp,dns}.
Detecting policy and datapath problems
- Is it policy? Remember the semantics: with no NetworkPolicy selecting a pod, all inbound and outbound connections are allowed. Once a policy selects it, only what a policy allows is permitted. So a service that broke right after a policy rollout is a suspect. With Cilium, list dropped flows for the pod:
hubble observe --pod <pod> --verdict DROPPED. The output shows source and destination namespace/pod:port and the verdict, so you can see the exact flow that a missing rule blocked. The lines look like this (pod names are illustrative; the layout follows the Cilium documentation):
May 4 13:23:47.852: default/xwing:42818 <> default/deathstar-c74d84667-cx5kp:80 Policy denied DROPPED (TCP Flags: SYN)
May 4 13:23:48.854: default/xwing:42818 <> default/deathstar-c74d84667-cx5kp:80 Policy denied DROPPED (TCP Flags: SYN)
Reading it: a timestamp, then source namespace/pod:port, the arrow, destination namespace/pod:port, the reason (Policy denied), the verdict DROPPED, and the TCP flags. A SYN dropped by policy and then sent again about a second later is a client retrying a connection that a rule never lets through, so the fix is the rule, not the application.
2. Is the policy even enforced? A NetworkPolicy created on a cluster whose plugin does not support it has no effect. Test with a deliberate deny rule in a test namespace and confirm the probe fails.
3. Is it DNS? Check the CoreDNS rcode counters and the request duration histogram first, because many "network" outages are failed lookups.
4. Is it the node datapath? Look at node interface drops and errors, and conntrack usage against its limit; a full conntrack table drops new connections on that node only.
5. Is it cross-node? Compare probe results same-node versus cross-node. If only cross-node fails, suspect the overlay (the tunnel the plugin builds to carry pod traffic between nodes) or the underlay (the physical or cloud network beneath it): a tunnel adds header bytes, so a full-size packet may no longer fit the MTU (maximum transmission unit, the largest packet a link carries) and is dropped, while small requests still pass; a firewall between nodes causes the same cross-node-only symptom.
Alerts
Alert on symptoms with rate and ratio: drop rate by reason rising, DNS error ratio, probe success under its target, and conntrack above a set fraction of its maximum. Page on service-level symptoms; keep drop counters as the diagnostic dashboard.
Pitfalls
- Per-pod metric labels that grow without bound as pods churn.
- Watching only application error rates, which cannot separate DNS, policy and datapath causes.
- Trusting that a policy works because it applied without an error.
How would you measure latency, jitter, packet loss and throughput between two points in a cloud environment? Compare the tools you would use at different layers and what each can miss.
Sample Answer
Direct answer
Measure each property with the probe that sees it, from both ends, over enough time to be meaningful, and keep passive counters running next to the active tests. Latency is the round-trip time (RTT) from ping or a timed HTTP request. Jitter (variation in delay between packets) and packet loss come from a UDP test stream or a long run of probes. Throughput comes from a TCP test with iperf3 (a free, widely used throughput tester). No single tool is complete: each one sees one layer and misses something the next layer would show.
What each tool measures and what it misses
| Property | Tool and layer | What it can miss |
|---|---|---|
| Latency | ping (ICMP echo, layer 3, the network layer where IP addresses and routing live) | ICMP can be rate limited, deprioritized or blocked by security groups (the cloud's virtual firewall rules), and under equal-cost multipath (ECMP, where routers spread flows across parallel paths by hashing the 5-tuple of addresses, protocol and ports) a ping may take a different path than your application's flows |
| Latency | curl -w timings (TCP handshake plus first byte, layer 7, the application layer where HTTP lives) | Mixes network and server time: a slow application looks like a slow network |
| Jitter | iperf3 -u (UDP) receiver jitter, ping mdev (despite the name, the standard deviation of the RTTs, so a measure of their spread) | Averages hide bursts. Two formal definitions exist, with a numeric example under the table. RFC 3550 defines a smoothed estimator, J = J + ( |
| Loss | ping -c, iperf3 -u lost/total datagrams, mtr per hop | Few probes give a noisy figure (worked example below). mtr documents that routers may give ICMP echo lower priority, so loss shown at a middle hop that does not continue to the last hop is usually rate limiting, not real loss |
| Throughput | iperf3 TCP, then -P for parallel streams and -R for the reverse direction | One flow is bounded by window divided by RTT, so a single-stream test on a long path under-reports what many flows achieve. Providers also cap bandwidth per instance and per flow: AWS documents single-flow traffic limited to 5 Gbps unless instances share a cluster placement group (an option that packs instances physically close together; 10 Gbps), and burst credits (extra bandwidth smaller instance types may use only while a saved-up allowance lasts) that run out |
| Passive (all four) | ss -ti on the host (socket statistics: per-TCP-connection RTT, retransmit count and window size), provider counters such as AWS Elastic Network Adapter (ENA, the virtual network card) ethtool -S eth0 fields bw_in_allowance_exceeded, bw_out_allowance_exceeded, pps_allowance_exceeded | Passive data covers only real traffic that happened. Virtual-network drops caused by instance allowances are invisible to a ping inside the guest but show in those counters |
Worked jitter example. Four packets sent at 20 ms spacing arrive with one-way transit times of 40, 46, 41 and 45 ms. D, the change in transit between consecutive packets, is 6, 5 and 4 ms. The RFC 3550 estimator starts at J = 0 and goes J = 0 + (6 - 0)/16 = 0.375, then 0.375 + (5 - 0.375)/16 = 0.664, then 0.664 + (4 - 0.664)/16 = 0.873 ms. A de-jitter buffer (the receiver's short wait that smooths out uneven arrival for voice or video) needs the RFC 5481 view instead: delay minus the minimum delay gives 0, 6, 1 and 5 ms, so a buffer of about 6 ms absorbs all four packets.
Method rules that apply in any cloud: test both directions (paths and limits can differ), run from the same zone and across zones/regions so you can separate a local fault from a path fault, run long enough to include a busy period, and repeat probes on a schedule so you have a baseline rather than one number. One-way delay needs synchronized clocks on both ends, so use RTT unless you control time sync.
Worked example you can reproduce
This runs in any Linux container started with the NET_ADMIN capability (the permission to change network settings; iperf3, iproute2, iputils-ping and curl installed). tc netem (the Linux network emulator) delays packets on the loopback interface by 20 ms plus or minus 5 ms and drops 1 percent. Because every packet crosses loopback in each direction, expect an RTT near 40 ms and about 1.99 percent round-trip ping loss (1 - 0.99 x 0.99).
ip link set dev lo mtu 1500
tc qdisc add dev lo root netem delay 20ms 5ms loss 1%
iperf3 -s -D
python3 -m http.server 8080 --bind 127.0.0.1 &
ping -c 200 -i 0.2 127.0.0.1 | tail -3
iperf3 -c 127.0.0.1 -u -b 10M -t 20 | grep receiver
curl -s -o /dev/null -w 'connect=%{time_connect}s ttfb=%{time_starttransfer}s total=%{time_total}s\n' http://127.0.0.1:8080/
Output from one run (netem's randomness is not seeded, so your numbers will differ slightly):
200 packets transmitted, 194 received, 3% packet loss, time 40517ms
rtt min/avg/max/mdev = 33.127/44.944/65.456/7.373 ms
[ 5] 0.00-20.04 sec 23.6 MBytes 9.89 Mbits/sec 4.293 ms 150/17267 (0.87%) receiver
connect=0.043648s ttfb=0.093469s total=0.093511s
How to read it. The receiver's Lost/Total Datagrams column reads 150/17267: of the 17,267 datagrams it accounted for, 150 never arrived, 150 / 17,267 = 0.87 percent one way, with 4.293 ms of jitter. That is close to the 1 percent configured: 1 percent of 17,267 would be 173 lost, and the natural spread for that many packets is about 13 (sqrt(17,267 x 0.01 x 0.99)), so 150 is within two standard deviations of 173. The grep receiver keeps only the receiver's summary line, because the receiver is the end that counts what arrived. The ping shows 3 percent loss against an expected 1.99 percent: with only 200 probes the standard error is sqrt(0.0199 x 0.9801 / 200), about 0.99 percentage points, so 3 percent is one standard error high and is not evidence of a second problem. That is why a loss figure needs thousands of probes or a long window before you act on it. The curl connect time of 43.6 ms is one RTT (the TCP handshake needs a round trip), agreeing with the ping average of 44.9 ms, and the 93 ms first byte adds a second round trip for the request.
Trade-offs and pitfalls
- Reporting averages. Report percentiles (95th, 99th) for latency and the maximum for loss bursts, since users feel the tail.
- Active tests consume bandwidth and, across regions, can incur data-transfer charges. Run heavy iperf3 tests off-peak and short, keep light probes continuous.
- A clean ICMP result does not clear the application: test with the same protocol and port the application uses (
mtr -T -P 443sends TCP SYN probes to a chosen port). - Treat a throughput number as a ceiling for that test, not a promise: a TCP result tells you about the host's windows and the path together.
Inside a Kubernetes cluster, how would you measure per-flow network latency, spot microbursts and tell queueing delay from other causes, without adding heavy overhead across thousands of containers?
Sample Answer
Direct answer
Measure on each node, inside the kernel, with eBPF (small programs the kernel verifies and then runs at hook points, so no packet copying and no sidecar in every pod). Read the round-trip time (RTT) the kernel already tracks for every TCP connection, fold it into histograms in the kernel, and export only per-node, per-workload-pair summaries. Catch microbursts with byte counters kept in 1 ms bins on the node's network interface. Separate queueing from other causes by comparing three signals: kernel RTT, application request latency, and a same-node control path. Queueing is the explanation only when RTT rises in step with the burst bins and recovers the moment they end.
Background in plain language
End-to-end delay is propagation (distance), serialization (putting bits on the wire), processing (host and switch CPUs) and queueing (waiting in a buffer because the output port is busy). Of these four, queueing is the delay that grows directly with load on a link (host processing can also grow when a CPU is saturated, which is why the ordered checks below rule it out). A microburst is traffic that arrives faster than the output port can drain for a few milliseconds, fills the buffer, then disappears, so a counter polled every 10 to 60 seconds shows a near-idle link.
Worked example, computed with python3: four senders each push 10 Gb/s at one 10 Gb/s switch port for 1 ms. Ingress is 4 x 10 Gb/s x 1 ms = 5.0 MB; the port drains 1.25 MB in that millisecond, so 3.75 MB queue up. If the port's buffer is 4 MB, all 3.75 MB are held and the last byte waits 3.75 MB / (10 Gb/s) = 3.0 ms, with no loss. If the buffer is only 2 MB, the queue is capped at 2 MB, so the worst wait is 2 MB / (10 Gb/s) = 1.6 ms and 1.75 MB are dropped. Averaged over 10 seconds the same burst is 5 MB / 10 s = 4 Mb/s, which is 0.04% of the link. Averaged over the 1 ms bin it is 400% of line rate. Same event, 10,000 times apart, purely from the measurement interval.
Per-flow latency without heavy overhead
- One agent per node, not per container. A DaemonSet (Kubernetes runs one copy on every node) loads the eBPF programs. Thousands of containers share one agent per node, so overhead scales with nodes, not containers.
- Use RTT the kernel already computes.
ss -iprints the per-connectionrttandrttvar(mean deviation) in milliseconds, per the ss man page. The bcctcprtttool does the same job as a histogram: it traces TCP RTT, can break histograms down by local or remote address (-b,-B), and is described by its authors as a way to tell whether latency comes from a user process or the physical network. A production agent applies that idea continuously.
A real line fromss -ti dst 127.0.0.1:9000on a Linux test host (a loopback connection, so the values are tiny):
cubic wscale:10,10 rto:201 rtt:0.016/0.008 mss:32741 pmtu:65535 rcvmss:536 advmss:65483 cwnd:10 bytes_acked:1 segs_out:2 segs_in:1 send 164Gbps lastsnd:1010 lastrcv:1010 lastack:1010 pacing_rate 327Gbps delivered:1 app_limited rcv_space:65495 rcv_ssthresh:65495 minrtt:0.016 snd_wnd:65483
Read the fields that matter here: cubic is the congestion-control algorithm; rtt:0.016/0.008 is the smoothed RTT (0.016 ms) and its mean deviation, rttvar (0.008 ms), both in milliseconds per the ss man page; minrtt is the lowest RTT seen, a baseline; cwnd:10 is the congestion window in segments; wscale:10,10 the send and receive window-scale factors. On a real pod-to-pod path across nodes the rtt value is larger and rttvar shows how much it wobbles.
- Aggregate before exporting. Key a log-scale histogram by (source workload, destination workload, destination node), not by pod, so pod churn does not explode label cardinality (the number of distinct label combinations, each of which is a separate stored time series). Map IPs to workloads using the Kubernetes API or the CNI (container network interface plugin) data. Export every 15 seconds.
- Raw per-flow records only when interesting. Emit a full record for 1 in 100 connections plus any connection whose RTT exceeds a threshold. Hubble (Cilium's flow observability layer) is useful for flow verdicts and drops, but its documented
tcpmetrics handler counts TCP flags, so it does not give RTT; the eBPF layer does. - Non-TCP traffic has no kernel RTT. Use a small active probe mesh: each node probes a rotating sample of peer nodes once a second, which also provides the baseline.
Spotting microbursts
- Host side: an eBPF program on the node's egress interface adds packet bytes to a counter selected by the kernel's monotonic clock divided into 1 ms bins. Every interval it exports the maximum bin and the number of bins above, say, 80% of line rate. That is a few numbers per interface per interval, not a packet stream.
- Switch side: the queue forms at the switch's output port, so queue depth or buffer high-watermark telemetry from the switch shows the effect, if the platform exports it. The host bins show the cause (which nodes transmit in lockstep, as in a shuffle or all-to-one read).
- Align both on timestamps (clocks synchronized with NTP or PTP, network time protocols, PTP being the more precise) before correlating.
Telling queueing from other causes (ordered checks)
- Is it the network at all? Compare kernel RTT with application request latency. Kernel RTT flat while request latency rises points at the application, CPU throttling, or garbage collection.
- Run the control path. RTT between two pods on the same node never crosses the NIC or a switch. If that is also slow, the cause is host CPU, softirq load (the kernel's deferred packet-processing work, which runs on a CPU) or connection tracking (the kernel's table of connections that firewall and NAT rules depend on), not the fabric.
- Check load correlation. Queueing has a signature: RTT jumps by queue depth divided by line rate (3 ms in the 4 MB-buffer case above, at most 1.6 ms in the 2 MB-buffer case), affects every flow sharing that egress port at once, and returns to baseline within milliseconds of the burst ending. A cause like a bad cable or a CPU stall does not track the burst bins.
- Look at loss. Retransmits that appear only when the burst exceeds the buffer (in the 2 MB case, the 1.75 MB overflow) fit queueing; steady retransmits with flat bins fit a faulty link.
- Group by path. Only flows hashed onto one uplink (ECMP, equal-cost multipath) affected means a hot port, not a cluster-wide problem.
Worked trace (illustrative numbers, built on the 4 MB-buffer case above, where a burst adds 3 ms of queueing). Users report request latency rising from 8 ms to 12 ms. Check 1: kernel RTT for the same flows went from 0.4 ms to 3.4 ms, so about 3 ms of the 4 ms rise is in the network path, not the application. Check 2: the same-node pod pair stays at 0.05 ms, so host CPU and connection tracking are fine. Check 3: the 3 ms jumps line up with the 1 ms bins above 80% of line rate and fall back to 0.4 ms within a few milliseconds of each burst. Check 4: retransmits appear only in the few largest bursts, the ones that overflow the buffer. Check 5: only flows hashed onto uplink B are affected. Conclusion: queueing at one egress port, so spread the lockstep senders or add buffer or capacity there.
Trade-offs and pitfalls
- Kernel RTT is a smoothed per-connection estimate: it hides a single 3 ms spike on a long connection. That is why the 1 ms bins exist.
- eBPF needs a privileged agent and a kernel with the needed hooks; canary it on a few nodes, and put a hard cap on map memory and export rate.
- Do not mirror packets or inject sidecars cluster-wide for this: cost scales with total traffic or pod count, which is the overhead the question rules out.
- Recommendation: node-level eBPF histograms plus 1 ms bin counters as the always-on layer. It would change to targeted packet capture only on the two or three nodes the histograms implicate.
That is every published Network Monitoring and Performance question for Cloud Engineer so far. Browse the other topics in this category, or practice this one interactively.