Network Monitoring and Performance Questions
Network telemetry and performance operations: SNMP polling and traps (including counter wraparound and SNMPv3 access), NetFlow, sFlow and IPFIX flow export, sampling and its accuracy, streaming telemetry (gNMI), and eBPF or packet-capture telemetry; interface-level metrics (utilization, errors, discards, queue depth, microbursts), active synthetic probing alongside passive counters, link-flap detection, baselining and anomaly detection on network signals including elephant-flow spotting, network SLIs and SLOs, alerting, alert-storm suppression and NOC dashboards, telemetry pipeline design, storage, retention, downsampling and cardinality for network data (including securing the collection path and handling bursty remote sites), BGP and link-state monitoring including prefix hijack and route-leak detection, and network capacity monitoring, percentile utilization and bandwidth headroom planning. Also covers measuring and tuning network-level latency, jitter, packet loss and throughput (bufferbloat, queueing, TCP tuning for long paths). Excludes the generic metrics, logs and traces stack and alert design, the layered fault-isolation method and packet-capture troubleshooting, TCP and protocol fundamentals, application and CDN latency engineering, cloud VPC design and security detection.
Design a pipeline to collect, enrich and analyze network flow records arriving at hundreds of thousands per second, including flow logs from cloud environments, with real-time dashboards and hourly rollups. Cover collectors, buffering, processing, storage and failure behavior.
Sample Answer
Direct answer
Run a horizontally scaled set of stateless-ish collectors that decode flows and publish them to a durable streaming log (a replicated queue such as Kafka, split into partitions: independent ordered lanes so many workers can read in parallel). Behind that log, an enrichment stage adds site, application, autonomous system number (ASN, the number that identifies the network owning an address range) and sampling-corrected bytes (routers often export one packet in every N, so bytes are multiplied by N to estimate the real total), and a columnar analytical store holds a few days of raw records plus rollups: one-minute rollups for live dashboards and hourly rollups for long retention. The streaming log is the shock absorber: when the store, the enricher or a collector falls behind, data waits in the log instead of being dropped, and the one place data can still be lost silently (UDP between the router and the collector) is watched with sequence-number gap detection.
Vocabulary bridge
Network readers: Kafka and ClickHouse are the data tools here. A topic is a named log; a consumer is a program reading it, and consumer lag is how far behind the newest record that reader is. A dead-letter topic is a side log for records that cannot be processed. A TTL (time-to-live) deletes rows after a set age. A tumbling window is a fixed, non-overlapping time bucket (every minute, for example); event time is the timestamp inside the record (when the flow ended) as opposed to when the pipeline saw it; a watermark is the pipeline's rule for how long to wait for late records before closing a window. Backpressure means a slow reader making the sender slow down. UDP has none, which is why the log sits behind the collectors. Protobuf is a compact binary record format; JSON is the readable alternative.
Data engineers: an exporter is a router or switch that sends flow records. Before a collector can decode data records it needs the exporter's template (a description of the fields and their order), sent periodically, which is why collector restarts and load balancing matter. An observation domain is an exporter's identifier scope, and a sequence number counts what it has sent, so a gap means loss.
Assumptions and sizing
These figures are design assumptions to be replaced by measurements from your exporters. The point is to show the method.
| Quantity | 300,000 records/s | 1,000,000 records/s |
|---|---|---|
| Wire size at 60 bytes per exported record | 18 MB/s (144 Mbit/s) | 60 MB/s (480 Mbit/s) |
| Decoded and enriched record at 200 bytes | 60 MB/s | 200 MB/s |
| Partitions at a 10 MB/s design target per partition, doubled for headroom | 12 | 40 |
| Log retention of 6 hours of enriched data | 60 MB/s x 21,600 s = 1.30 TB | 200 MB/s x 21,600 s = 4.32 TB |
| Raw records kept 3 days | 77.8 billion rows | 259.2 billion rows |
Raw retention beyond a few days is rarely worth the storage, so the long-retention view is a rollup. A rollup keyed by source site, destination site and application cannot exceed 200 x 200 x 30 = 1.2 million rows per hour for 200 sites and 30 applications. Against 300,000 x 3,600 = 1.08 billion raw records in the same hour, that is a reduction of at least 900 times.
Architecture
The diagram reads left to right: routers send UDP flow packets to collectors, cloud flow logs arrive as files through a queue and a loader, both feed the raw log, enrichment writes an enriched log, and that log feeds the raw store, one-minute rollups for dashboards and hourly rollups for retention.
flowchart LR
R[Routers and switches<br/>IPFIX, NetFlow, sFlow] -->|UDP| C[Collector pool<br/>decode and tag]
V[Cloud flow logs<br/>object storage] --> Q[Notification queue] --> L[Log loader<br/>parse and normalize]
C --> K[(Streaming log<br/>flows-raw)]
L --> K
K --> E[Enrichment<br/>site, ASN, app, sampling]
E --> K2[(Streaming log<br/>flows-enriched)]
K2 --> S[(Columnar store<br/>3 day raw)]
S --> H[Hourly rollups<br/>long retention]
K2 --> M[1 minute rollups<br/>live dashboards]
K -.->|poison records| D[(Dead-letter topic)]
Collectors
- What they run. A flow collector that speaks sFlow v5, NetFlow v5 and v9, and IPFIX (IP Flow Information Export, the standard flow format). goflow2 is one open-source example: it decodes those protocols, takes the sampling rate from the Option Data Set, and can write protobuf or JSON to Kafka. Many parallel UDP sockets per host are configured with its listen option.
- Template state is the catch. NetFlow v9 and IPFIX send a template that says how to read the data records. In IPFIX, template IDs are unique only within a transport session and observation domain (RFC 7011), and exporters over UDP re-send templates periodically. A collector that restarts or a load balancer that sprays one exporter's packets across collectors cannot decode records until a template reaches that instance, which can take minutes. Hash by exporter source IP so each exporter lands on one collector, and add a standby that takes over the same exporter addresses.
- Resource cost. Measure, do not guess: replay captured export traffic at 1.5 times the target rate against one collector and read the kernel's UDP receive-buffer error counter (RcvbufErrors in /proc/net/snmp on Linux) plus CPU. Size the pool with clear headroom below the rate at which that counter starts to rise, and raise the socket receive buffer before adding hosts.
Buffering
Decode at the collector and publish to a topic keyed by exporter, so one router's records stay in order. Keying has a cost: all of one exporter's records land in a single partition, so a very busy exporter becomes a hot partition and delays the other exporters hashed to it. At the 200-byte decoded size and the 10 MB/s per-partition target in the sizing table, one partition carries about 50,000 records per second (10,000,000 / 200). If any single exporter can exceed that, key by exporter plus a hash of the flow key (source, destination, ports, protocol) so its records spread across partitions; the rollups sum by time and site, so they do not need per-exporter ordering. Replicate the log (three copies is a typical choice; the 1.30 TB and 4.32 TB figures above are one copy, so three replicas need about 3.9 TB and 13 TB of disk) and retain at least as long as the longest outage you want to ride out without loss: 6 hours in the sizing above. Consumer lag per stage is the main health signal.
Processing and enrichment
- Normalize to one schema across router flows and cloud flow logs: time, source, destination, ports, protocol, bytes, packets, direction, exporter or resource ID, sampling rate.
- Scale for sampling at this stage: store raw bytes and sampling-corrected bytes (bytes x rate) side by side.
- Enrich by joining against reference data held in memory and refreshed every few minutes: IP to site, application and owner (from the IP address management system and configuration database), interface index to interface name (the INPUT_SNMP field is an index, not a name), and ASN where the exporter does not send one.
- Bucket for rollups by flow end time. Long flows are exported at the active timeout, so a record's bytes are attributed to the minute it ends in. Say so on the dashboard: a one-minute chart smears a long transfer.
- Poison records (undecodable, impossible timestamps) go to a dead-letter topic with a counter, never block a partition.
Storage and rollups
Raw records live in a columnar analytical store (ClickHouse in the example below) with a 3-day time-to-live (TTL), and live dashboards read one-minute rollups fed from the enriched topic, so panels never scan raw rows. A tested ClickHouse pattern for the hourly rollup is a materialized view that writes aggregate states into an AggregatingMergeTree table:
CREATE TABLE flows_raw
(
ts DateTime,
exporter LowCardinality(String),
src_site LowCardinality(String),
dst_site LowCardinality(String),
app LowCardinality(String),
bytes UInt64,
packets UInt64,
sampling_rate UInt32
)
ENGINE = MergeTree
PARTITION BY toDate(ts)
ORDER BY (exporter, ts)
TTL ts + INTERVAL 3 DAY;
CREATE TABLE flows_hourly
(
hour DateTime,
src_site LowCardinality(String),
dst_site LowCardinality(String),
app LowCardinality(String),
est_bytes AggregateFunction(sum, UInt64),
est_packets AggregateFunction(sum, UInt64)
)
ENGINE = AggregatingMergeTree
ORDER BY (hour, src_site, dst_site, app);
CREATE MATERIALIZED VIEW flows_hourly_mv TO flows_hourly AS
SELECT
toStartOfHour(ts) AS hour,
src_site, dst_site, app,
sumState(bytes * sampling_rate) AS est_bytes,
sumState(packets * sampling_rate) AS est_packets
FROM flows_raw
GROUP BY hour, src_site, dst_site, app;
How the statements work: flows_raw holds every record and its TTL line deletes rows after 3 days. flows_hourly stores AggregateFunction(sum, UInt64) columns, which hold a partial running sum (a state) rather than a number, so partial sums from many inserts can be combined correctly later; AggregatingMergeTree merges rows with the same ORDER BY key into one row in the background. The materialized view is a trigger that runs its SELECT on every batch inserted into flows_raw, multiplies bytes by the sampling rate, and writes sumState(...) into flows_hourly. Reading it back uses the matching merge functions (sumMerge(est_bytes)), which combine the stored states into the final number; a plain sum would not work on state columns. Inserting three rows (1,500 and 3,000 bytes at sampling rate 1000 for dc1 to branch-014 in the 10:00 hour, plus 1,200 bytes at rate 100 for another pair in the 11:00 hour) returns 4,500,000 estimated bytes and 120,000 estimated bytes for the two hours, as expected from (1,500 + 3,000) x 1000 and 1,200 x 100.
Because this view adds on every insert, replaying records from the log would double count. When a replay is needed, delete the affected hours from the rollup and recompute them from raw records, rather than re-feeding the view.
Cloud flow logs
Cloud flow logs (AWS VPC Flow Logs as the example) are batched files, not UDP packets: publish them to object storage, send a bucket notification to a queue, and have a loader parse each file into the same normalized schema. Four differences matter:
- Latency. AWS states flow logs "do not capture real-time log streams", so cloud data arrives minutes behind router flows. Because the hourly materialized view adds each row to its hour as the row is inserted, late cloud rows land in the correct hour with no recompute; keep the live dashboard labelled as minutes behind.
- Per-interface logging doubles traffic. Logs are produced for each network interface, so one flow between two instances appears at both ends. Pick one side (typically the source interface, or the flow direction field) before summing.
- Address fields. Through a network address translation (NAT) gateway or transit gateway the plain source and destination fields are the gateway's addresses; the pkt-srcaddr and pkt-dstaddr fields hold the original packet addresses. Enrich on the packet addresses.
- Also excluded from the logs: DHCP, ARP, metadata and DNS-server traffic, so some "missing" flows are by design.
Failure behaviour
| Failure | What happens | Detection and response |
|---|---|---|
| Collector host dies | UDP packets sent to it are lost until failover | Gaps in the IPFIX sequence number per exporter and observation domain (RFC 7011: it counts data records and lets a collector see missed ones), alert on gap rate |
| Collector overloaded | Kernel drops UDP before the process reads it | RcvbufErrors rising, collector CPU; add collectors, grow receive buffer |
| Streaming log broker lost | Replicas keep serving; producers retry | Under-replicated partition alert, producer error rate |
| Enricher or store slow or down | Consumer lag grows, data waits in the log | Lag alert; log retention (6 hours here) is the time budget to fix |
| Poison record | One record fails decode | Dead-letter topic with counter, pipeline keeps moving |
| Exporter template lost | Records undecodable until the next template | Count of "no template" decodes per exporter; hash-to-one-collector rule reduces it |
| Late cloud data | Rows arrive after their hour has ended | The hourly view adds them to the right hour on insert, so nothing is recomputed; in the windowed 1,000,000 records per second design, rows later than the watermark go to the correction path |
Scaling to 1,000,000 records per second
At this rate writing every record is the cost driver (259.2 billion raw rows over 3 days). Add a stream-processing stage between the enriched topic and the store that aggregates into one-minute tumbling windows keyed by (exporter, source site, destination site, application) and writes only the aggregates. The saving depends on how many distinct keys a minute really contains, which you measure from your own traffic. Suppose (illustratively) a minute has 500,000 distinct keys: the writes fall from 1,000,000 rows per second to 500,000 / 60 = 8,333 rows per second, a 120 times reduction. With 100,000 distinct keys it would be 600 times, and with 50,000, 1,200 times, while raw records go to cheap object storage in a columnar file format for 3 days. Use event time (the flow end time, as in the bucketing step above) with a watermark that covers how late a record can arrive after that time: a flow that goes quiet is exported only after its idle timeout has passed, the exporter adds a cache-scan and queueing delay, and cloud flow logs arrive minutes behind. Size the watermark from the measured lag per source, and send anything later to the correction path (recompute the affected hours from raw records). A long flow exported at the active timeout carries an end time close to its export time, so the active timeout does not set the watermark.
Trade-offs
- Decoding at the collector keeps the log small and schema-stable but couples collector upgrades to protocol support. Publishing raw packets and decoding downstream decouples them at several times the log volume.
- A coarser rollup key cuts cost but removes the ability to answer new questions after the fact. Keep raw records long enough to cover the investigation window you actually use.
Pull-based metrics collection is popular, but many network devices do not speak it and sit behind NAT or firewalls. How would you monitor them with a pull-based system, and when would pull be the wrong fit?
Sample Answer
A pull-based system (Prometheus is the common example) has the monitoring server fetch metrics over HTTP from each target. Each fetch is called a scrape: one HTTP request to the target and a read of the current metrics in the reply. Most network devices speak SNMP (Simple Network Management Protocol), gNMI (a streaming management protocol) or only send flows and syslog, so you put a translator, called an exporter, between them and the server.
Making non-Prometheus devices scrapable
- SNMP devices: run
snmp_exporter. Prometheus scrapes the exporter at/snmp?target=<device>withauthandmoduleparameters, and the exporter does the SNMP polling. One instance can serve thousands of devices. Its documentation notes community strings are sent unencrypted and are not secrets in SNMP v1 and v2c, so use SNMP v3 where the device supports it. The relabel pattern that turns each device address into a target parameter is in the snippet below. - gNMI devices: run
gnmicwith its Prometheus output (type: prometheus), which exposes a scrape endpoint built from gNMI updates, cached for a configurable expiration. - Reachability and path quality: run
blackbox_exporter, which probes over HTTP, HTTPS, DNS, TCP, ICMP and gRPC via/probe?target=...&module=.... On Linux, ICMP (ping) probes need one of three things: theCAP_NET_RAWcapability (permission to build raw network packets), membership of a group listed in the kernel settingnet.ipv4.ping_group_range(which allows unprivileged pings), or root. This matters only when the exporter runs as a non-root user.
metrics_path: /snmp
params: {auth: [public_v2], module: [if_mib]}
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: device
- target_label: __address__
replacement: snmp-exporter.mgmt.example:9116
Reading it line by line (the rules run top to bottom):
metrics_path: /snmpmakes Prometheus request the exporter's/snmppath instead of the default/metrics.paramsbecome the query string:?auth=public_v2&module=if_mib.- Rule 1 copies
__address__(the device address from the job'sstatic_configstargetslist, not shown) into__param_target, which becomes&target=<device>in the URL. This is how the exporter learns which device to poll. - Rule 2 copies the same value into an ordinary label,
device, so stored series say which device they describe. Labels starting with__are discarded after relabeling, which is why this copy is needed. - Rule 3 has no source label; it just overwrites
__address__with the exporter's host and port. This is the line that sends the HTTP request to the exporter instead of to the device.
On a test run with targets leaf-12.example and leaf-13.example (illustrative names), Prometheus's targets API showed it fetching http://snmp-exporter.mgmt.example:9116/snmp?auth=public_v2&module=if_mib&target=leaf-12.example, and the series carried device="leaf-12.example".
(Config validated with promtool check config; the public_v2 auth and if_mib module are in the exporter's shipped example snmp.yml. Use an SNMP v3 auth profile in production.)
Devices behind NAT or firewalls. The server needs a route to the exporter, and the exporter needs a route to the device. So:
- Place the exporter, and a Prometheus (or an agent-mode Prometheus that only scrapes and forwards) inside the protected segment, close to the devices, and send data out through one outbound HTTPS remote-write connection (remote-write is Prometheus's way of streaming samples to another store; agent mode is a Prometheus mode that only scrapes and forwards, with no local querying). This opens one egress rule instead of an inbound hole per device.
- Where the server must stay outside, PushProx lets Prometheus traverse a firewall or NAT (named in the Prometheus pushing guide): a PushProx client inside the protected network keeps an outbound connection to a proxy, and Prometheus asks the proxy, which relays each scrape through that connection, so no inbound hole is needed.
- Do not use the Pushgateway for this. The Prometheus guidance is that its only valid use is the outcome of service-level batch jobs, because it becomes a single point of failure, removes the
uphealth signal (the metric Prometheus records for every target: 1 when the last scrape succeeded, 0 when it failed), and keeps pushed series forever until deleted.
When pull is the wrong fit
- Event data: SNMP traps, syslog and flow export (NetFlow or IPFIX) are push by design. Run collectors for them; do not convert events to polled gauges.
- Fast detection: a scrape every 30 seconds gives a detection floor of up to 30 seconds. For sub-second link or session events use device-originated streaming (gNMI on-change, where the device sends an update the moment a value changes instead of waiting to be asked) or BFD (Bidirectional Forwarding Detection, defined in RFC 5880 for low-latency detection: two neighbours exchange tiny hello packets and declare the path down after a set number are missed, so detection time is the interval times that count). BFD is the tool when a link or neighbour failure must be noticed faster than any polling interval.
- Short-lived sources that start and finish between scrapes cannot be pulled.
- Very large per-device tables (full BGP tables) are costly to walk repeatedly over SNMP; subscribe via streaming instead.
- Targets you cannot reach even through a local collector, because there is nowhere to place one.
In short: pull for state and counters you can reach through a collector, push or stream for events and speed.
Design the metric names and labels for interface metrics across 1,000 network devices so operators can query by device role, region and interface type without exploding the number of series. What would you refuse to label by?
Sample Answer
The rule behind everything below: each distinct combination of label values creates a separate time series (one metric name plus one exact set of label values, stored as a stream of timestamped numbers), so every label multiplies series count. The number of distinct values a label can take is its cardinality, and Prometheus guidance is not to use labels with high cardinality (many or unbounded values, such as user IDs or email addresses). It also says to put units in the metric name in base units, with _total on counters.
Metric names (base units, _total for counters, no label names in the name):
network_interface_receive_bytes_totalandnetwork_interface_transmit_bytes_totalnetwork_interface_receive_errors_totalandnetwork_interface_transmit_errors_totalnetwork_interface_receive_discards_totalandnetwork_interface_transmit_discards_totalnetwork_interface_oper_status(a gauge carrying the SNMP interface operational status value)
That is 7 series per interface, plus onenetwork_interface_infoseries. An info series is a metric whose value is always 1 and whose labels carry slow-changing text such as speed and description, so the text is stored once instead of on every series.
Labels on every series (all bounded):
| Label | Source | Values |
|---|---|---|
device | inventory hostname | about 1,000 |
interface | interface name (ifName) | about 48 per device |
device_role | inventory | a short fixed list (spine, leaf, access, edge, wan) |
region | inventory | a short fixed list |
if_class | derived from interface name | physical, lag, loopback, svi |
device_role and region come from the source of truth (the inventory system) and are attached as target labels (labels Prometheus adds to everything it scrapes from one device), not read from the device, so a renamed switch cannot change them silently. Operators then query by role and region without extra series, because those labels are constant per device:
sum by (region, device_role) (rate(network_interface_receive_bytes_total{if_class="physical"}[5m])) * 8
Reading it: a counter only ever grows, so rate(...[5m]) turns it into a per-second rate averaged over the last 5 minutes (a counter that rises by 7,500,000 bytes every minute gives 125,000 bytes per second). * 8 converts bytes to bits, so the result is bits per second. sum by (region, device_role) adds the per-interface rates into one total per region and role.
Series count. 1,000 devices x 48 interfaces x 8 series = 384,000, which stays 384,000 however many role and region combinations exist, because those labels do not split a device's series.
What I refuse to label by:
- VLAN ID on interface series. If 10% of ports (4,800 of 48,000) carry 100 VLANs each, those ports alone grow from 4,800 x 8 = 38,400 series to 4,800 x 100 x 8 = 3.84 million, and the fleet total becomes 3,840,000 + 345,600 (the other 43,200 ports, unchanged) = 4,185,600, about 10.9x the original 384,000.
- Client or peer IP address, or MAC address. If every port were an access port with 30 MACs, that is 384,000 x 30 = 11.52 million series, and they churn as clients move. That is flow-data territory, not metrics.
- Flow 5-tuple fields (the five values that identify a flow: source IP, destination IP, source port, destination port, protocol). Unbounded; send to the flow store.
- Free-text interface description or alias. Operators edit these, and every edit starts a new series and orphans the old history. Keep the text on the info series and join when you need it. A join matches each rate series to the one info series with the same
deviceandinterfaceand copies a label across:rate(network_interface_receive_bytes_total{device="leaf-12"}[5m]) * 8 * on (device, interface) group_left (ifAlias) network_interface_info. Multiplying by the info series' value of 1 leaves the number unchanged, andgroup_left (ifAlias)attaches the description to the result. (Evaluated withpromtool test ruleson illustrative series: a counter rising 7,500,000 bytes a minute returned 1,000,000 bit/s withifAlias="uplink-to-spine1"attached; the hostname and description are illustrative.) - Firmware version or serial on every series. Put it on a per-device info series.
Enforcement in the scrape config (validated with promtool check config):
- job_name: snmp_interfaces
sample_limit: 2000
metric_relabel_configs:
- source_labels: [__name__]
regex: ifHCInOctets
target_label: __name__
replacement: network_interface_receive_bytes_total
- regex: ifAlias|ifDescr
action: labeldrop
Reading the snippet line by line:
job_name: snmp_interfacesnames the scrape job (a scrape is one fetch of a device's metrics). This job carries the counters; the info series comes from a separate job without thelabeldroprule below, otherwise the text labels would be stripped from it too.metric_relabel_configsis a list of rewrite rules applied to every scraped sample after the scrape and before it is stored.- Rule 1:
source_labels: [__name__]is what to read (__name__is the hidden label that holds the metric name),regex: ifHCInOctetsis what to match (the 64-bit received-bytes counter from the IF-MIB, the standard SNMP interface table),target_label: __name__is what to write, andreplacementis the new value. With noactiongiven the rule does a replace, so the metric is renamed. - Rule 2:
action: labeldropdeletes every label whose name matchesifAlias|ifDescr(either name) from every sample, so editable free text can never become a label. sample_limit: 2000is the safety net in the next paragraph.
sample_limit makes Prometheus treat a scrape with more samples than the limit (counted after metric relabeling) as failed. 2,000 is about 5x the roughly 400 samples a 48-port device produces, so a device that suddenly exports a runaway table fails loudly instead of flooding storage. The exporter's own metric names (here the IF-MIB names) are renamed at ingestion, so dashboards depend on one scheme whichever protocol (SNMP or streaming telemetry) feeds it.
Review rule. Any new label needs a stated bound on its values and an owner; anything unbounded is a log or flow field instead.
How would you monitor network connectivity for a microservices platform on Kubernetes? Which signals would you collect at pod and node level, and how would you detect policy or datapath problems?
Sample Answer
Direct answer
Monitor in layers and let each layer name a different failure: pod and node signals (is the datapath on this host healthy), cluster services (DNS and service routing), and active probes (does pod A reach service B across nodes, now). Detect policy problems by looking at the verdict the enforcing network plugin gave each flow, not by guessing from application errors. The Kubernetes NetworkPolicy resource is only enforced by the network plugin (the CNI, Container Network Interface), so the plugin's flow and drop visibility is your best source.
Signals to collect
| Layer | Signal | Source | What it tells you |
|---|---|---|---|
| Pod | Flow verdicts (forwarded vs dropped) and drop reasons | Hubble (Cilium's flow-observability component) metrics hubble_flows_processed_total (label verdict) and hubble_drop_total (labels reason, protocol), when Cilium is the plugin | Policy denials versus other drops |
| Pod | TCP flag counts, for example SYN (the first packet of a TCP handshake) with no matching reply | hubble_tcp_flags_total | Connections that start and never complete |
| Pod | DNS queries and responses by response code | hubble_dns_responses_total (label rcode) | Lookups failing inside the cluster |
| Cluster | DNS request duration and response codes | CoreDNS (the cluster's DNS server) coredns_dns_request_duration_seconds and coredns_dns_responses_total{rcode}, served on port 9153 at /metrics | Slow or failing resolution (a rising SERVFAIL rate, the DNS server failed to answer, or NXDOMAIN rate, the name does not exist) |
| Node | Interface counters, TCP statistics, connection-tracking (conntrack, the kernel's per-connection state table) stats | node_exporter collectors netdev, netstat, conntrack (enabled by default) | NIC errors and drops, retransmits, conntrack pressure |
| Path | Synthetic probes between pods, across nodes and to external dependencies | blackbox_exporter (probes HTTP, HTTPS, DNS, TCP, ICMP, gRPC) | End-to-end reachability and latency now, independent of real traffic |
| Service | Request success rate and latency from the application or mesh | App or proxy metrics | Whether users are affected |
Read the table top to bottom as the order of suspicion in the next section: pod verdicts and DNS first, then node counters, then path probes. The Hubble metrics are disabled by default and are switched on in the Helm values with hubble.metrics.enabled, for example {drop,flow,tcp,dns}.
Detecting policy and datapath problems
- Is it policy? Remember the semantics: with no NetworkPolicy selecting a pod, all inbound and outbound connections are allowed. Once a policy selects it, only what a policy allows is permitted. So a service that broke right after a policy rollout is a suspect. With Cilium, list dropped flows for the pod:
hubble observe --pod <pod> --verdict DROPPED. The output shows source and destination namespace/pod:port and the verdict, so you can see the exact flow that a missing rule blocked. The lines look like this (pod names are illustrative; the layout follows the Cilium documentation):
May 4 13:23:47.852: default/xwing:42818 <> default/deathstar-c74d84667-cx5kp:80 Policy denied DROPPED (TCP Flags: SYN)
May 4 13:23:48.854: default/xwing:42818 <> default/deathstar-c74d84667-cx5kp:80 Policy denied DROPPED (TCP Flags: SYN)
Reading it: a timestamp, then source namespace/pod:port, the arrow, destination namespace/pod:port, the reason (Policy denied), the verdict DROPPED, and the TCP flags. A SYN dropped by policy and then sent again about a second later is a client retrying a connection that a rule never lets through, so the fix is the rule, not the application.
2. Is the policy even enforced? A NetworkPolicy created on a cluster whose plugin does not support it has no effect. Test with a deliberate deny rule in a test namespace and confirm the probe fails.
3. Is it DNS? Check the CoreDNS rcode counters and the request duration histogram first, because many "network" outages are failed lookups.
4. Is it the node datapath? Look at node interface drops and errors, and conntrack usage against its limit; a full conntrack table drops new connections on that node only.
5. Is it cross-node? Compare probe results same-node versus cross-node. If only cross-node fails, suspect the overlay (the tunnel the plugin builds to carry pod traffic between nodes) or the underlay (the physical or cloud network beneath it): a tunnel adds header bytes, so a full-size packet may no longer fit the MTU (maximum transmission unit, the largest packet a link carries) and is dropped, while small requests still pass; a firewall between nodes causes the same cross-node-only symptom.
Alerts
Alert on symptoms with rate and ratio: drop rate by reason rising, DNS error ratio, probe success under its target, and conntrack above a set fraction of its maximum. Page on service-level symptoms; keep drop counters as the diagnostic dashboard.
Pitfalls
- Per-pod metric labels that grow without bound as pods churn.
- Watching only application error rates, which cannot separate DNS, policy and datapath causes.
- Trusting that a policy works because it applied without an error.
Write a Python script using only the standard library that measures TCP connect latency to a list of endpoints concurrently and reports p50, p95 and p99 of successful connects. Handle timeouts and failures sensibly.
Sample Answer
Direct answer
Run every attempt as an asyncio task, cap how many run at once with a semaphore, time only the connection setup with time.perf_counter() (a high-resolution timer for short durations; only the difference between two calls means anything) around asyncio.open_connection, wrap it in a timeout, classify each failure (timeout, refused, DNS or other OS error), and compute percentiles over successful connects only while reporting the failure rate beside them. A p99 computed after silently dropping failures flatters the endpoint, so failures are always reported next to the percentiles.
The script
#!/usr/bin/env python3
"""Concurrent TCP connect-latency probe (standard library only)."""
import argparse
import asyncio
import math
import sys
import time
from collections import Counter
def percentile(sorted_vals, p):
"""Nearest-rank percentile: the smallest value with at least p% of samples at or below it."""
if not sorted_vals:
return None
rank = max(1, math.ceil(p / 100 * len(sorted_vals)))
return sorted_vals[rank - 1]
def parse_endpoint(text):
host, sep, port = text.strip().rpartition(":")
if not sep or not port.isdigit():
raise ValueError(f"bad endpoint {text!r}, expected host:port or [v6]:port")
return host.strip("[]"), int(port)
async def probe(host, port, timeout, payload, sem):
"""One attempt. Returns (kind, connect_ms, mbps); kind is 'ok' or a failure class."""
async with sem:
t0 = time.perf_counter()
try:
reader, writer = await asyncio.wait_for(
asyncio.open_connection(host, port), timeout)
except asyncio.TimeoutError:
return "timeout", None, None
except ConnectionRefusedError:
return "refused", None, None
except OSError as exc: # DNS failure, unreachable, reset, ...
return type(exc).__name__, None, None
connect_ms = (time.perf_counter() - t0) * 1000
mbps = None
try:
if payload:
t1 = time.perf_counter()
writer.write(b"x" * payload)
await asyncio.wait_for(writer.drain(), timeout)
await asyncio.wait_for(reader.readexactly(payload), timeout)
secs = time.perf_counter() - t1
mbps = (2 * payload * 8) / secs / 1e6 # bytes sent plus echoed back
except (asyncio.TimeoutError, asyncio.IncompleteReadError, OSError):
pass # connect succeeded, so keep its latency; throughput stays None
finally:
writer.close()
try:
await asyncio.wait_for(writer.wait_closed(), 1)
except (asyncio.TimeoutError, OSError):
pass
return "ok", connect_ms, mbps
async def run(endpoints, attempts, timeout, concurrency, payload):
sem = asyncio.Semaphore(concurrency)
plan = [ep for ep in endpoints for _ in range(attempts)]
done = await asyncio.gather(*(probe(h, p, timeout, payload, sem) for h, p in plan))
results = {ep: [] for ep in endpoints}
for ep, row in zip(plan, done):
results[ep].append(row)
return results
def report(results, payload):
print(f"{'endpoint':24} {'n':>4} {'ok':>4} {'fail%':>6} {'p50':>8} {'p95':>8} {'p99':>8} failures"
+ (" mbps_p50" if payload else ""))
for (host, port), rows in results.items():
lat = sorted(r[1] for r in rows if r[0] == "ok")
fails = Counter(r[0] for r in rows if r[0] != "ok")
pct = lambda p: "-" if not lat else f"{percentile(lat, p):.1f}"
line = (f"{host + ':' + str(port):24} {len(rows):>4} {len(lat):>4} "
f"{100 * (len(rows) - len(lat)) / len(rows):>6.1f} "
f"{pct(50):>8} {pct(95):>8} {pct(99):>8} {dict(fails) or '-'}")
if payload:
mb = sorted(r[2] for r in rows if r[2] is not None)
line += f" {percentile(mb, 50):.1f}" if mb else " -"
print(line)
def main():
ap = argparse.ArgumentParser()
ap.add_argument("endpoints", nargs="+", help="host:port or [v6]:port")
ap.add_argument("-n", "--attempts", type=int, default=50)
ap.add_argument("-t", "--timeout", type=float, default=2.0)
ap.add_argument("-c", "--concurrency", type=int, default=20)
ap.add_argument("--payload-bytes", type=int, default=0,
help="also send this many bytes to an echo service and time the round trip")
a = ap.parse_args()
try:
eps = [parse_endpoint(e) for e in a.endpoints]
except ValueError as exc:
sys.exit(str(exc))
results = asyncio.run(run(eps, a.attempts, a.timeout, a.concurrency, a.payload_bytes))
report(results, a.payload_bytes)
if __name__ == "__main__":
main()
Usage: python3 tcpprobe.py 192.0.2.10:443 [2001:db8::1]:443 -n 100 -c 20 -t 2. Add --payload-bytes 65536 against an echo service (a server that sends back whatever it receives) to get the per-connection round-trip throughput the same way.
What the script prints
#!/usr/bin/env python3
"""Minimal TCP echo server on 127.0.0.1:9001 for testing tcpprobe.py."""
import asyncio
async def handle(reader, writer):
try:
while data := await reader.read(65536):
writer.write(data)
await writer.drain()
except OSError:
pass
writer.close()
async def main():
server = await asyncio.start_server(handle, "127.0.0.1", 9001)
async with server:
await server.serve_forever()
asyncio.run(main())
Save the script above as tcpprobe.py and this one as echo.py. Start the echo server with python3 echo.py &, then run python3 tcpprobe.py 127.0.0.1:9001 127.0.0.1:9002 -n 100; nothing listens on port 9002, so every attempt there is refused. Add --payload-bytes 65536 to the same command to exercise the throughput column.
The output of that command (loopback, so the times are tiny and will differ on every run; the values below are illustrative):
endpoint n ok fail% p50 p95 p99 failures
127.0.0.1:9001 100 100 0.0 0.5 0.8 1.2 -
127.0.0.1:9002 100 0 100.0 - - - {'refused': 100}
Reading the columns: n is attempts, ok is successful connects, fail% is (n minus ok) divided by n, p50, p95 and p99 are connect times in milliseconds over the successful attempts only, and failures counts each failure class. The second row shows why ok and fail% sit beside the percentiles: with no successes there is nothing to take a percentile of, so the cells show - and the failure column carries the story. With --payload-bytes 65536 one more column, mbps_p50, shows the median round-trip throughput.
How the asyncio pieces fit
asyncio runs many tasks in one thread: a task runs until it hits an await, then hands control to another, so hundreds of connection attempts overlap without threads.
async with semtakes one permit from the semaphore (a counter of free slots, here-c); a task that finds none waits, and the permit is returned automatically when the block ends, even after an error.asyncio.wait_for(x, timeout)cancelsxand raisesasyncio.TimeoutErrorif it has not finished in time; that is how a silent host becomes atimeoutrow instead of hanging the run.writer.drain()waits until the data handed towritehas been pushed out to the network, andreader.readexactly(n)waits for exactly n bytes back (and raisesIncompleteReadErrorif the connection closes first).asyncio.gather(...)starts all the probe tasks and returns their results in the order given, which is what letszip(plan, done)pair each result with its endpoint.
Why it is built this way
- What is timed.
open_connectionreturns when the TCP three-way handshake (client sends SYN, server answers SYN-ACK, client sends ACK) completes, so the figure is roughly one network round trip plus local overhead. If the host is a name, DNS resolution is inside that time. Probe addresses, or accept that resolution is part of the figure. - Concurrency is capped. Without the semaphore, hundreds of simultaneous attempts queue inside the client and inflate the numbers it is trying to measure.
-cbounds the client-side load; the tasks for all endpoints share it. - Failures are classified, not merged.
refusedmeans the host answered with a reset (path works, service down).timeoutmeans no answer within the limit (filtered or lost).gaierroris a DNS failure. They have different owners. - Percentiles use the nearest-rank rule (rank = ceil(p/100 x n)). With few successes p99 equals the maximum, which is why
nandokare printed. - Throughput is optional and honest. It needs an echo service to read the bytes back, covers the bytes in both directions, and a failure there keeps the connect sample.
- Cleanup is explicit.
writer.close()runs in afinally, so a failed read does not leak sockets.
Self-check of the percentile function
from tcpprobe import percentile
vals = sorted([12, 15, 11, 14, 13, 90, 12, 16, 11, 250])
print(vals)
for p in (50, 95, 99):
print(p, percentile(vals, p))
print(percentile([], 50))
Output when run:
[11, 11, 12, 12, 13, 14, 15, 16, 90, 250]
50 13
95 250
99 250
None
With 10 samples the 5th value (13) is the p50, and the p95 and p99 ranks both round up to 10, the maximum. That is the arithmetic by hand: ceil(0.95 x 10) = 10.
Pitfalls
- Probe from where users are; loopback or same-rack numbers say nothing about the WAN.
- Alert on a sustained change in p95 or failure rate over several runs, never on one sample. Percentiles of 20 attempts are noisy.
- Repeated short connections can trip rate limits or intrusion detection on the target, so coordinate and keep the attempt rate modest.
Unlock Full Question Bank
Get access to all 8 Network Monitoring and Performance interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.