Network Monitoring and Performance Questions
Network telemetry and performance operations: SNMP polling and traps (including counter wraparound and SNMPv3 access), NetFlow, sFlow and IPFIX flow export, sampling and its accuracy, streaming telemetry (gNMI), and eBPF or packet-capture telemetry; interface-level metrics (utilization, errors, discards, queue depth, microbursts), active synthetic probing alongside passive counters, link-flap detection, baselining and anomaly detection on network signals including elephant-flow spotting, network SLIs and SLOs, alerting, alert-storm suppression and NOC dashboards, telemetry pipeline design, storage, retention, downsampling and cardinality for network data (including securing the collection path and handling bursty remote sites), BGP and link-state monitoring including prefix hijack and route-leak detection, and network capacity monitoring, percentile utilization and bandwidth headroom planning. Also covers measuring and tuning network-level latency, jitter, packet loss and throughput (bufferbloat, queueing, TCP tuning for long paths). Excludes the generic metrics, logs and traces stack and alert design, the layered fault-isolation method and packet-capture troubleshooting, TCP and protocol fundamentals, application and CDN latency engineering, cloud VPC design and security detection.
Design a pipeline to collect, enrich and analyze network flow records arriving at hundreds of thousands per second, including flow logs from cloud environments, with real-time dashboards and hourly rollups. Cover collectors, buffering, processing, storage and failure behavior.
Sample Answer
Direct answer
Run a horizontally scaled set of stateless-ish collectors that decode flows and publish them to a durable streaming log (a replicated queue such as Kafka, split into partitions: independent ordered lanes so many workers can read in parallel). Behind that log, an enrichment stage adds site, application, autonomous system number (ASN, the number that identifies the network owning an address range) and sampling-corrected bytes (routers often export one packet in every N, so bytes are multiplied by N to estimate the real total), and a columnar analytical store holds a few days of raw records plus rollups: one-minute rollups for live dashboards and hourly rollups for long retention. The streaming log is the shock absorber: when the store, the enricher or a collector falls behind, data waits in the log instead of being dropped, and the one place data can still be lost silently (UDP between the router and the collector) is watched with sequence-number gap detection.
Vocabulary bridge
Network readers: Kafka and ClickHouse are the data tools here. A topic is a named log; a consumer is a program reading it, and consumer lag is how far behind the newest record that reader is. A dead-letter topic is a side log for records that cannot be processed. A TTL (time-to-live) deletes rows after a set age. A tumbling window is a fixed, non-overlapping time bucket (every minute, for example); event time is the timestamp inside the record (when the flow ended) as opposed to when the pipeline saw it; a watermark is the pipeline's rule for how long to wait for late records before closing a window. Backpressure means a slow reader making the sender slow down. UDP has none, which is why the log sits behind the collectors. Protobuf is a compact binary record format; JSON is the readable alternative.
Data engineers: an exporter is a router or switch that sends flow records. Before a collector can decode data records it needs the exporter's template (a description of the fields and their order), sent periodically, which is why collector restarts and load balancing matter. An observation domain is an exporter's identifier scope, and a sequence number counts what it has sent, so a gap means loss.
Assumptions and sizing
These figures are design assumptions to be replaced by measurements from your exporters. The point is to show the method.
| Quantity | 300,000 records/s | 1,000,000 records/s |
|---|---|---|
| Wire size at 60 bytes per exported record | 18 MB/s (144 Mbit/s) | 60 MB/s (480 Mbit/s) |
| Decoded and enriched record at 200 bytes | 60 MB/s | 200 MB/s |
| Partitions at a 10 MB/s design target per partition, doubled for headroom | 12 | 40 |
| Log retention of 6 hours of enriched data | 60 MB/s x 21,600 s = 1.30 TB | 200 MB/s x 21,600 s = 4.32 TB |
| Raw records kept 3 days | 77.8 billion rows | 259.2 billion rows |
Raw retention beyond a few days is rarely worth the storage, so the long-retention view is a rollup. A rollup keyed by source site, destination site and application cannot exceed 200 x 200 x 30 = 1.2 million rows per hour for 200 sites and 30 applications. Against 300,000 x 3,600 = 1.08 billion raw records in the same hour, that is a reduction of at least 900 times.
Architecture
The diagram reads left to right: routers send UDP flow packets to collectors, cloud flow logs arrive as files through a queue and a loader, both feed the raw log, enrichment writes an enriched log, and that log feeds the raw store, one-minute rollups for dashboards and hourly rollups for retention.
flowchart LR
R[Routers and switches<br/>IPFIX, NetFlow, sFlow] -->|UDP| C[Collector pool<br/>decode and tag]
V[Cloud flow logs<br/>object storage] --> Q[Notification queue] --> L[Log loader<br/>parse and normalize]
C --> K[(Streaming log<br/>flows-raw)]
L --> K
K --> E[Enrichment<br/>site, ASN, app, sampling]
E --> K2[(Streaming log<br/>flows-enriched)]
K2 --> S[(Columnar store<br/>3 day raw)]
S --> H[Hourly rollups<br/>long retention]
K2 --> M[1 minute rollups<br/>live dashboards]
K -.->|poison records| D[(Dead-letter topic)]
Collectors
- What they run. A flow collector that speaks sFlow v5, NetFlow v5 and v9, and IPFIX (IP Flow Information Export, the standard flow format). goflow2 is one open-source example: it decodes those protocols, takes the sampling rate from the Option Data Set, and can write protobuf or JSON to Kafka. Many parallel UDP sockets per host are configured with its listen option.
- Template state is the catch. NetFlow v9 and IPFIX send a template that says how to read the data records. In IPFIX, template IDs are unique only within a transport session and observation domain (RFC 7011), and exporters over UDP re-send templates periodically. A collector that restarts or a load balancer that sprays one exporter's packets across collectors cannot decode records until a template reaches that instance, which can take minutes. Hash by exporter source IP so each exporter lands on one collector, and add a standby that takes over the same exporter addresses.
- Resource cost. Measure, do not guess: replay captured export traffic at 1.5 times the target rate against one collector and read the kernel's UDP receive-buffer error counter (RcvbufErrors in /proc/net/snmp on Linux) plus CPU. Size the pool with clear headroom below the rate at which that counter starts to rise, and raise the socket receive buffer before adding hosts.
Buffering
Decode at the collector and publish to a topic keyed by exporter, so one router's records stay in order. Keying has a cost: all of one exporter's records land in a single partition, so a very busy exporter becomes a hot partition and delays the other exporters hashed to it. At the 200-byte decoded size and the 10 MB/s per-partition target in the sizing table, one partition carries about 50,000 records per second (10,000,000 / 200). If any single exporter can exceed that, key by exporter plus a hash of the flow key (source, destination, ports, protocol) so its records spread across partitions; the rollups sum by time and site, so they do not need per-exporter ordering. Replicate the log (three copies is a typical choice; the 1.30 TB and 4.32 TB figures above are one copy, so three replicas need about 3.9 TB and 13 TB of disk) and retain at least as long as the longest outage you want to ride out without loss: 6 hours in the sizing above. Consumer lag per stage is the main health signal.
Processing and enrichment
- Normalize to one schema across router flows and cloud flow logs: time, source, destination, ports, protocol, bytes, packets, direction, exporter or resource ID, sampling rate.
- Scale for sampling at this stage: store raw bytes and sampling-corrected bytes (bytes x rate) side by side.
- Enrich by joining against reference data held in memory and refreshed every few minutes: IP to site, application and owner (from the IP address management system and configuration database), interface index to interface name (the INPUT_SNMP field is an index, not a name), and ASN where the exporter does not send one.
- Bucket for rollups by flow end time. Long flows are exported at the active timeout, so a record's bytes are attributed to the minute it ends in. Say so on the dashboard: a one-minute chart smears a long transfer.
- Poison records (undecodable, impossible timestamps) go to a dead-letter topic with a counter, never block a partition.
Storage and rollups
Raw records live in a columnar analytical store (ClickHouse in the example below) with a 3-day time-to-live (TTL), and live dashboards read one-minute rollups fed from the enriched topic, so panels never scan raw rows. A tested ClickHouse pattern for the hourly rollup is a materialized view that writes aggregate states into an AggregatingMergeTree table:
CREATE TABLE flows_raw
(
ts DateTime,
exporter LowCardinality(String),
src_site LowCardinality(String),
dst_site LowCardinality(String),
app LowCardinality(String),
bytes UInt64,
packets UInt64,
sampling_rate UInt32
)
ENGINE = MergeTree
PARTITION BY toDate(ts)
ORDER BY (exporter, ts)
TTL ts + INTERVAL 3 DAY;
CREATE TABLE flows_hourly
(
hour DateTime,
src_site LowCardinality(String),
dst_site LowCardinality(String),
app LowCardinality(String),
est_bytes AggregateFunction(sum, UInt64),
est_packets AggregateFunction(sum, UInt64)
)
ENGINE = AggregatingMergeTree
ORDER BY (hour, src_site, dst_site, app);
CREATE MATERIALIZED VIEW flows_hourly_mv TO flows_hourly AS
SELECT
toStartOfHour(ts) AS hour,
src_site, dst_site, app,
sumState(bytes * sampling_rate) AS est_bytes,
sumState(packets * sampling_rate) AS est_packets
FROM flows_raw
GROUP BY hour, src_site, dst_site, app;
How the statements work: flows_raw holds every record and its TTL line deletes rows after 3 days. flows_hourly stores AggregateFunction(sum, UInt64) columns, which hold a partial running sum (a state) rather than a number, so partial sums from many inserts can be combined correctly later; AggregatingMergeTree merges rows with the same ORDER BY key into one row in the background. The materialized view is a trigger that runs its SELECT on every batch inserted into flows_raw, multiplies bytes by the sampling rate, and writes sumState(...) into flows_hourly. Reading it back uses the matching merge functions (sumMerge(est_bytes)), which combine the stored states into the final number; a plain sum would not work on state columns. Inserting three rows (1,500 and 3,000 bytes at sampling rate 1000 for dc1 to branch-014 in the 10:00 hour, plus 1,200 bytes at rate 100 for another pair in the 11:00 hour) returns 4,500,000 estimated bytes and 120,000 estimated bytes for the two hours, as expected from (1,500 + 3,000) x 1000 and 1,200 x 100.
Because this view adds on every insert, replaying records from the log would double count. When a replay is needed, delete the affected hours from the rollup and recompute them from raw records, rather than re-feeding the view.
Cloud flow logs
Cloud flow logs (AWS VPC Flow Logs as the example) are batched files, not UDP packets: publish them to object storage, send a bucket notification to a queue, and have a loader parse each file into the same normalized schema. Four differences matter:
- Latency. AWS states flow logs "do not capture real-time log streams", so cloud data arrives minutes behind router flows. Because the hourly materialized view adds each row to its hour as the row is inserted, late cloud rows land in the correct hour with no recompute; keep the live dashboard labelled as minutes behind.
- Per-interface logging doubles traffic. Logs are produced for each network interface, so one flow between two instances appears at both ends. Pick one side (typically the source interface, or the flow direction field) before summing.
- Address fields. Through a network address translation (NAT) gateway or transit gateway the plain source and destination fields are the gateway's addresses; the pkt-srcaddr and pkt-dstaddr fields hold the original packet addresses. Enrich on the packet addresses.
- Also excluded from the logs: DHCP, ARP, metadata and DNS-server traffic, so some "missing" flows are by design.
Failure behaviour
| Failure | What happens | Detection and response |
|---|---|---|
| Collector host dies | UDP packets sent to it are lost until failover | Gaps in the IPFIX sequence number per exporter and observation domain (RFC 7011: it counts data records and lets a collector see missed ones), alert on gap rate |
| Collector overloaded | Kernel drops UDP before the process reads it | RcvbufErrors rising, collector CPU; add collectors, grow receive buffer |
| Streaming log broker lost | Replicas keep serving; producers retry | Under-replicated partition alert, producer error rate |
| Enricher or store slow or down | Consumer lag grows, data waits in the log | Lag alert; log retention (6 hours here) is the time budget to fix |
| Poison record | One record fails decode | Dead-letter topic with counter, pipeline keeps moving |
| Exporter template lost | Records undecodable until the next template | Count of "no template" decodes per exporter; hash-to-one-collector rule reduces it |
| Late cloud data | Rows arrive after their hour has ended | The hourly view adds them to the right hour on insert, so nothing is recomputed; in the windowed 1,000,000 records per second design, rows later than the watermark go to the correction path |
Scaling to 1,000,000 records per second
At this rate writing every record is the cost driver (259.2 billion raw rows over 3 days). Add a stream-processing stage between the enriched topic and the store that aggregates into one-minute tumbling windows keyed by (exporter, source site, destination site, application) and writes only the aggregates. The saving depends on how many distinct keys a minute really contains, which you measure from your own traffic. Suppose (illustratively) a minute has 500,000 distinct keys: the writes fall from 1,000,000 rows per second to 500,000 / 60 = 8,333 rows per second, a 120 times reduction. With 100,000 distinct keys it would be 600 times, and with 50,000, 1,200 times, while raw records go to cheap object storage in a columnar file format for 3 days. Use event time (the flow end time, as in the bucketing step above) with a watermark that covers how late a record can arrive after that time: a flow that goes quiet is exported only after its idle timeout has passed, the exporter adds a cache-scan and queueing delay, and cloud flow logs arrive minutes behind. Size the watermark from the measured lag per source, and send anything later to the correction path (recompute the affected hours from raw records). A long flow exported at the active timeout carries an end time close to its export time, so the active timeout does not set the watermark.
Trade-offs
- Decoding at the collector keeps the log small and schema-stable but couples collector upgrades to protocol support. Publishing raw packets and decoding downstream decouples them at several times the log volume.
- A coarser rollup key cuts cost but removes the ability to answer new questions after the fact. Keep raw records long enough to cover the investigation window you actually use.
Telemetry from remote cellular-managed sites drops out and then arrives in bursts. How would you design collection so you lose little data, avoid false alerts when a site goes quiet, and avoid overwhelming collectors on reconnect?
Sample Answer
Direct answer
Make the site responsible for not losing data and the cloud responsible for not being flooded. At each site, a small agent timestamps every sample when it is collected, writes it to a disk-backed queue, and pushes batches outbound (cellular links often sit behind carrier address translation, where the carrier shares public addresses among many devices, so a cloud collector cannot reliably open a connection to the site to pull). On reconnect the agent sends live data first and drains the backlog at a capped rate with random delay. Ingest is idempotent so retries are safe. Alerts distinguish a quiet site (data late, heartbeat late) from a down site by a longer, measured timer, and are evaluated with a lag so bursts do not cause flapping.
Design
- Edge buffer. Timestamp at collection (the gNMI specification likewise defines a notification timestamp as the time the device collected the data from its underlying source, or otherwise the time the target generated the Notification message, not the time a collector received it). Tag each batch with a per-site sequence number so the cloud can detect gaps. Cumulative counters (interface byte counters) survive loss as lost resolution only, since the delta across the gap still gives the right total; gauges (CPU, signal strength) lose the missing moments, so keep 1-minute minimum, average and maximum aggregates as a fallback when the buffer is full and drop the finest data first.
- Sizing. Assume 500 series per site at one sample per 10 seconds: 50 samples per second, at 16 bytes each 800 B/s or 69.12 MB per day. A 72 hour buffer is about 207 MB.
- Idempotent ingest (receiving the same data twice has the same effect as receiving it once). Key each sample by (site, series, timestamp) so a retried batch overwrites itself. Follow the standard retry contract (the Prometheus remote-write specification: senders must retry on 5xx with backoff, may retry on 429, and must not retry other 4xx). The last rule matters: a malformed batch that is retried forever blocks the queue behind it, so route it to a dead-letter store (a side holding area for batches that cannot be processed, kept for inspection) and alert on it. In those codes, 5xx means a server-side error and 429 means "too many requests, slow down".
- Reconnect without a stampede. Two lanes: live samples first, backlog second. Each site has a token bucket (a rate limiter) allowing 5 times its live rate: with 800 B/s live the bucket refills at 4,000 B/s, so the backlog gets the remaining 3,200 B/s. Reconnects start after a random delay, with exponential backoff (waiting longer after each failed attempt, for example 1, 2, 4, 8 seconds) and jitter (a random extra wait so sites do not retry in step) on failures. At the collector, a durable queue sits between the receiver and storage so a surge is absorbed rather than dropped, and admission control returns 429 when the queue is deep.
- Alerts that understand silence. Three signals: a separate heartbeat (tiny message every 30 seconds), data freshness, and an out-of-band check (for example the cellular management platform's online status). Page on missing heartbeat only after a time longer than normal dropouts (take the 95th percentile of past outage durations plus margin: if 95% of past dropouts ended within 20 minutes and you add a 10 minute margin, page at 30 minutes of silence; a 3 minute dropout then never pages), using a pending period before firing (Prometheus calls this the
forclause) and a hold-down after recovery (keep_firing_for, which the documentation describes as preventing flapping and false resolutions caused by missing data). Evaluate threshold alerts on data older than the usual delay, and when a backlog arrives re-evaluate history for the record but page only on current conditions.
Worked example
1,000 sites, a regional carrier outage of 4 hours:
- Backlog per site: 800 B/s x 14,400 s = 11.52 MB. Fleet: 11.5 GB.
- Unthrottled at once to a 200 Mbit/s (25 MB/s) collector: 11.52 GB / 25 MB/s is about 461 s of saturated ingest, during which live data queues behind the backlog.
- With the cap: each site sends 4,000 B/s, of which 800 is live, so the backlog drains at 3,200 B/s and clears in 11.52 MB / 3,200 B/s = 3,600 s (1 hour). The 4,000 B/s cap already includes the 800 B/s live share, so fleet load is 1,000 x 4,000 B/s = 4 MB/s = 32 Mbit/s in total: 1,000 x 3,200 B/s = 25.6 Mbit/s of backfill plus 1,000 x 800 B/s = 6.4 Mbit/s live. That is well under the 200 Mbit/s collector, against 200 Mbit/s fully saturated in the unthrottled case.
- Spreading reconnects over 300 s starts only 3.3 sites per second.
Verification
In a lab, cut one modem for 4 hours and compare samples produced against samples ingested using the sequence numbers (target: zero gaps). Cut for 3 minutes (no page expected) and for longer than the timer (one page, one recovery, no flapping). Reconnect 50 simulated sites together and watch queue depth and live-data delay.
Trade-offs and pitfalls
- Backlog drains slower than a stampede, so history is complete later. That is the price for protecting live visibility.
- A long heartbeat timer means slower detection of a real failure. Set it per site criticality, and for critical sites add a second path.
- Clock drift at the edge corrupts event time. Sync time at the site and alert on drift.
Passive counters say the network is fine but users report slowness. How would you set up synthetic active checks to complement them: what to probe, how often, from where, and how to use the results?
Sample Answer
Direct answer
Build a ladder of probes (synthetic test requests run on a timer by the Prometheus blackbox exporter, a service that makes the probe and reports the result as metrics) that each test one more layer, run them from inside every branch toward the data center, and alert on the success ratio and latency percentile over a window, not on a single failed probe. A ladder of five probes at 15-second intervals from each branch (local gateway, hub router, DNS, TCP connect, HTTP request) costs 50 probes per second for 150 branches and pinpoints the failing layer. Use the results next to the counters: when passive data says healthy and active data says broken, the fault is on a layer or path the counters do not describe (a dependency, a policy, a provider path, a queue in one traffic class).
What to probe (the ladder)
| Probe | Module (blackbox exporter) | Question it answers | If it fails while the one above passes |
|---|---|---|---|
| Local gateway, ICMP (Internet Control Message Protocol) echo | icmp_probe | Is the branch LAN and gateway alive? | The switch, cabling, or gateway at the branch |
| Hub router, ICMP echo | icmp_probe | Is the WAN path up, and what are round-trip time (RTT) and loss? | WAN circuit, provider, or routing |
| DNS A-record query to the resolver | dns_a | Can the branch resolve the application name? | Resolver, or DNS traffic filtered |
| TCP connect to the application's virtual IP address (VIP, one address that fronts several servers) on port 443 | tcp_connect | Can a connection be opened? | Firewall policy, load balancer, or path for that port |
| HTTP GET of a health URL | http_2xx | Does the application answer correctly? | Application or TLS, not the network |
The Prometheus blackbox exporter provides the http, tcp, icmp, dns and grpc probers. The module names in the table are the ones defined in blackbox.yml below, and each is built on one prober type. Each probe returns probe_success (1 or 0) and probe_duration_seconds, and the ICMP prober returns probe_icmp_duration_seconds split by phase (resolve, setup, rtt). Use the rtt phase for latency.
How often
Detection time is roughly interval x number of failures you require, plus any for wait. At 15 seconds, the 5-minute window holds 20 probes, so the rule avg_over_time(probe_success[5m]) < 0.9 needs 3 failures (2 of 20 gives exactly 0.9 and does not fire), then a 2-minute for wait (the alert must stay true for that long before it fires). Step by step: with probes every 15 s and the outage starting at the first failed probe at t = 0, failures land at t = 0, 15 and 30 s. Two failures leave 18 of 20 successes, which is exactly 0.9 and does not fire; the third leaves 17 of 20 = 0.85, so the condition turns true at about t = 30 s. Adding the 2-minute wait gives about 150 s. Rule evaluation adds its own delay. Replaying this rule with promtool test rules, with probe_success going to 0 at a known time, the alert fired 150 s after the first failed probe when the rules were evaluated every 15 s, and 180 s after it with the default 1-minute evaluation_interval (the first evaluation after the condition turns true falls at 60 s, plus the 2-minute wait). Set the rule group's interval to 15 s for this rule, or accept the extra half minute. avg_over_time(probe_success[5m]) averages the 1s and 0s of the last 5 minutes into a success ratio. Probe less often (60 s) where an outage can wait, more often only for the few services where waiting 150 seconds to page is too slow. Cost is linear: 150 branches x 5 probes / 15 s = 50 probes per second.
From where
- Inside each branch, on a small probe host or container on the user VLAN, so the probe takes the same path and policy as users.
- From the hub toward branches, for the return direction (asymmetric routing and one-way faults).
- From a cloud vantage toward public services, to separate your WAN from the internet.
- At least two independent targets per layer (for example two hub routers, two DNS servers) so a dead target is not mistaken for a dead path.
Configuration (Prometheus scraping the blackbox exporter)
The exporter's modules (blackbox.yml): each block names a module and says which prober to use and how long to wait before counting a failure.
modules:
icmp_probe:
prober: icmp
timeout: 3s
tcp_connect:
prober: tcp
timeout: 3s
http_2xx:
prober: http
timeout: 5s
dns_a:
prober: dns
timeout: 3s
dns:
query_name: "app.example.com"
query_type: "A"
Prometheus then calls /probe?target=<target>&module=<module> on the exporter once per scrape. The relabel rules build that URL from the target's labels, in this order: the device address in __address__ is copied to __param_target (a label named __param_X becomes the URL parameter X, so this is target=); the label probe_module from the targets file is copied to __param_module, which becomes module= and picks the blackbox module; the address is also copied to instance so stored series are labelled with the probed device; and finally __address__ is overwritten with the exporter's own address, 127.0.0.1:9115, where the request is really sent. file_sd_configs (file-based service discovery) reads the target list from the YAML file below, which can change without restarting Prometheus:
scrape_configs:
- job_name: branch_probes
scrape_interval: 15s
scrape_timeout: 10s
metrics_path: /probe
file_sd_configs:
- files: [targets/branch-probes.yml]
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [probe_module]
target_label: __param_module
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: 127.0.0.1:9115
# targets/branch-probes.yml (one entry per probe)
- targets: ['10.20.1.1']
labels: {site: branch-014, probe_module: icmp_probe, probe: lan_gateway}
- targets: ['10.0.0.1']
labels: {site: branch-014, probe_module: icmp_probe, probe: hub_icmp}
- targets: ['10.0.8.53:53']
labels: {site: branch-014, probe_module: dns_a, probe: app_dns}
- targets: ['10.0.8.10:443']
labels: {site: branch-014, probe_module: tcp_connect, probe: app_tcp}
- targets: ['https://app.example.com/health']
labels: {site: branch-014, probe_module: http_2xx, probe: app_http}
groups:
- name: synthetic-path
rules:
- alert: BranchProbeFailing
expr: avg_over_time(probe_success{probe="app_tcp"}[5m]) < 0.9
for: 2m
labels:
severity: page
- alert: BranchPathSlow
expr: quantile_over_time(0.95, probe_icmp_duration_seconds{probe="hub_icmp", phase="rtt"}[10m]) > 0.15
for: 5m
labels:
severity: ticket
quantile_over_time(0.95, ...) takes the 95th percentile (p95: 95 percent of the values are at or below it) of the last 10 minutes of round-trip times. The 150 ms threshold assumes a branch whose normal round trip to the hub is far below that. Set it from two weeks of the branch's own baseline plus a margin, not from a guess.
How to use the results
- Alert on ratio over time, as above, never on one failed probe.
- Set a service-level objective (SLO) on the same signal: for example 99.5 percent probe success per site per 30 days, an error budget (the amount of failure the objective permits) of 43,200 minutes x 0.5 percent = 216 minutes.
- Localize with the ladder and the counters:
| Active result | Passive result | Likely location |
|---|---|---|
| Hub ping slow or lossy | Uplink utilization near line rate | Congestion: look at queues and QoS (quality of service) |
| Hub ping slow or lossy | Uplink utilization low, counters clean | Provider path or a policing or queueing fault in one class |
| Ping fine, TCP connect fails | Flows show packets leaving but none returning | Firewall policy or asymmetric routing (replies take a different path than requests and are dropped by a stateful device) |
| TCP fine, HTTP fails or slow | Network counters clean | Application tier |
| All probes from one branch fail, other branches fine | That branch's uplink up | Branch gateway or provider circuit |
| All branches fail the same target | Hub counters clean | The target or hub-side service |
- Feed changes: run the same ladder before and after a change window and compare p95 RTT.
Pitfalls
- Probes measure the probe's class of traffic. If voice is queued separately, an ICMP probe may look fine while voice suffers: send probes marked with the same DSCP (differentiated services code point, the QoS marking in the IP header) as the traffic you care about. The blackbox exporter has no DSCP setting for its icmp, tcp or http probers, so apply the marking outside it (a firewall rule on the probe host that marks the exporter's outgoing packets) or use a probe tool that sets DSCP itself.
- Do not alert on every probe type at page severity. Page on the TCP probe, ticket on latency, and use the others for diagnosis.
- A probe host that is itself unhealthy produces false outages. Alert on the probe host's own metrics and on a probe that checks something known good.
Architect network observability across cloud, on-prem and edge covering about 50,000 devices, with sub-second alerting for critical failures and long-term trend retention. Where do you standardize, what do you collect, and how do you keep alert volume under control?
Sample Answer
Scale first. Assume about 40 interfaces and 6 series each across 50,000 devices: 12 million series, 400,000 samples per second at a 30-second interval (computed: 50,000 x 40 x 6 = 12,000,000 series, divided by 30 s). Raw storage at 2 bytes per sample (ESTIMATE; Prometheus documents 1 to 2) is about 2.1 TB for 30 days: 400,000 x 2 B = 800,000 B/s, x 86,400 s = 69.12 GB per day, x 30 days = 2.07 TB. One server will not do that, so the design is regional collection plus central query.
Where to standardize
- Identity and labels from one inventory (source of truth). Every device, in every domain, carries
device,site,region,device_role,team,envfrom the inventory, never from the device. Alert routing and dashboards depend on these labels being identical everywhere. - Metric names and units (base units,
_totalon counters), the same scheme for cloud, on-prem and edge, so a dashboard works across all three. - Transports: every device class needs a baseline: SNMP (Simple Network Management Protocol) v3 for counters and state, IPFIX (IP Flow Information Export, records of who talked to whom) for flows, syslog for events. Where devices support it, gNMI (a streaming management protocol: the device sends data to the collector instead of being polled) carries counters and state changes, with OpenConfig (vendor-neutral data models) so the same path means the same thing across vendors; SNMP v3 is the fallback. The sub-second machinery below (BFD, on-change streaming) applies only to the few thousand critical paths. Cloud networks feed the same store from the provider's metrics and flow logs.
- Severity taxonomy:
criticalmeans page a human now,warningmeans ticket,infomeans dashboard only.
What to collect
- Counters: interface rates, errors, discards, operational status, every 30 to 60 seconds, per region.
- State changes: link, BGP session and BFD state via on-change streaming events.
- Flows: sampled IPFIX per site, aggregated regionally.
- Synthetic probes between sites and to key services (ICMP, DNS, HTTP).
- Syslog and configuration-change events.
- Retention: 30 days raw, 5-minute rollups for 13 months, hourly beyond (long-term store through remote write). A rollup is a pre-computed lower-resolution copy: one averaged point per 5 minutes replaces ten 30-second samples, so old data costs far less to store and query.
Sub-second alerting for critical failures, stated honestly. Scraping cannot do it: polling every 30 seconds adds up to 30 seconds. Sub-second detection happens on the device (BFD (Bidirectional Forwarding Detection), designed in RFC 5880 for low-latency failure detection, and hardware link-down). To get the alert to people quickly, subscribe to critical state with gNMI STREAM / ON_CHANGE, which the specification defines as sending an update when the value changes, into an event pipeline that feeds the alert manager. Apply this only to a few thousand critical paths (core links, border, WAN edges), not all 50,000 devices. An illustrative timeline: at t = 0 the link fails. BFD with a 50 ms interval and a miss count of 3 (illustrative settings; detection time is the interval times the count, per RFC 5880) declares the session down at about 0.15 s. The device sends the ON_CHANGE update (expected within milliseconds on hardware that supports it; an assumption to measure, not a guarantee), the event pipeline turns it into an alert and hands it to Alertmanager, and the critical route then waits its group_wait (5 s in the config below, so related alerts can join the page and inhibition can apply) before the first notification. So detection is sub-second, while the page arrives after roughly group_wait plus delivery time; lowering group_wait buys speed at the cost of more duplicate pages. End-to-end latency (device event to page) is a number to MEASURE by injecting a link failure in a lab, not to assume.
Alert volume control (Alertmanager, config validated with amtool check-config)
Four terms used below: grouping bundles alerts that share chosen labels into one notification; inhibition automatically mutes alerts whose cause is already alerting; a mute time interval is a recurring named window in which a route sends nothing; a silence is a one-off manual mute with matchers and an expiry, created in the UI or API.
route:
receiver: noc-queue
group_by: [alertname, region, device_role]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- matchers: ['severity="critical"']
continue: true
receiver: noc-page
group_by: [region, site]
group_wait: 5s
- matchers: ['team="security"']
receiver: secops-queue
- matchers: ['severity="warning"', 'team="neteng"']
receiver: neteng-ticket
mute_time_intervals: [weekly-core-maintenance]
inhibit_rules:
- source_matchers: ['alertname="DeviceDown"']
target_matchers: ['alertname=~"InterfaceDown|HighLatency|BGPSessionDown"']
equal: [device]
- source_matchers: ['alertname="SiteIsolated"']
target_matchers: ['alertname=~"DeviceDown|InterfaceDown"']
equal: [site]
time_intervals:
- name: weekly-core-maintenance
time_intervals:
- times: [{start_time: '02:00', end_time: '04:00'}]
weekdays: ['tuesday']
(Receiver definitions are omitted here; amtool needs them present.) Reading it:
routeis the root of the routing tree.receiveris the default destination andgroup_bylists the labels whose values define a group: alerts with the samealertname,regionanddevice_rolego out as one notification.group_wait: 30sis how long a new group waits before its first notification, so related alerts can arrive and be bundled.group_interval: 5mis how long to wait before notifying about new alerts added to a group already notified.repeat_interval: 4his how often a still-firing, unchanged group is reminded.- Under
routes, child routes are tried top to bottom and the first match wins unless the route setscontinue: true. The first one matches every critical alert whatever its owner, sends it tonoc-page, groups it byregion, siteinstead, overrides the wait withgroup_wait: 5sso pages go out sooner, and setscontinue: trueso a security-owned critical alert is also offered to the next route instead of stopping here (checked withamtool config routes test: critical alerts for team neteng or noc resolve tonoc-page, and critical for team security resolves tonoc-pageandsecops-queue). The next sends security-owned alerts to a queue. The third sends NetEng warnings to tickets and refers by name to a mute window. inhibit_rules: when an alert matchingsource_matchersis firing, alerts matchingtarget_matchersare suppressed, but only if both carry the same value for each label inequal. SoInterfaceDownis muted only for the device that is down, not for every device.time_intervalsdefines the named windowweekly-core-maintenance: Tuesdays 02:00 to 04:00 in UTC, which is the default; addlocation(an IANA timezone name such as Europe/London) to the interval if the window should follow local time.- Dedup and grouping: hundreds of alerts from one site outage collapse into one notification per
region, sitegroup. - Inhibition: when a device is down, its interface, latency and BGP alerts are suppressed (matching on the same
device); when a site is isolated, device and interface alerts for thatsiteare suppressed. - Maintenance windows: recurring work uses
mute_time_intervals(note the root route cannot have mute times, so they sit on child routes); one-off work uses a silence created for the change window. - Role-based routing and escalation: NOC receives every critical alert as a page (anything unmatched lands in the
noc-queueticket queue, which does not page). NetEng gets warnings as tickets. Security gets only alerts labelledteam="security". If a NOC page is unacknowledged for a set time (the paging tool's escalation policy, for example 15 minutes), it escalates to the NetEng on-call and then to the owning team lead. - Alert quality bar: review monthly the share of pages that led to human action, and demote or delete rules that fall below your bar (for example half).
Dashboards by audience: NOC gets a site-state map and the active alert list; NetEng gets per-device and per-interface trends, capacity forecasts and change timelines; Security gets flow anomalies and configuration changes; leadership gets availability per region.
Network telemetry arrives at very different rates: interface counters, flow aggregates and routing events. What granularity and retention plan would you set for each, what do you keep at full resolution, what do you roll up, and why?
Sample Answer
Choose granularity and retention from the questions each data type has to answer, because the three sources behave very differently: interface counters are steady and cumulative, flow records are bulky and per-conversation, and routing events are rare and order-sensitive.
Interface counters (SNMP or streaming telemetry)
- SNMP (Simple Network Management Protocol) counters such as bytes in and out only ever increase, so the collector computes a rate from two readings. Poll every 30 seconds on uplinks and core links, and every 60 seconds on access ports. A coarser poll loses no bytes, only visibility of short bursts.
- Use the 64-bit counters (ifHCInOctets and ifHCOutOctets). RFC 2863 requires 64-bit octet counters on interfaces faster than 20 Mbit/s, because a 32-bit octet counter wraps in about 34 seconds at 1 Gbit/s (RFC 2863). At 10 Gbit/s the same arithmetic gives 2^32 x 8 / 10^10, about 3.4 seconds, so a 30-second poll of a 32-bit counter on a 10G link is wrong most of the time.
- Keep full resolution (30 s) for 30 days, 5-minute rollups for 13 months (long enough to compare this month with the same month last year), and 1-hour rollups longer if capacity planning needs it.
- A rollup replaces many readings with a few summary numbers. For example, one 5-minute bucket holds 10 readings at 30 s. With the illustrative values 20, 25, 22, 30, 28, 24, 26, 95, 27 and 23 Mbit/s the rollup stores average 32, maximum 95 and 95th percentile 95 (the value that at least 95 percent of the readings do not exceed, using the nearest-rank method: sort the readings and take the one at position ceil(0.95 x count), which for 10 readings is position 10, the largest. A tool that interpolates between readings, such as numpy's default percentile, reports about 65.75 for these ten, so state which method your rollup uses). Roll up with average, maximum and the 95th percentile computed from the raw points at rollup time. Never store only the average: a 5-minute average of 40% can hide a 20-second burst at 100% that caused drops. And never average 95th percentiles later; a percentile of percentiles is not the percentile of the underlying data. Example: half a day has 100 readings of which 90 are 10 and 10 are 100 (its 95th percentile is 100), and the other half has 100 readings all equal to 20 (95th percentile 20). Averaging the two gives 60, but the 95th percentile of all 200 readings is 20, because only 10 of 200 readings (5 percent) are high (nearest-rank, position 190 of 200; an interpolating tool reports about 24, still nowhere near 60). Keep the raw points to compute the percentile, or store enough to recompute it.
Flow data (NetFlow v9, RFC 3954, or IPFIX, the IETF standard that followed it; both are export formats for the same kind of flow record, and a collector usually accepts either)
- A flow record says who talked to whom, on which ports, how many bytes. Volume scales with the number of conversations, not the number of links, so on fast links export sampled flows (for example 1 in N packets) and multiply byte counts by N when reporting. Sampling hides very small flows, which matters for security work but not for capacity work.
- Keep raw flow records for 7 days (enough to investigate an incident someone reports on Monday about Friday), aggregate at the collector into 1-minute buckets keyed by source prefix (a block of addresses such as 10.1.0.0/16), destination prefix, protocol, destination port and interface, and keep those aggregates for 13 months. Top-talker views and capacity questions come from the aggregates.
Routing events (BGP session changes, OSPF adjacency changes, link up/down)
- These arrive as syslog messages (the standard way devices send text log lines), SNMP traps (unsolicited alerts a device sends when something changes) or streaming on-change updates, and there are few of them. An OSPF adjacency is the working neighbour relationship between two routers, so an adjacency change means a neighbour was lost or gained. Keep every event at full fidelity (each one stored unaltered), with its device timestamp, for 13 months. Do not roll them up: a rollup of events destroys the one thing you need, the exact order and timing relative to a traffic dip.
- Alongside, sample slow-moving state (route count per peer, adjacency count) every 5 minutes as an ordinary metric.
What I keep at full resolution, what I roll up, and why
| Data | Full resolution kept | Rolled up | Reason |
|---|---|---|---|
| Interface counters | 30 days at 30-60 s | 5 min (13 months), 1 h beyond | Troubleshooting needs seconds; trends need months |
| Flow records | 7 days raw | 1-minute aggregates for 13 months | Raw volume is large; the question is usually "which prefix or port" |
| Routing events | All events, 13 months | Never | Rare and order-sensitive |
The cost driver to watch is series count times polling frequency. A series is one measured value over time, such as bytes out on one port. With an illustrative fleet of 1,000 devices, 48 ports and 2 counters (in and out) that is 96,000 series; polled every 30 s that is 3,200 samples per second, or 2,880 samples per series per day and 86,400 per series over 30 days of full resolution. The 13-month tier at 5 minutes with average, maximum and p95 is 288 x 3 x 395 = 341,280 values per series, kept 13 times longer but at one tenth the time resolution. So spend resolution on uplinks and core links first, and use the coarser interval where an hour of lost detail would never change a decision.
Unlock Full Question Bank
Get access to all Network Monitoring and Performance interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.