Observability and Monitoring Architecture Questions
Building visibility into infrastructure and services: metrics, logs, and traces, dashboards and alerting, SLIs/SLOs, and the design of an observability stack. Covers instrumenting systems for actionable signal, reducing alert noise, and diagnosing production issues from telemetry. Infrastructure-wide observability, distinct from network-specific monitoring.
Your team keeps getting paged at 2 a.m. for a disk space alert that clears itself ten minutes later before anyone can act on it. How would you redesign the alert so it stops paging on transient spikes but still catches real capacity problems?
Sample Answer
Direct answer
A threshold that fires on a single instantaneous sample is the root problem: it reacts to noise, not to a sustained condition. The fix is to require the condition to hold for a minimum duration meaningfully longer than the known false-alarm's own duration before it pages, and to separate "worth looking at during business hours" from "wake someone up," instead of treating every threshold breach as page-worthy.
Structured elaboration
Concretely:
- Change the check from "disk usage over 90% right now" to "disk usage over 90% for at least 15 consecutive minutes." Since this specific false alarm already resolves on its own in about ten minutes, the sustained window needs real margin past that, not a duration equal to it, or the same class of spike could still accidentally satisfy the rule.
- Add a second, lower-severity threshold, for example a ticket or a chat notification at 85% sustained for an hour, that goes to the team's queue instead of paging, so a genuinely slow leak still gets noticed before it becomes urgent.
- Use trend, not just level, where you can: usage climbing 5% per hour is a different problem than sitting flat at 91% for months on a disk that runs nearly full by design.
- Route the actual page only for the sustained, high-severity case, and route everything else to a lower-urgency channel with a runbook link, so on-call learns to trust that a page always means act now.
Worked example
A minute-by-minute simulation makes the trade-off concrete. Baseline disk usage sits at 70%. A transient spike jumps to 93% at minute 10, holds for exactly 10 minutes (matching the alert's own known pattern), and drops back to 71% at minute 20, nothing wrong. Separately, a real leak starts climbing 1% per minute from minute 30 onward and first crosses 90% at minute 49, staying above it from then on.
- Instant threshold (pages the moment any sample is at or above 90%): fires immediately at minute 10 for the spike, a false page, and fires again at minute 49 for the real leak. Across the run it fires on 42 separate minutes total.
- Sustained-15-minute threshold: never fires during the 10-minute spike, since 10 consecutive minutes does not reach the 15-minute bar. It does confirm the real leak, paging at minute 63, 15 consecutive minutes after the climb first crossed 90% at minute 49, a 14-minute detection delay traded for silencing the known false-alarm pattern with real margin.
That 14-minute delay is the actual cost of the redesign. If this were a volume that could fill from 90% to 100% faster than that under a genuinely fast leak, a flat 15-minute duration would be too slow and would need shortening or pairing with a rate-of-change check instead.
Trade-offs and pitfalls
- A "for" duration trades detection speed for noise reduction: too short and you are back to paging on spikes, too long and a genuine fast fill (a runaway log during an incident) eats into your response time. Size the duration to the fastest realistic growth rate you actually need to catch in time.
- Static percentage thresholds do not generalize across disks of very different sizes; 90% full on a 20 TB data volume is very different headroom than 90% on a 20 GB boot volume. Consider an absolute free-space floor in addition to a percentage.
- Silently raising the threshold to make the pages stop is the wrong fix and the most common mistake here; it just narrows the window before a real full-disk incident.
What the interviewer probes next
They will usually push on how you would gain confidence in a new duration setting before it goes live in production, and on how your same alerting policy needs to bend for a volume that is intentionally expected to run near full, like a cache tier.
One of your Linux servers has a load average of 8 on a 4 core box, but top shows CPU usage sitting around 20 percent. Walk me through how you would figure out what is actually driving that load.
Sample Answer
Direct answer
Load average and CPU percentage measure different things: load average also counts processes waiting on I/O in the run queue, not just processes waiting for CPU. On a 4 core box, a load of 8 means, on average, two processes are queued per core, a real overload signal, but since CPU sits at only 20% they are not queued waiting for CPU time. That combination almost always means something is I/O bound (disk, network, or a lock), not CPU bound. The fix is to find which processes are stuck in that waiting state and what they are waiting on.
Structured elaboration
Walk the toolchain in order:
uptime: confirms the load trend (1/5/15 min averages) is real and not a one-off spike.vmstat 1 5: check thewa(I/O wait) column. Highwawith lowus/syconfirms the CPU is idle waiting on I/O, not busy computing.iostat -x 1 5: look at%utilandawaitper device to find which disk is saturated.ps auxfiltered on process state, to find processes in uninterruptible sleep (stateD), the OS-level signature of a process blocked on I/O that cannot even be killed until the I/O completes.
Worked example
Ran the exact filter against a sample process table to demonstrate the technique, not a live host:
== processes stuck in uninterruptible sleep (D state) ==
USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
mysql 1122 12.4 8.2 812340 84200 ? Dl 09:14 4:41 /usr/sbin/mysqld
backup 2200 0.4 0.2 15200 3100 ? D 09:41 0:00 tar czf /backups/db.tar.gz /data
web 980 3.1 1.5 95200 22100 ? S Jul24 1:02 nginx: worker process
Command: ps aux | awk 'NR==1 || $8 ~ /^D/' (in a real ps aux listing, STAT is field 8, right after TTY). Run against the table above it drops the healthy nginx worker (state S) and keeps only the two processes actually blocked on I/O: a MySQL process and a backup tar job, both in state D/Dl. That narrows an 8 load average down to two concrete suspects and explains high load with idle-looking CPU.
Trade-offs and pitfalls
top's default view hideswa; you have to check the CPU detail line or usevmstat/mpstatto see it.- A high load average from many short-lived processes (a fork bomb, a cron storm) can look similar to an I/O problem at a glance;
vmstat'sr(runnable) column versusb(blocked) column tells them apart. - Network-attached storage (NFS, iSCSI) can cause
Dstate processes while iostat shows the local disk as idle. Checknfsstator network latency too.
What the interviewer probes next
They will usually push on what you would do once you have identified the blocked process (kill it, or is it un-killable and you need to fix the underlying storage), and whether you would have caught this earlier with a proactive alert instead of a manual investigation.
How would you configure a monitoring check so a Windows service only pages after it has been down for two consecutive check intervals, and routes to a different on-call group than a Linux daemon alert?
Sample Answer
Direct answer
You get this with two independent mechanisms: the check's own retry and interval settings to require consecutive failures before declaring a hard state, and the notification configuration to route by host group rather than one global contact list. In Nagios terms that is max_check_attempts plus retry_interval; in Zabbix terms that is a trigger evaluated over recent history plus a separate Action per host group.
Structured elaboration
Nagios detects the failure through its state machine: a check result that fails moves the service from OK to a SOFT non-OK state (no notification yet); Nagios then re-checks at the faster retry_interval instead of the normal check_interval. Only once max_check_attempts consecutive checks have failed does the service become HARD and a notification fire. contact_groups then decides who gets that notification, so pointing a Windows service at one contact group and a Linux daemon check at a different one is a routing decision, entirely independent of the retry logic above.
Zabbix reaches the same outcome differently: a trigger expression evaluated against recent item history (for example the last 2 collected values) decides the PROBLEM state, and a separate Action, scoped to the host's group, decides who is notified and how.
Worked example
Nagios service definition:
define service {
host_name winsvc01
service_description Critical Windows Service
check_command check_nt!SERVICESTATE!-d SHOWALL -l MyService
max_check_attempts 2
check_interval 5
retry_interval 1
contact_groups windows-oncall
}
max_check_attempts 2 means the service must fail on two consecutive checks (one SOFT, then one HARD) before Nagios notifies, which is "down for two consecutive check intervals." retry_interval 1 makes that second, confirming check happen fast, one minute after the first failure, rather than waiting a full normal check_interval. contact_groups windows-oncall routes the page to a different rotation than the Linux daemon's contact_groups linux-oncall.
Zabbix equivalent: a trigger using the count() history function to require two consecutive failed values, paired with an Action scoped to the Windows host group:
count(/winsvc01/service.state,#2,"eq",1)=2
This reads: of the last 2 collected values (#2) of the service.state item, count how many equal 1 (the down state); if that count is 2, both of the last two checks were down. Pair it with a separate Action whose condition is "Host group = Windows Servers" pointing at the Windows on-call escalation, distinct from the Action tied to the Linux host group.
Trade-offs and pitfalls
- The retry mechanism trades detection latency for noise reduction: with
retry_interval 1andmax_check_attempts 2, confirmation itself only takes about a minute once the first failure is caught, but that first failure can occur anywhere inside the normal 5-minutecheck_interval, so worst-case time from "the service actually went down" to "on-call gets paged" is just under 6 minutes, not a flat 10. - Routing by host group only works if your inventory (which host belongs to which group) is accurate and kept current; a misclassified host pages the wrong team, which erodes trust in the whole system fast.
- Do not forget dependency checks: if the Windows host itself is unreachable, you want one page for "host down," not one for every service on it, which is what Nagios host/service dependencies and Zabbix trigger dependencies are for.
What the interviewer probes next
They will usually ask how you would avoid alert storms when an entire host group goes down at once, pointing at dependencies or maintenance-mode suppression, and how you would test that the escalation actually reaches the right rotation before finding out during a real incident.
Compare Nagios, Zabbix, and Prometheus with node exporter as the monitoring stack for a 500 host estate that mixes bare metal Windows and Linux servers, network gear, and a few cloud VMs. What would push you toward one over the others?
Sample Answer
Direct answer
For a 500 host hybrid estate with bare-metal Windows and Linux, network gear, and a handful of cloud VMs, the real decision driver is agent model and protocol reach, not feature checklists. Nagios and Zabbix both natively speak SNMP for network gear and have mature agents for Windows, while Prometheus with node_exporter is pull-based, Linux-native, and has weak native support for Windows and none for SNMP without extra exporters. That single fact usually decides it for a genuinely mixed on-prem estate.
Structured elaboration
- Nagios: agent-based (NRPE or NSClient++) or agentless (SSH, SNMP) checks, a huge plugin ecosystem built over decades, but configuration is mostly flat text files, which gets unwieldy past a few hundred hosts without a config-generation layer on top such as Puppet or Ansible templates.
- Zabbix: has its own lightweight native agent for both Windows and Linux, first-class SNMP polling built in, SNMP trap intake via its bundled trap-receiver script paired with Net-SNMP's
snmptrapd, a real web UI with database-backed config that scales more gracefully to hundreds of hosts than flat Nagios files, and built-in trend storage so you do not need a separate time-series database. - Prometheus with node_exporter: node_exporter only covers Linux/Unix host metrics well; Windows needs a separate
windows_exporter, and network gear needssnmp_exporteras a bolt-on translator, since Prometheus itself only speaks its pull-based HTTP scrape protocol, not SNMP. It shines for containerized and cloud-native workloads with dynamic service discovery, a poor fit for static bare-metal inventory that barely changes month to month.
Given this estate, mostly static Windows/Linux bare metal plus network gear plus a few cloud VMs, Zabbix is usually the pragmatic default: one agent model covers both operating systems, SNMP is native rather than bolted on, and the host count is static enough that Prometheus's dynamic discovery strength is not worth much here. If the cloud VM footprint were the majority instead of a handful, that calculus flips toward Prometheus.
Worked example
Put concrete numbers on the 500 hosts: say 380 Linux bare metal, 80 Windows bare metal, 30 network switches and UPS units, and 10 cloud VMs (380 + 80 + 30 + 10 = 500). Zabbix's single native agent covers all 460 Windows and Linux hosts, its native SNMP polling covers the 30 network devices, and the same agent covers the 10 cloud VMs too, one tool for all 500. Prometheus with node_exporter alone only natively reaches the 380 Linux hosts; the 80 Windows hosts need windows_exporter, and the 30 network devices need snmp_exporter. That is still three separate exporter types, node_exporter, windows_exporter, and snmp_exporter, deployed across the estate just to match the coverage Zabbix gets from one agent.
Trade-offs and pitfalls
- Nagios's flat-file config becomes a real operational burden past a few hundred hosts without a templating layer; teams often underestimate this until they are maintaining thousands of object definitions by hand.
- Zabbix's backing database (usually MySQL or PostgreSQL) becomes something you now have to keep highly available and capacity-plan for, a new operational dependency you did not have with flat-file Nagios.
- Running Prometheus here means running and maintaining three separate exporters, node, windows, and snmp, plus Prometheus itself plus Alertmanager, more moving parts than a single Zabbix or Nagios install for the same coverage.
- These are not mutually exclusive in practice; plenty of shops run Zabbix for the bare-metal estate and Prometheus for the Kubernetes workloads side by side, which is often the honest answer rather than picking one tool for everything.
What the interviewer probes next
They will usually dig into whether the cloud VM slice should just use the same agent as everything else or be treated as a separate Prometheus-scraped island, and they will want a concrete cutover sequence rather than just a target-state recommendation.
You are asked to build a capacity trend for disk usage across 200 servers so management can plan storage purchases. What data would you collect, how would you calculate the trend, and what would trigger a purchase order?
Sample Answer
Direct answer
Collect a time series of used capacity per volume, for example daily df samples or whatever your monitoring system already stores, fit a trend line to it, and project forward to when it crosses your action threshold. You need enough history to smooth out noise (weekly patterns, one-off cleanups) but recent enough to reflect current growth, and the output should be a date, not just a percentage, because a date is what triggers a purchase order.
Structured elaboration
Collect at least daily df-style samples of used GB per volume, not just percent, since percent alone hides how many GB a jump actually represents on a large disk. Fit a trend line (ordinary least squares is enough for this) to get a growth rate in GB per day, then project forward to find the day the fit crosses your capacity ceiling.
What triggers the purchase order: pick a lead time longer than your procurement cycle. If buying and provisioning new storage takes 3 weeks, trigger the order when the trend crosses "90% full in 5 weeks," not when the disk is already at 90%.
Worked example
Executed example using 14 days of sampled disk usage on a 500 GB volume, fitting a least-squares line and projecting forward:
measured samples (GB used): [350.5, 350.4, 355.1, 357.4, 355.1, 358.0, 362.2,
363.5, 366.0, 366.7, 371.3, 373.2, 374.1, 377.7]
fitted growth rate: 2.105 GB/day (true underlying rate in this simulation was 2.0 GB/day)
fitted intercept (day 0 est.): 349.3 GB
90% full (450 GB) projected at day 47.9 from day 0
days remaining from today (day 13) to the 90% mark: 34.9 days
The fit (ordinary least squares) recovers the true 2.0 GB/day growth rate closely, 2.105 estimated, even with daily measurement noise. That is the point: a single day's jump or dip should not be read as a trend change, the line across many days is what you act on.
Trade-offs and pitfalls
- A straight-line fit assumes linear growth; a service about to onboard a large new customer or double its retention window will blow past a linear projection, so pair the trend with awareness of planned changes, not just historical data.
- Too little history (a few days) makes the slope noisy and unreliable; too much history (a year) can hide a recent acceleration by averaging it away. Recompute on a rolling window, for example the trailing 30 days, and re-evaluate weekly.
- An aggregate trend across 200 servers hides the one server about to fill up next week; you want both a fleet-wide summary for planning and a per-host projection for the "who pages tonight" question.
What the interviewer probes next
They will usually push on whether the fit is really linear or whether a single unusual day, a big import, a bulk cleanup, is quietly steering it, and on how you would roll a whole fleet of these projections into something a non-technical stakeholder can act on without reading a chart.
Unlock Full Question Bank
Get access to all 10 Observability and Monitoring Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.