Monitoring, Logging, and Observability Questions
Understanding running systems through their signals. Covers metrics, logs, and traces, instrumentation, dashboards, alerting design, and log analysis and correlation for debugging production. Emphasizes designing observability so problems are detectable and diagnosable before users are affected.
What would you monitor to know a customer-facing web service is healthy, and which of those signals would you prioritize if you could only page on a handful of them? Walk through how you'd decide what's essential versus nice-to-have.
Sample Answer
Direct answer
I'd start from user experience outward: request success rate, latency (p95/p99), and traffic or throughput are the three signals that most directly reflect whether real users are having a good or bad time right now, so those are what I'd page on if I could only pick a handful. Everything else (queue depth, resource usage, thread-pool saturation, dependency health) is valuable for diagnosis and capacity planning, but it's typically a leading indicator or a root-cause detail rather than something a first responder needs to be paged on directly.
What I'd monitor, and how I'd prioritize
| Signal | What it tells you | Prioritize for paging? |
|---|---|---|
| Error rate (4xx/5xx, by endpoint) | Are requests actually failing | Yes, core paging signal |
| Latency (p50/p95/p99) | Are successful requests still slow enough to feel broken | Yes, core paging signal |
| Traffic / throughput | Is the shape of load itself abnormal (a sudden drop can mean upstream routing broke, not that everything's fine) | Yes, core paging signal, often paired with the two above |
| Saturation (CPU, memory, connection/thread pool, queue depth) | Are we close to a resource limit that will cause the above to degrade soon | Diagnostic and leading indicator, usually not a page on its own |
| Dependency health (DB, cache, external APIs) | Is a downstream system the actual cause of our own degradation | Diagnostic, feeds root cause once paged on our own signals |
This is the familiar "four golden signals" framing (latency, traffic, errors, saturation), narrowed here to the three that most directly track user-visible harm, with saturation kept as a fast diagnostic step rather than a primary page.
Worked example: deciding what pages versus what doesn't, for a checkout service
If I could only page on a handful of signals for this service, I'd page on error rate crossing a sustained threshold (checkout failing for real users), p95/p99 latency crossing a threshold (checkout technically succeeding but painfully slow), and a sudden traffic drop (which often means something upstream, like a CDN or DNS issue, broke before requests even reach us, and a pure error-rate alert would miss it because there's no request to error on).
I'd deliberately not page on CPU or memory alone: high CPU that isn't yet causing elevated latency or errors is useful to know about, and worth watching as a leading indicator for a proactive look, but paging on it directly tends to produce alerts that fire before there's any actual user impact, which is exactly the kind of noisy, not-yet-actionable signal that trains people to ignore pages. The right response to rising CPU with no user-facing symptom yet is usually "look at this during business hours and consider scaling," not "wake someone up."
Trade-offs and pitfalls
- Paging on too many signals defeats the purpose of prioritizing at all; if everything can page, the team is back to full alert fatigue with extra steps. The "handful" constraint in the question is doing real work: it forces a genuine choice about what's essential.
- Traffic-drop alerts need a sensible baseline that accounts for real traffic patterns (day of week, time of day, known low-traffic periods); a naive fixed threshold will false-positive constantly on legitimate quiet periods.
- Saturation metrics are tempting to page on because they feel proactive, but a resource that's near its limit and staying stable isn't yet a user-facing problem. The discipline is to treat saturation as a diagnostic and capacity-planning signal, escalating to a page only once it actually starts producing errors or latency.
Product stakeholders are asking for 99.999% availability, but your historical data shows you've been running at 99.90%. How do you approach that conversation? Walk through how you'd figure out what's realistic, what closing the gap would actually cost, and how you'd present the trade-off.
Sample Answer
Direct answer: Don't argue with the target directly, translate it into downtime and cost first, then let the numbers make the case. 99.999% and 99.90% sound close as percentages but are wildly different in engineering effort, so the conversation should start with "here's what each number actually means in hours per month, and here's what closing that gap costs," not with "that's unrealistic."
Structured elaboration
Step 1: quantify the current and requested gap
- Translate both numbers into allowed downtime over a fixed period so stakeholders can compare them concretely instead of abstractly.
- Segment by user journey: does every feature need five nines, or is that only true for the checkout path while, say, the recommendations widget could tolerate far less?
Step 2: figure out what closing the gap costs
- 99.9% failures are usually single points of failure fixable with redundancy (a second AZ, a retry policy, better health checks).
- 99.999% failures are usually deep systemic issues: multi-region active-active, consensus protocol correctness, extensive chaos and failover testing, on-call structured around instant automated failover rather than human response. This is a different order of engineering investment, not an incremental tightening of the same knobs.
- Put a rough cost shape on it: infrastructure (multi-region redundancy, cross-region data replication), engineering time (redesign, extensive testing), and ongoing operational cost (more complex on-call, more failure modes to reason about).
Step 3: present the trade-off, not a refusal
- Bring a recommendation, not just a "no": propose a tiered SLO where the critical path gets closer to the ask and non-critical paths stay at a lower, still-honest target.
- Frame it as a business decision with a number attached ("closing this gap is roughly this much more engineering investment for this much less downtime"), and let the stakeholder weigh it against the actual cost of the downtime they're trying to eliminate.
Worked example
Convert both availability targets into monthly downtime, using a 30-day month:
Total minutes in 30 days=30×24×60=43,200 minutesAt 99.90% availability, allowed downtime is:
(1−0.999)×43,200=43.2 minutes per monthAt 99.999% availability, allowed downtime is:
(1−0.99999)×43,200=0.432 minutes per month≈26 seconds per monthSo the ask is to go from "43 minutes of downtime is acceptable per month" to "26 seconds is acceptable per month," a reduction by a factor of 100. That is the number that belongs in front of stakeholders: not "99.999% vs 99.90%," which reads as a small gap, but "you're asking us to reduce acceptable downtime by 100x," which reads as what it actually is.
Trade-offs & pitfalls
- Presenting only the technical difficulty without a cost estimate lets the conversation drift back into "just make it more reliable," because the stakeholder has no number to weigh against their own priorities. Always attach a cost shape, even a rough one, to the technical explanation.
- Refusing outright ("that's not realistic") without an alternative reads as obstruction. Bringing a tiered proposal (five nines on the critical path, a lower honest target elsewhere) turns the conversation from a refusal into a negotiation.
- The same negotiation skill runs in the other direction too: sometimes the ask is to relax an SLO on a non-critical feature to buy delivery speed, and the same three-step structure applies, quantify the current target, quantify what's being traded for the relaxation, and bring a recommendation rather than a bare yes or no.
- A related variant is a stakeholder who wants an undefined "real-time" latency target with no explicit number and no budget attached. The same translate-to-concrete-cost approach applies: get them to state an actual number (or propose one based on what's technically and financially reasonable), then price it the same way, because "real-time" without a number isn't a target you can build against or budget for.
What is a runbook, and what does a good one actually need to contain to be useful when someone's paged at 3am? Sketch what you'd want in one for a failed database migration.
Sample Answer
Direct answer
A runbook is a step-by-step operational document for handling a specific, known failure mode: what to check first, what commands to run, when to roll back versus push forward, and who to call if it gets worse. A good one is written so that someone half-awake at 3am who has never touched this exact system before can follow it without reconstructing context from scratch. The test of a good runbook is whether a different engineer than its author can execute it correctly under pressure.
What a runbook needs to contain
- Scope and severity: what specific failure this covers, and the severity/priority it corresponds to.
- Preconditions: what access, credentials, or tools you need before starting.
- Immediate triage steps: the first few things to check, in order, to confirm the diagnosis.
- Remediation steps: copy-pasteable commands with the expected output at each step, not prose descriptions of what to do.
- Verification: how to confirm the fix actually worked, not just that the command ran.
- Rollback path: a safe way back if remediation makes things worse, including its own preconditions (e.g. "requires a backup from the last 24 hours").
- Escalation: who to page next and when, by name/role/contact, not just "escalate if needed."
- Post-incident: where to file the incident ticket, and a note to update the runbook itself if a step was wrong or missing.
Worked example: failed database migration runbook
- Scope: prod schema migration failed mid-deploy. Severity: P1 if writes are blocked, P2 if only the migration job failed cleanly.
- Triage (first 5 minutes): tail the migration tool's log for the exact error; check
SELECT count(1) FROM pg_stat_activity WHERE state <> 'idle';to see if the migration left long-running locks; check the app's error dashboard for whether requests are actually failing yet. - Remediation, case A (migration failed cleanly, nothing partially applied): re-run the migration tool in dry-run mode first, then apply.
- Remediation, case B (partially applied, schema now inconsistent): put the app in read-only/maintenance mode to stop new writes, then decide between manually completing the migration or rolling back.
- Rollback (case B, if completing isn't safe): confirm the most recent backup timestamp, pause replication, restore with
pg_restore --clean --no-owner <backup>, then run a smoke test against a few critical read/write paths before removing maintenance mode. - Verification: re-run the app's smoke tests, spot-check row counts on affected tables against the pre-migration baseline.
- Escalation: if triage doesn't identify the cause within 10 minutes, or rollback is being considered, page the on-call DBA by name/rotation, not just "the DBA team."
- Post-incident: file the incident ticket with the timeline and root cause, and if any step above was missing or wrong, fix the runbook in the same pass as the incident writeup, not "later."
Trade-offs and pitfalls
- A runbook that's too generic ("check the logs, investigate, fix it") isn't actually a runbook, it's a checklist item pretending to be one; specificity is what makes it useful at 3am when judgment is impaired by fatigue.
- Runbooks rot: a step that references a tool or dashboard that got replaced months ago is worse than no runbook, because it wastes time and erodes trust in the whole document. Tie runbook review to any change in the system it covers, not a fixed calendar cadence alone.
- Over-indexing on "never improvise" can be as dangerous as no runbook at all: a good runbook documents when to deviate (e.g. "if replication lag exceeds a set threshold, stop and escalate instead of continuing") rather than pretending every failure mode was anticipated.
- Untested runbooks are a liability. The rollback path above should actually be exercised in a drill, not just written down, since commands like
pg_restore --cleanbehave differently depending on schema ownership and extensions that may not match what was true when the runbook was written.
You're generating terabytes of logs per day and need a long-term retention strategy. How would you think about storing older logs cheaply while still being able to search them for forensic investigations and run batch analytics over them?
Sample Answer
Split the problem into a short, fully-indexed hot tier for day-to-day operational search, and a much cheaper cold tier of compressed columnar files on object storage for the long tail, with a lightweight catalog (not a full-text index) that maps time ranges and a few coarse identifiers to the specific files a query needs. Forensic point-lookups ("find everything about this one request from eight months ago") and batch analytics ("scan a year of logs for a pattern") are different access patterns and should be served differently: the catalog gets you to the right files for a point-lookup, while a batch engine scanning the columnar files directly serves analytics, and neither needs the cold tier to be a fully-indexed search cluster.
Framework
Tiering. Hot tier (days, full search index) for active operations. Cold tier (the long retention window) as compressed, partitioned columnar files (Parquet/ORC), partitioned by date and service so both access patterns can prune to the relevant subset without scanning everything.
The catalog, not full-text indexing, is what makes cold-tier forensics tractable. Index a small, deliberately chosen set of identifiers, most usefully trace_id or request_id, to (file, partition) pointers, so a forensic point-lookup for one specific request goes: catalog lookup for the ID, straight to the handful of files that contain it, rather than a full scan of a year of data. Sizing this catalog correctly matters a lot, and naive "index every line" quickly stops paying for itself, shown below.
Batch analytics uses the same columnar files directly, via a scan engine (Spark/Trino/Athena-style) with partition pruning and predicate pushdown, since analytics workloads (aggregate over a time range, find a pattern across many requests) don't need row-level lookup, they need efficient columnar scanning, which Parquet/ORC already provide without any extra indexing.
Lifecycle automation, not manual cleanup, moves data through the tiers and eventually deletes it per retention policy, with an explicit legal-hold flag that can override the TTL for specific data under investigation or compliance hold.
Worked example
Assume 2 TB/day of raw logs, and (stated as a planning assumption to validate against real data, not a measured fact) roughly 8x compression converting to a columnar, dictionary-encoded, zstd-compressed format:
82 TB/day=0.25 TB/day=250 GB/day compressedOver a 1-year retention window, compressed footprint versus keeping raw for comparison:
250 GB×365=91,250 GB≈91.25 TB (compressed, 1 year) 2 TB×365=730 TB (raw, uncompressed, 1 year) 730/91.25=8.0×which checks out against the assumed 8x ratio, as it must (the two numbers are the same assumption expressed two ways).
Catalog sizing is where the real design decision lives. At an average log line size of 500 bytes, 2 TB/day is:
500 bytes/line2×1012 bytes=4×109 lines/day (4 billion)A naive "index every line" catalog, at a compact 40 bytes per index entry (an ID plus a file/partition pointer):
4×109×40 bytes=1.6×1011 bytes=160 GB/dayThat's 160 GB/day just for the index, against 250 GB/day for the compressed log data itself, the same order of magnitude as the data it's supposed to be a lightweight pointer into. Indexing at the individual-line grain has quietly stopped being cheap.
The fix is to index at trace_id grain instead of line grain, since many log lines share one trace_id. Assuming an average of 20 lines per trace:
An 8 GB/day catalog is a 20x reduction from the naive per-line version, and a small fraction (about 3%) of the 250 GB/day of compressed log data it points into, which is the shape a catalog should actually have: cheap relative to the data, not comparable to it. This is the concrete reasoning that separates "we added an index" from "we added an index that's actually worth what it costs."
Trade-offs and pitfalls
- Indexing at too fine a grain (as the naive per-line calculation shows) can erase most of the cost savings tiering was supposed to deliver; indexing at too coarse a grain (say, only by day and service, nothing more specific) makes forensic point-lookups require scanning entire daily partitions instead of jumping straight to the relevant files. The
trace_id-level index above is a deliberate middle ground, chosen because it matches how forensic investigations actually query (by request, not by arbitrary line), not because finer or coarser is wrong in the abstract. - The 8x compression figure is a planning assumption and needs validating against a real pilot on your actual log shape: logs with a lot of free-text/stack-trace content compress differently than uniform structured fields, and the catalog-sizing math above scales directly off whatever the real compression ratio turns out to be.
- Legal holds and lifecycle automation can conflict directly: an automated TTL-based deletion job and a compliance requirement to preserve specific records indefinitely need an explicit override mechanism (a hold flag checked before any delete), not an assumption that "we'll remember to exclude that data" during a routine cleanup job.
- Batch analytics over the cold tier is inherently higher-latency than a hot-tier query; that has to be communicated as an explicit trade-off to anyone used to interactive search, since "why is this query taking minutes instead of being instant like the hot dashboard" is a predictable point of friction if it isn't set as an expectation up front.
You inherit a dashboard with 40 panels that the on-call team has basically stopped looking at because it's too noisy to be useful during an incident. How would you go about fixing it?
Sample Answer
Direct answer
Treat this like triaging technical debt: first classify every panel by whether it maps to an actual triage decision, aggressively cut or consolidate everything that doesn't, and validate the result against real on-call engineers before calling it done, not just against your own judgment.
Approach
Step 1: audit and classify
Talk to two or three people who've actually used the dashboard during an incident. For every panel, tag it: critical (maps to a golden signal or SLO), useful for drill-down, rarely used, or dead weight with no clear owner or purpose anyone can name.
Step 2: consolidate around golden signals
Replace many single-entity panels with a small set of golden-signal panels (rate, errors, latency, saturation) plus a "top N by errors or latency" table instead of one panel per service or endpoint.
Step 3: use template variables for the long tail
Instead of one panel per service or region, add a dashboard variable and let the engineer select what to drill into. The front page stays small; context-switching happens on demand rather than by scrolling past 35 panels that don't apply to the current incident.
Step 4: redesign layout around the incident workflow
Row 1: is it broken (SLO burn, error rate, active alerts). Row 2: how broken (latency, saturation, dependency health). Row 3 and below: drill-down detail. This mirrors how someone actually works an incident, overview first, detail on demand, instead of a flat grid of 40 equally-weighted panels.
Step 5: validate, don't just ship
Run the new version past on-call for a rotation or two, or a tabletop exercise against the last few real incidents, and explicitly ask whether this panel set would have gotten them to root cause faster. Cut anything that doesn't survive that test.
Worked example
A concrete, honest instance of the consolidation pattern: suppose 15 of the 40 panels are the same latency chart repeated once per microservice. Replacing those 15 with one templated latency panel (a service selector variable) plus one "top 5 services by p99" ranked table removes 13 panels (15 down to 2) while preserving, and arguably improving, the same information, since the ranked table surfaces the worst offender automatically instead of requiring a scroll through 15 charts to spot it. The same consolidation pattern applied to a repeated error-count-per-service panel family would remove a comparable number. The exact final panel count for the full 40 depends on how much of the original sprawl is this kind of per-entity repetition versus genuinely distinct signals, which is precisely what step 1's audit is for.
Trade-offs and pitfalls
- Cutting aggressively without checking with on-call can remove a panel that's rarely used but critical for one specific failure mode, like a replication-lag panel that's silent 99% of the time but is the first thing needed during that one incident type. Validate against past incidents, not just usage frequency.
- Template variables trade a small amount of always-visible context for a click. That's usually the right trade for an overview dashboard, but a signal that's genuinely page-worthy, like overall SLO burn, should stay pinned rather than sit behind a variable.
- Common wrong turn: fixing the dashboard once without establishing an ownership or review process, which is how it reached 40 panels in the first place. Without governance, the same sprawl recurs within a year.
- Don't conflate "unused" with "unnecessary." Some panels are unused specifically because the failure mode they'd catch hasn't happened yet, not because the panel is dead weight.
Unlock Full Question Bank
Get access to all Monitoring, Logging, and Observability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.