Monitoring, Logging, and Observability Questions
Understanding running systems through their signals. Covers metrics, logs, and traces, instrumentation, dashboards, alerting design, and log analysis and correlation for debugging production. Emphasizes designing observability so problems are detectable and diagnosable before users are affected.
Design a sampling strategy for an e-commerce service that sees big traffic spikes during promotions. How would you combine cheap steady-state sampling with something that reliably captures high-latency and error traces, and how would you protect your budget when traffic suddenly spikes 10x?
Sample Answer
Run two sampling mechanisms together, not one: a cheap, deterministic head-based sample for steady-state coverage, plus an always-keep rule for anything tagged as an error or high-latency, applied either at the SDK or at a short-lived buffer in the collector. Protect the budget during a spike by making the head-sample rate dynamic, inversely proportional to observed traffic, so the absolute volume of sampled traces stays roughly constant even when request volume jumps 10x.
Framework
Layer 1: deterministic head sampling for steady-state. Hash the trace ID (not the user ID, which can be skewed or low-cardinality) and keep it if the hash falls under the current rate threshold. This is cheap (decided at the start of the request, no buffering needed) and gives uniform, reproducible coverage.
Layer 2: always-keep rules for the traces that matter most. Any trace with an error status or latency above a threshold should bypass the head-sample rate entirely, either by having the SDK tag it as "keep" once the outcome is known, or by buffering spans briefly at a collector and making the keep/drop decision once the full trace's outcome is visible (tail-based sampling). This is what guarantees you don't lose the rare, expensive traces to a low steady-state sample rate.
Layer 3: a dynamic controller that protects the budget. Monitor request rate (or trace-ingest rate) and adjust the head-sample rate so that absolute sampled volume, not the percentage, stays within budget. When traffic returns to baseline, ramp the rate back up gradually rather than snapping back, to avoid oscillation.
flowchart LR
A[Request] --> B[SDK: head sample\nhash trace_id]
A --> C[SDK: tag error/slow]
B --> D[Collector]
C --> D
D --> E{Keep?}
E -->|head-sampled in, or tagged error/slow| F[Store trace]
E -->|neither| G[Drop]
H[Dynamic controller\nwatches request rate] -.adjusts rate.-> B
Worked example
Baseline traffic: 2,000 req/sec. Promotion spike: 10x, 20,000 req/sec. Set an overall trace-write budget of 40 traces/sec (a deliberate cost decision, not derived from anything else, and the number the rest of this design has to respect).
Splitting the budget. Suppose the baseline error rate is 0.5%. At baseline, that's already 0.005×2,000=10 error traces/sec on its own. If the plan is "always keep 100% of errors" plus a head sample for everything else, the error traffic alone consumes 10 of the 40-trace/sec budget even before any head-sampled successful traffic is counted, and during an incident (when error rate rises above baseline), that share only grows. A naive 90/10 split (36 traces/sec for head-sampled traffic, 4 traces/sec reserved for errors) doesn't survive contact with this number, since baseline errors alone are 10/sec, 2.5x the reserved 4/sec. That's a genuine finding, not a footnote: "always keep all errors" is not free, and the budget has to be sized around the error volume you actually see, with its own cap and probabilistic fallback once errors themselves spike beyond that cap, rather than assumed to be a small side allocation.
Sizing the head rate correctly, and adjusting it during the spike. Say the error/latency-tagged traces get a dedicated 15 traces/sec slice (covering the observed 10/sec baseline plus headroom), leaving 25 traces/sec for head-sampled normal traffic. At baseline:
head ratebaseline=2,000 req/sec25 traces/sec=1.25%At the 10x spike, holding the same 25 traces/sec allocation requires:
head ratespike=20,000 req/sec25 traces/sec=0.125%The head rate has to drop by exactly the same 10x factor the traffic grew by, to hold absolute volume constant, which is the mechanical rule the dynamic controller implements: rate is inversely proportional to observed request rate for a fixed target volume.
Trade-offs and pitfalls
- The error-budget arithmetic above is the real trap in this design: teams reach for "just always keep every error trace" as an obviously-correct rule, and it is correct in isolation, but it has to be sized against actual error volume or it silently blows the total budget, especially during the exact incidents where error volume is elevated and you most need the traces.
- A centrally-computed dynamic rate has control-loop lag: it reacts to a spike, it doesn't prevent the first few seconds of it. Pair the central controller with a local, fast token-bucket rate limiter (a mechanism that refills a fixed-size pool of "tokens" at a steady rate and only lets a trace through if it can spend one, so it absorbs a short burst up to the pool's size and then throttles once the pool runs dry) at the SDK/collector so a sudden burst doesn't blow through the budget before the controller catches up.
- Deterministic hashing on trace ID gives uniform coverage; hashing on user ID or another business key instead can silently under-sample whatever's correlated with that key (a specific high-volume tenant, say), which defeats the "representative steady-state coverage" goal.
- Tail-based sampling (buffering until the outcome is known) is what makes "always keep errors" possible, but it costs memory and adds latency to the sampling decision; that buffer needs its own bounded size and eviction policy so it doesn't become a new failure mode under exactly the load spike it's meant to handle.
What's the difference between head-based and tail-based sampling for traces? Give a concrete situation where the extra complexity of tail-based sampling is actually worth it, and one where you'd stick with head-based.
Sample Answer
Direct answer
Head-based sampling decides whether to keep a trace at its very first span, before anything about the request's outcome is known: a cheap rule (usually a hashed trace ID compared against a threshold) picks a fixed percentage of traffic, and every downstream service just honors that one decision. Tail-based sampling holds every span of a trace until the whole thing finishes, then decides based on what actually happened (an error, a slow span, a specific status), which guarantees you keep the traces you'd actually want to debug, at the cost of buffering and routing every trace's spans to one place long enough to make that call. Use head-based when uniform, cheap coverage of "normal" traffic is what you need; use tail-based when missing a specific rare failure is unacceptable and you can afford the extra infrastructure.
Structured elaboration
How the decision differs mechanically
| Head-based | Tail-based | |
|---|---|---|
| Decision point | At the trace's first span, before the request has even resolved | After the trace completes, once its outcome is known |
| What it needs | A shared random value or deterministic hash, propagated in context to every downstream service | Buffering (or storing) every span until the trace is complete, then evaluating a policy against the assembled trace |
| Coordination across services | None beyond honoring the one propagated decision | A trace's spans usually have to reach the same collector instance to be evaluated together |
| What it can guarantee | A uniform, reproducible sample rate of overall traffic | Capture of whatever outcome you define as worth keeping (error, high latency, specific status) |
Where head-based sampling itself has variants
Fixed-rate/deterministic head sampling (hash the trace ID, keep it if the hash falls under a threshold) and reservoir sampling (maintain an unbiased random sample of a bounded size over a rolling window, useful when total trace volume is unpredictable or bursty) are both still head-based in the sense that matters here: the keep/drop decision is made without any knowledge of the outcome. That shared blind spot, not something reservoir sampling fixes, is the actual dividing line versus tail-based sampling.
Concrete situations
- Head-based is enough: a well-behaved internal health-check or low-risk read endpoint with a near-zero error rate, where you mainly want representative latency percentiles across normal traffic and don't need to guarantee capture of any specific rare event. A flat, cheap deterministic sample gives you that with no buffering infrastructure at all.
- Tail-based earns its complexity: a payment-authorization service where a specific downstream provider's rejection happens on well under 1% of requests, but every one of those traces matters for debugging a customer-facing failure. At a low head-sample rate, most of those rare error traces simply never get kept; tail-based sampling, evaluated once the trace's outcome (error status) is known, guarantees you capture all of them regardless of how rare they are.
Very high-throughput production services
At extreme volume, pure tail-based sampling (buffering every span of every trace until it completes) becomes expensive in its own right, since buffering cost scales with total traffic, not with the sampled fraction. The common production pattern, and what the OpenTelemetry Collector's tail-sampling processor implements, is a hybrid: a low fixed head-sample rate for baseline, representative coverage of normal traffic, plus a tail-based always-keep policy layered on top for errors and latency outliers. That keeps the guaranteed-capture property where it matters most while bounding the buffering cost to the traffic you'd realistically want full visibility into.
Worked example
Take a service where a rare downstream error occurs independently on roughly 1 in 2,000 requests, and a head-based sample rate of r=1%. For any single occurrence, the chance a head-based sample happens to keep that specific trace is just r=1%, since the sampling decision is made before the error even happens and has no idea it's about to occur. Over k independent occurrences of this same rare error, the chance that head-based sampling misses every single one of them is:
P(miss all k)=(1−r)kFor k=20 occurrences of the error:
P(miss all 20)=(1−0.01)20=0.9920≈0.818So even after 20 separate occurrences of this specific rare failure, a 1% head-based sample rate still has roughly an 82% chance of having captured none of them, purely because the sampling decision never looked at the outcome. Tail-based sampling with an always-keep-errors policy makes that same probability exactly 0%, by construction, since its decision is made only after the error is already known to have happened.
Trade-offs and pitfalls
- Tail-based sampling's guarantee only holds for whatever you defined the keep policy around (error status, a latency threshold). It won't catch an outcome you didn't think to write a rule for, so the policy needs revisiting as new failure modes show up.
- Buffering spans for a tail-based decision adds memory pressure and latency to the sampling pipeline, and requires routing a trace's spans consistently to one collector, which is real infrastructure a pure head-based setup doesn't need.
- Aggressive head-based sampling (a low fixed rate) degrades any metric computed only from the sampled traces, like latency percentiles built from kept traces alone, since a small sample size increases variance on tail percentiles specifically even though the sample is representative on average; prefer computing percentiles from a metrics pipeline that sees 100% of requests, not from sampled traces, when precision matters.
- Deciding sampling deterministically from the trace ID, rather than independently per span, is what makes a trace's sampling decision "sticky": every span belonging to the same trace ID gets the same keep/drop outcome. Losing that property (say, sampling independently at each service) produces incomplete, unusable traces where some spans of a request are kept and others silently dropped.
The same service needs a dashboard for the on-call engineering team and a very different one for a business stakeholder. How would the metrics, visualization choices, time ranges, and alerting differ between the two, and what's a panel or alert you'd deliberately leave off one of them?
Sample Answer
Direct answer
The SRE dashboard and the business dashboard should be built from the same underlying events but answer different questions: the SRE dashboard needs to answer "what's broken right now and where," so it's dense, high-resolution, and short-windowed; the business dashboard needs to answer "is the product healthy from a user or revenue standpoint," so it's sparse, trend-oriented, and longer-windowed. Alerting follows the same split: SRE alerts are operational and verbose, business alerts are outcome-level and rare.
How the two dashboards differ
| Dimension | SRE / on-call dashboard | Business stakeholder dashboard |
|---|---|---|
| Metrics | Per-endpoint error rate, p50/p95/p99 latency, saturation (CPU, memory, queue depth), dependency health | Conversion rate, revenue per hour, active users, overall uptime percentage |
| Visualization | Time series with drill-down, heatmaps for latency, per-host/region breakdowns | Single big numbers, simple trend lines, week-over-week or month-over-month comparisons |
| Time range | Minutes to hours, short enough to catch and react to an active incident | Days to months, long enough to see real trend versus daily noise |
| Alerting | High verbosity, pages on technical thresholds (e.g. p99 over 2s for 5 minutes), links to runbooks | Low verbosity, notifies on business impact only (e.g. conversion down sharply versus baseline), plain-language summary |
A flow from shared telemetry to two different views
flowchart LR
A[Raw telemetry: request events] --> B[High-resolution metrics store]
A --> C[Business event stream]
B --> D[SRE dashboard]
C --> E[Hourly and daily rollups]
E --> F[Business dashboard]
D -.drill-down link.-> F
Both dashboards should compute from the same raw events (not two independently instrumented pipelines) so that when a stakeholder asks "why did conversion drop," the SRE dashboard's drill-down actually explains it, instead of the two numbers just disagreeing.
What I'd deliberately leave off each
Off the SRE dashboard: revenue and business-outcome panels. They're not actionable for an on-call engineer during an incident and just add noise to a screen that needs to be scanned in seconds.
Off the business dashboard: raw error-rate and per-host technical panels, and just as important, anything that exposes PII. A stakeholder-facing dashboard is more likely to be screen-shared in a broader meeting or exported into a slide deck than an internal SRE tool, so panels built directly from raw log or event data (which might carry user IDs, emails, or IP addresses) need to be aggregated or scrubbed before they reach that surface, not just visually simplified.
Trade-offs and pitfalls
- Building genuinely separate pipelines for the two dashboards (instead of one source feeding both) is the most common mistake: it's less work up front but guarantees the two views eventually disagree, which destroys trust in both.
- SRE alerts are tuned for recall (better to page on a false positive than miss a real incident); business alerts are tuned for precision (a business stakeholder paged on noise will stop trusting the channel fast). Using the same alerting sensitivity for both is a common failure mode.
- A dashboard trying to serve both audiences at once usually serves neither well: too dense for a business reader, too shallow for on-call triage. Linking between them is good; merging them into one view usually isn't.
Binary up/down SLIs miss a lot of what users actually experience as degraded service. How would you design an SLI that captures 'the product feels slow or broken' rather than just 'the product is down', and how would you validate that it actually tracks real user pain instead of just moving a threshold around?
Sample Answer
Direct answer
Replace the binary "is it up" check with a request-level ratio SLI: define a "good event" as one that meets both a correctness condition (no error) and a latency condition (under a threshold chosen to reflect when users actually start to feel pain), then track the fraction of good events over valid events. The threshold can't be picked by intuition alone: it has to be validated by correlating it against an independent signal of real user frustration, otherwise you've just moved the goalposts without proving they mean anything.
Structured elaboration
Defining the SLI (an SLI, service level indicator, is the specific measured ratio; an SLO is the target you set for it)
SLI=#{valid events}#{good events},good event⟺(status=2xx)∧(latency≤T)This is deliberately simple and auditable: anyone can look at the raw event stream and check whether an event counted as good. Composite scores that blend many signals with hand-tuned weights are harder to defend later ("why is latency weighted 0.4 and errors 0.6?") and harder to debug when the number moves.
Choosing the threshold T, not by feel
Pick T by correlating latency buckets against an independent pain signal you did not use to build the SLI itself: abandonment rate, support-ticket rate, or a conversion drop. If the same team that owns the SLI also picks T from "what feels slow," they will unconsciously tune it to hit whatever SLO they were already targeting, which is the exact "just moving a threshold around" failure the question is pointing at.
Validating it actually tracks pain
Bucket historical requests by latency, compute the pain-signal rate in each bucket, and check whether latency and pain move together (a strong positive correlation) rather than being flat or noisy. If they don't correlate, the "feels slow" complaint isn't really about server-side latency and a different SLI (client-side rendering, third-party script blocking, etc.) is needed instead.
Avoiding self-referential validation
The threshold must be re-validated periodically against fresh data (traffic mix, device mix, and user expectations all drift), and ideally against a held-out set of incidents the SLI wasn't tuned on: does the SLI actually dip during known bad periods, and stay flat during known-fine periods?
Worked example
Suppose five latency buckets, each with an observed support-ticket rate for that band (this is a stated, self-contained example, not measured production data; the point is to show the procedure for validating a threshold, which is what the interview is testing):
| Latency bucket midpoint (ms) | Ticket rate (%) |
|---|---|
| 100 | 0.2 |
| 300 | 0.5 |
| 600 | 1.8 |
| 1200 | 4.9 |
| 2500 | 11.0 |
Computing the Pearson correlation between latency and ticket rate:
r=∑i(xi−xˉ)2∑i(yi−yˉ)2∑i(xi−xˉ)(yi−yˉ)With xˉ=940 and yˉ=3.68, this evaluates to r≈0.998: latency and complaint rate move together almost perfectly in this example, which is what would justify treating latency as a valid proxy for pain here.
The biggest jump in ticket rate happens between the 300ms and 600ms buckets (0.5% to 1.8%, more than 3x). That's the elbow: it's evidence for setting T somewhere in the 400-500ms range, rather than picking a round number like 1000ms because it "sounds reasonable." This is exactly the kind of derivation a validation step should produce; a real rollout would run this against actual production data, not five hand-picked points.
Trade-offs & pitfalls
- A single threshold applied globally can hide cohort-specific pain (mobile users on slow networks feel a 600ms response very differently from users on fast broadband); segmenting the SLI by cohort catches this at the cost of more series to track and alert on.
- Correlation is not causation: a rising ticket rate at high latency could be caused by a third factor (e.g., the same deploy that slowed requests also broke a feature). Treat the correlation as evidence to investigate, not as final proof.
- If the "independent" pain signal is itself derived from the same infrastructure being measured (e.g., using error logs from the same service as the validation signal), the validation isn't actually independent and the trap the question warns about resurfaces.
- Re-tuning T too often erodes trust in the SLO; re-validate on a fixed cadence (e.g., quarterly or after a major traffic-mix shift) rather than whenever the number is inconvenient.
You've got an anomaly-detection alert on a per-entity rate, think per-customer or per-tenant request volume, and it fires constantly for your smallest customers while staying quiet for your largest ones. How would you redesign the detection so it's sensitive to real anomalies at both ends of that size range?
Sample Answer
Direct answer
A flat percentage threshold ("alert if volume drops 30%") is the bug: it treats a small customer's normal statistical noise as equally meaningful as a large customer's genuine anomaly, when the two have very different natural variability. The fix is to scale the sensitivity to the expected variance at each customer's volume, so the alert reacts to how surprising a change is, not just how big it looks as a percentage.
Structured elaboration
Why static percentage thresholds fail across a size range
Count-based data like requests-per-hour naturally fluctuates more, in relative terms, the smaller the volume is. A customer normally sending 50 requests/hour swinging to 35 in a given hour is well within ordinary noise; a customer normally sending 100,000 requests/hour dropping by the same 30% is a massive, almost certainly real event. A single relative threshold can't tell those apart, because it was never designed to account for volume-dependent variance in the first place.
Redesigning the detection
- Scale sensitivity to expected variance, not to a flat percentage. For count data, a common and reasonably good approximation is that variance is close to the mean (a Poisson-like assumption), which means the standard deviation grows with the square root of volume, not linearly with it. Comparing a change in units of that standard deviation (a z-score) instead of raw percentage automatically adapts sensitivity to volume.
- Add a minimum absolute floor. Even with a z-score approach, an extremely small customer going from 1 event to 0 can register as a huge relative or statistical swing that isn't actually meaningful; require both a statistical threshold AND a minimum absolute change before alerting.
- Require persistence. A single anomalous window can be a fluke; requiring the anomaly to hold across a couple of consecutive windows filters out one-off noise for both ends of the size range.
Handling non-Poisson reality (beyond what most interviews require)
Real traffic has diurnal and weekly seasonality and is usually over-dispersed relative to a pure Poisson model (its true variance is larger than its mean), so a production system typically layers a seasonal baseline or an EWMA-smoothed rolling variance, or a more robust statistic like the median absolute deviation, on top of the basic idea rather than using raw Poisson variance directly. The core insight, scaling sensitivity to expected variance instead of using a flat percentage, is the part worth leading with; the seasonal/robust refinements are worth mentioning as the natural next step, not the starting point.
Worked example
Assume request counts are approximately Poisson-distributed, so standard deviation is approximately the square root of the mean, and compare a small customer (mean 50 requests/hour) against a large one (mean 100,000 requests/hour) under both the old 30% static threshold and a redesigned z-score-based one.
σ≈μ,z=σμ−xA 30% drop for the small customer (50 to 35) versus the large customer (100,000 to 70,000), measured in standard deviations:
zsmall=500.3×50=7.0715≈2.12,zlarge=100,0000.3×100,000=316.230,000≈94.9That gap, 2.12 versus 94.9 standard deviations for the same "30% drop," is exactly the miscalibration in the question: for the small customer, a 30% swing is only mildly unusual (roughly a 2-sigma event, which happens by chance often enough to explain the constant noisy alerts), while for the large customer it's a statistical impossibility under normal variation, meaning a real 30%-threshold rule was tuned to something that's noise-level for small customers and would almost never even get the chance to fire for large ones since a genuinely damaging but smaller drop wouldn't reach 30%.
Now redesign using a fixed z-score threshold of 4 (a much rarer, roughly-equally-surprising event at both ends) and translate it back into request counts:
xsmall=μsmall−4σsmall=50−4(7.07)≈21.7 xlarge=μlarge−4σlarge=100,000−4(316.2)≈98,735(a 1.27% drop)So the redesigned rule fires for the small customer only once volume drops below about 22 (a genuinely rare event, not routine noise), and for the large customer it fires on a drop of just 1.27%, far more sensitive than the old flat 30% rule ever was, and appropriately so, since a 1.27% drop at that volume is exactly as statistically surprising as the small customer's much larger relative swing.
Trade-offs and pitfalls
The minimum absolute floor is not optional even after adding the statistical scaling: a customer going from 2 events to 0 can register as an enormous z-score purely because the Poisson approximation breaks down at very low counts, so pair the statistical threshold with a hard floor like "don't alert unless the absolute drop is at least N events." A second pitfall is trusting the pure Poisson assumption too far into production: real traffic has daily/weekly seasonality that a naive Poisson model doesn't capture, and treating a predictable Monday-morning dip as anomalous will just recreate the noise problem in a new form, so the production version needs a seasonal or rolling baseline underneath the same z-score logic. Finally, explainability matters more here than raw model sophistication; an on-call engineer needs to be able to look at why an alert fired and reconstruct the reasoning, which favors a transparent statistical approach like this over an opaque ML anomaly detector the team can't debug at 3am.
Unlock Full Question Bank
Get access to all Monitoring, Logging, and Observability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.