Monitoring, Logging, and Observability Questions
Understanding running systems through their signals. Covers metrics, logs, and traces, instrumentation, dashboards, alerting design, and log analysis and correlation for debugging production. Emphasizes designing observability so problems are detectable and diagnosable before users are affected.
You've got a request path that goes through three services in sequence, each with its own availability target. Users only care whether the whole request succeeded. How would you think about the end-to-end SLO, and how would you split the error budget across the teams that own those three services?
Sample Answer
Direct answer: For a serial dependency chain, the end-to-end availability is the product of each service's individual availability, not a simple average, because the request only succeeds if every hop succeeds. Once you have that end-to-end number, split the resulting error budget across the owning teams proportional to each service's own contribution to the total failure rate, with a small shared pool held back for failures that don't cleanly attribute to one team.
Structured elaboration
Why multiplication, not averaging
- If service A, B, and C are called in sequence and the user needs all three to succeed, the probability all three succeed is the product of their individual success probabilities. This is the standard independent-events multiplication rule, and it's the right model whenever failures are (approximately) independent across the three services, meaning one service failing doesn't itself cause another to fail via a shared cause.
Allocating the resulting budget
- Compute each service's own error contribution, ei=1−Si, then allocate the end-to-end budget proportionally: a service with a larger individual error rate gets a larger slice of the end-to-end budget, since it's naturally consuming more of it.
- Hold back a small shared pool (a common starting point is around 10% of the end-to-end budget) for incidents that don't have a single clear owner, a shared network layer, a common auth service, or an incident where root cause is still being investigated when the budget starts burning.
- Track each team's consumption as a ratio against their own allocation, not as a raw count, so a team with a naturally smaller allocation isn't unfairly flagged for consuming the same absolute number of failures as a team with a larger one.
Operational policy
- Give each team a burn-rate alert against their own slice, so a team gets paged and can react before the end-to-end SLO itself is at risk, catching problems at the source rather than only at the aggregate.
- If the end-to-end budget is burning but no single team's individual slice shows an obvious spike, that's the signal to draw from the shared pool and open a cross-team investigation rather than waiting for one team's dashboard to point at itself.
Worked example
Assume three services in sequence with individual monthly availabilities:
SA=0.999,SB=0.9995,SC=0.9999End-to-end availability:
Se=SA×SB×SC=0.999×0.9995×0.9999=0.99840065End-to-end error budget:
Ee=1−Se=0.00159935(≈0.15994%)Individual error contributions:
eA=1−SA=0.001,eB=1−SB=0.0005,eC=1−SC=0.0001 eA+eB+eC=0.0016Note that eA+eB+eC=0.0016 is extremely close to the exact Ee=0.00159935, this is the standard small-error approximation: for small individual error rates, the product's complement is very close to the sum of the individual complements, which is exactly why teams commonly reason about "additive" error budgets even though the underlying math is multiplicative.
Allocating the (rounded) 0.0016 end-to-end budget proportionally by contribution, after holding back a 10% shared pool:
shared pool=0.10×0.0016=0.00016 allocatable=0.0016−0.00016=0.00144Proportional shares (by each service's fraction of total individual error: A is 0.001/0.0016=62.5%, B is 0.0005/0.0016=31.25%, C is 0.0001/0.0016=6.25%):
allocA=0.625×0.00144=0.0009,allocB=0.3125×0.00144=0.00045,allocC=0.0625×0.00144=0.00009Service A, with the highest individual error rate, gets the largest allocation (0.09% of requests), which matches intuition: the least-reliable link in the chain gets the most room, and it's also the link that should be under the most reliability pressure to shrink its own eA over time.
Trade-offs & pitfalls
- Proportional allocation by raw error contribution can feel unfair to a team whose service is used by many other chains besides this one, their "share" here doesn't reflect their reliability work elsewhere. Consider whether allocation should be per-chain or aggregated across all the chains a service participates in.
- The independence assumption behind straight multiplication breaks down for correlated failures (a shared database, a shared network path, a common auth service all three depend on). A failure there hits all three services simultaneously, and treating it as three independent budget hits overstates how "used up" the chain's reliability actually is versus how concentrated the actual root cause is. This is exactly what the shared pool exists to absorb.
- A hard proportional split with no floor can effectively give a very reliable service (eC small) almost no budget at all, meaning even one legitimate incident consumes a huge fraction of their tiny allocation. A minimum floor per team avoids penalizing the already-most-reliable service disproportionately.
What is alert fatigue, and how would you go about preventing it on a team you're leading?
Sample Answer
Direct answer
Alert fatigue is what happens when on-call engineers get so many low-value, noisy, or duplicate pages that they start treating all alerts as probably-not-real, including the ones that matter. It's a trust problem as much as a technical one: once someone has been paged repeatedly in a night for something that turned out to be nothing, the next page, which might be the real incident, gets a slower, more skeptical response.
How I'd prevent it on a team I'm leading
- Deduplication and grouping: alerts that share a root cause (same service, same error type) should collapse into a single incident with a count, not fire a separate page per occurrence. This is usually a config change in the alerting tool (fingerprinting by service and error signature) rather than a code change.
- Severity tuning tied to required response time: not every alert deserves a page. A three-tier split (page now, notify during business hours, dashboard-only) forces every new alert to justify why it needs to interrupt someone's sleep.
- Actionable-by-default policy: no new paging alert ships without a linked runbook and a clear "what to check first." An alert with no next step is a dashboard panel that accidentally has a pager attached.
- Automated remediation for known, safe, repeatable fixes: if the same alert reliably resolves by restarting a stuck worker or clearing a queue, and that action is safe and idempotent, automate it and only page if the automated fix fails.
- A regular noise review: periodically look at which alerts fired most often and whether they led to real action; alerts that never lead to action get tuned or removed, not left running indefinitely out of habit.
Worked example
Suppose a team's on-call rotation is getting paged for "queue depth over 100" on a background job processor, firing several times a week, always self-resolving within a few minutes without anyone doing anything. Applying the framework above: first, check whether these spikes line up with a predictable traffic pattern (a nightly batch job, say) and if so, either raise the threshold above that expected peak or add a time-of-day exception. Second, if the queue really can back up unpredictably but always self-resolves within a known window without intervention, the alert should require a longer sustain window (e.g. "queue depth over 100 for 15 minutes") so it only fires when it isn't going to resolve on its own. Third, if manual intervention when it does page is always the same action (scale up worker count), that's a strong automated-remediation candidate: auto-scale on the same threshold, and only page if depth is still elevated after the auto-scale has had time to take effect.
Trade-offs and pitfalls
- Automated remediation without an audit trail or human confirmation for higher-severity cases can turn a noisy-alert problem into a silent-failure problem: the system "fixes" itself repeatedly while masking a root cause that's getting worse.
- Tuning thresholds purely to reduce page volume, without checking against real past incidents, risks quietly increasing false negatives; the goal is signal-to-noise, not just fewer pages.
- Alert fatigue prevention is not a one-time project. It needs an ongoing review cadence, because new alerts get added faster than old noisy ones get cleaned up if nobody owns the process.
In a long-lived system, how do you evolve a structured logging or metrics schema over time, for example adding a new field or changing what a field means, without breaking dashboards, alerts, and tooling that depend on the old schema?
Sample Answer
Direct answer
Default to additive-only changes (new optional fields with sane defaults), never silently repurpose an existing field's name or meaning, and when the meaning genuinely has to change, introduce it as a new versioned field and dual-emit both the old and new during a defined deprecation window so every consumer (dashboards, alerts, downstream jobs) has time to migrate before the old one disappears.
Structured elaboration
Additive changes are the default and the cheap case
Adding a brand-new field with a sensible default (or simply absent, if consumers already tolerate unknown fields) is safe: existing dashboards and alerts that don't reference it are unaffected, and new tooling can start using it immediately. Most schema evolution should fit this case; if it doesn't, that's a signal the change is more than "add a field."
Never repurpose a field in place
Changing what an existing field means (e.g., a latency field that used to be measured in milliseconds and is now measured in microseconds, keeping the same name) is the most dangerous kind of change, because it fails silently: old dashboards keep running the same query and now show numbers that are wrong by a constant factor, with no error to alert anyone. A rename or unit change should always get a new field name (latency_ms retired in favor of latency_us, both emitted for a transition period), never an in-place redefinition.
Version the schema explicitly
Tag every emitted record with a schema_version. Consumers that need to branch on shape (a downstream parser, a strict dashboard query) can check the version rather than guessing from field presence. This also gives you a clean place to document exactly which version introduced which change.
Deprecation as a process, not an event
- Announce the field's replacement and the planned sunset date.
- Dual-emit: write both the old and new field for a fixed window.
- Track actual usage of the old field (query logs, dashboard/alert definitions referencing it) to confirm consumers have migrated, not just assume they have.
- Only stop emitting the old field once usage has genuinely dropped to zero (or the sunset date passes and remaining consumers have been explicitly notified they'll break).
Testing the transition
Contract tests (automated checks that a producer's output still satisfies what a known consumer expects) and shadow validation (running the old and new emission side by side and diffing the derived metrics they produce) catch the case where the "safe" additive change turns out to interact badly with an existing aggregation, before it reaches production dashboards.
Worked example
A service currently emits {"latency": 245, ...} where latency is milliseconds, and the team wants to switch to microsecond precision.
Wrong approach (in-place redefinition): change the emitter to write {"latency": 245000, ...} under the same field name. A dashboard panel computing avg(latency) over the last hour now silently reports a number 1000x larger with zero errors or warnings; anyone glancing at the dashboard sees "avg latency: 245000ms" and either panics or, worse, doesn't notice because the panel has no sanity bound configured.
Correct approach: add latency_us alongside the existing latency field, dual-emit both for a stated transition window (e.g., until every dashboard query referencing latency has been rewritten to use latency_us, confirmed by grepping the dashboard/alert config repository for the old field name), then drop latency only after that grep returns zero references.
The key diagnostic in this example: the failure mode is not "the pipeline throws an error," it's "the pipeline keeps running and produces a wrong number that looks plausible." That's why additive-with-a-new-name is the default, not an optional extra step.
| Strategy | Backward compat risk | Consumer effort required | When to use |
|---|---|---|---|
| Additive field, new name | None | None (opt-in) | Default choice for any new signal or unit/meaning change |
| Field deprecation (dual-emit then drop) | Low, if the window is long enough and usage is tracked | Must update queries before sunset | Retiring a field that's being replaced |
| In-place semantic change (same name, new meaning) | High: silent, no error | None until someone notices wrong numbers | Avoid; only defensible for a field with zero known consumers |
Trade-offs & pitfalls
- Dual-emitting indefinitely accumulates cost and confusion; every deprecation needs an explicit sunset date, not an open-ended "eventually."
- Tracking actual field usage (rather than assuming consumers migrated because you announced it) is the step most teams skip, and it's exactly the step that prevents a surprise outage when the old field is finally dropped.
- Additive changes still need CI-enforced schema compatibility checks (backward/forward compatibility validation), because "just add a field" can still break a strict consumer that rejects unknown fields.
- A silent semantic change is strictly worse than a loud break: a query that errors gets noticed and fixed; a query that keeps returning a plausible-looking wrong number can go unnoticed for months.
You're generating terabytes of logs per day and need a long-term retention strategy. How would you think about storing older logs cheaply while still being able to search them for forensic investigations and run batch analytics over them?
Sample Answer
Split the problem into a short, fully-indexed hot tier for day-to-day operational search, and a much cheaper cold tier of compressed columnar files on object storage for the long tail, with a lightweight catalog (not a full-text index) that maps time ranges and a few coarse identifiers to the specific files a query needs. Forensic point-lookups ("find everything about this one request from eight months ago") and batch analytics ("scan a year of logs for a pattern") are different access patterns and should be served differently: the catalog gets you to the right files for a point-lookup, while a batch engine scanning the columnar files directly serves analytics, and neither needs the cold tier to be a fully-indexed search cluster.
Framework
Tiering. Hot tier (days, full search index) for active operations. Cold tier (the long retention window) as compressed, partitioned columnar files (Parquet/ORC), partitioned by date and service so both access patterns can prune to the relevant subset without scanning everything.
The catalog, not full-text indexing, is what makes cold-tier forensics tractable. Index a small, deliberately chosen set of identifiers, most usefully trace_id or request_id, to (file, partition) pointers, so a forensic point-lookup for one specific request goes: catalog lookup for the ID, straight to the handful of files that contain it, rather than a full scan of a year of data. Sizing this catalog correctly matters a lot, and naive "index every line" quickly stops paying for itself, shown below.
Batch analytics uses the same columnar files directly, via a scan engine (Spark/Trino/Athena-style) with partition pruning and predicate pushdown, since analytics workloads (aggregate over a time range, find a pattern across many requests) don't need row-level lookup, they need efficient columnar scanning, which Parquet/ORC already provide without any extra indexing.
Lifecycle automation, not manual cleanup, moves data through the tiers and eventually deletes it per retention policy, with an explicit legal-hold flag that can override the TTL for specific data under investigation or compliance hold.
Worked example
Assume 2 TB/day of raw logs, and (stated as a planning assumption to validate against real data, not a measured fact) roughly 8x compression converting to a columnar, dictionary-encoded, zstd-compressed format:
82 TB/day=0.25 TB/day=250 GB/day compressedOver a 1-year retention window, compressed footprint versus keeping raw for comparison:
250 GB×365=91,250 GB≈91.25 TB (compressed, 1 year) 2 TB×365=730 TB (raw, uncompressed, 1 year) 730/91.25=8.0×which checks out against the assumed 8x ratio, as it must (the two numbers are the same assumption expressed two ways).
Catalog sizing is where the real design decision lives. At an average log line size of 500 bytes, 2 TB/day is:
500 bytes/line2×1012 bytes=4×109 lines/day (4 billion)A naive "index every line" catalog, at a compact 40 bytes per index entry (an ID plus a file/partition pointer):
4×109×40 bytes=1.6×1011 bytes=160 GB/dayThat's 160 GB/day just for the index, against 250 GB/day for the compressed log data itself, the same order of magnitude as the data it's supposed to be a lightweight pointer into. Indexing at the individual-line grain has quietly stopped being cheap.
The fix is to index at trace_id grain instead of line grain, since many log lines share one trace_id. Assuming an average of 20 lines per trace:
An 8 GB/day catalog is a 20x reduction from the naive per-line version, and a small fraction (about 3%) of the 250 GB/day of compressed log data it points into, which is the shape a catalog should actually have: cheap relative to the data, not comparable to it. This is the concrete reasoning that separates "we added an index" from "we added an index that's actually worth what it costs."
Trade-offs and pitfalls
- Indexing at too fine a grain (as the naive per-line calculation shows) can erase most of the cost savings tiering was supposed to deliver; indexing at too coarse a grain (say, only by day and service, nothing more specific) makes forensic point-lookups require scanning entire daily partitions instead of jumping straight to the relevant files. The
trace_id-level index above is a deliberate middle ground, chosen because it matches how forensic investigations actually query (by request, not by arbitrary line), not because finer or coarser is wrong in the abstract. - The 8x compression figure is a planning assumption and needs validating against a real pilot on your actual log shape: logs with a lot of free-text/stack-trace content compress differently than uniform structured fields, and the catalog-sizing math above scales directly off whatever the real compression ratio turns out to be.
- Legal holds and lifecycle automation can conflict directly: an automated TTL-based deletion job and a compliance requirement to preserve specific records indefinitely need an explicit override mechanism (a hold flag checked before any delete), not an assumption that "we'll remember to exclude that data" during a routine cleanup job.
- Batch analytics over the cold tier is inherently higher-latency than a hot-tier query; that has to be communicated as an explicit trade-off to anyone used to interactive search, since "why is this query taking minutes instead of being instant like the hot dashboard" is a predictable point of friction if it isn't set as an expectation up front.
What is metric cardinality, and why can high-cardinality labels be dangerous for a metrics backend like Prometheus? What are a few concrete strategies you would use to keep cardinality under control while still preserving useful signal?
Sample Answer
Direct answer
Cardinality is the number of distinct time series a metric produces, which is the metric name multiplied by every unique combination of the label values attached to it. It's dangerous for a backend like Prometheus because each unique series consumes its own memory and disk in the time-series database regardless of how rarely it's queried, so a single label that takes millions of unique values can turn one metric definition into millions of series and exhaust the backend.
Why this happens mechanically
Prometheus and similar time-series databases store each unique series (metric name plus label set) as its own indexed chunk in memory and on disk. Adding a label isn't "one more column," it's a multiplier: a metric with three labels each taking 10 values has up to 1,000 series. Add a fourth label with 100,000 unique values and that multiplies to 100 million.
The classic mistake
The two labels people most often add without realizing the cost are user_id and request_id, because from an application-code perspective they feel like natural dimensions ("I want to see this metric per user"). From the time-series database's perspective, though, they're unbounded, ever-growing label values that are rarely reused, so every request effectively creates a permanent new series that's queried exactly once, if ever.
Strategies to keep cardinality under control while preserving signal
- Push high-cardinality dimensions to logs or traces instead of metrics. A metric answers "how many, how fast, in aggregate," while a log line or trace span is where you attach the specific
user_idorrequest_idfor per-entity debugging. - Use bounded, reusable label values (customer tier, region, status code) instead of unbounded identifiers.
- Pre-aggregate at the instrumentation point, emitting a counter bucketed by tier or cohort rather than by raw ID.
- Enforce it in CI: reject metric definitions with labels that don't come from a small, known enum, before they reach production.
- Monitor cardinality itself as a metric (series count per metric name) and alert on unexpected growth, so a new high-cardinality label gets caught within hours, not after the backend is already under memory pressure.
Worked example
Assume a request-latency histogram with labels method (4 values), status_code (6 values), and endpoint (20 values), where each observation also expands into multiple series per label combination (say 12 latency buckets plus _sum and _count, 14 series total):
Now add user_id as a label, with 100,000 active users:
That's the mechanism: adding one seemingly-reasonable label multiplies series count by its cardinality, turning a perfectly manageable roughly 6,700-series metric into a 672-million-series one, the kind of number that exhausts a Prometheus instance's memory well before it's ever queried usefully.
Trade-offs and pitfalls
- There's a real tension between wanting per-customer visibility and the fact that metrics can't afford per-customer labels at scale. The resolution is almost always to split the concern: aggregated metrics for alerting and dashboards, logs or traces (or a purpose-built analytics store) for per-entity investigation, not forcing one system to do both jobs.
- Hashing or bucketing an identifier reduces cardinality but also destroys the ability to look up a specific entity later. Only do this if coarse cohort visibility is genuinely what's needed, not entity-level debugging.
- Common wrong turn: discovering a cardinality problem in production and reflexively dropping every recently-added label, losing signal that was actually useful (like
status_code), instead of first identifying which specific label is the unbounded one.
Unlock Full Question Bank
Get access to all Monitoring, Logging, and Observability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.