Monitoring, Logging, and Observability Questions
Understanding running systems through their signals. Covers metrics, logs, and traces, instrumentation, dashboards, alerting design, and log analysis and correlation for debugging production. Emphasizes designing observability so problems are detectable and diagnosable before users are affected.
Implement a small decorator (in Python or a language of your choice) that generates a correlation ID for each incoming HTTP request, attaches it to the request and response headers, and makes it available to any structured logs written while handling that request.
Sample Answer
Wrap the handler so it reads an inbound correlation-id header if present (otherwise generates one), stores it somewhere every log call in that request can reach without threading it through every function signature, and stamps it back onto the response headers before returning. In Python, contextvars.ContextVar is the right tool for the "reach it from anywhere" part, because it's isolated per async task and per thread, unlike a plain module-level global.
Implementation
import contextvars
import functools
import logging
import uuid
_correlation_id: contextvars.ContextVar[str] = contextvars.ContextVar(
"correlation_id", default="-"
)
HEADER_NAME = "X-Correlation-ID"
class _CorrelationIdFilter(logging.Filter):
"""Injects the current request's correlation id into every log record."""
def filter(self, record: logging.LogRecord) -> bool:
record.correlation_id = _correlation_id.get()
return True
def configure_logging() -> logging.Logger:
logger = logging.getLogger("app")
logger.setLevel(logging.INFO)
handler = logging.StreamHandler()
handler.setFormatter(
logging.Formatter("%(levelname)s corr=%(correlation_id)s %(message)s")
)
handler.addFilter(_CorrelationIdFilter())
logger.handlers = [handler]
logger.propagate = False
return logger
logger = configure_logging()
def with_correlation_id(handler):
"""Decorator for a request handler with signature handler(request) -> response.
request/response are dict-like objects with a "headers" dict, matching the
shape used by most lightweight Python web frameworks (or an adapter around
a framework's native request/response objects).
"""
@functools.wraps(handler)
def wrapper(request):
incoming = request.get("headers", {}).get(HEADER_NAME)
corr_id = incoming if incoming else str(uuid.uuid4())
token = _correlation_id.set(corr_id)
try:
response = handler(request)
response.setdefault("headers", {})[HEADER_NAME] = corr_id
return response
finally:
_correlation_id.reset(token)
return wrapper
@with_correlation_id
def create_order(request):
logger.info("received request path=%s", request["path"])
logger.info("order validated")
return {"status": 201, "body": {"order_id": "o_1"}}
Key points
_correlation_idis aContextVar, not a plain global: each async task (or thread, for sync frameworks) gets its own isolated value, so concurrent in-flight requests never see each other's correlation id.- The logging filter (
_CorrelationIdFilter) reads the contextvar at log-emit time, which is what makes it "available to any structured log written while handling that request" without changing everylogger.info(...)call site to pass the id explicitly. - If the request already carries
X-Correlation-ID(it came from an upstream service that already started a trace), the decorator reuses it instead of minting a new one, so the id stays consistent across the whole call chain rather than getting reset at every hop. token = _correlation_id.set(...)plus_correlation_id.reset(token)in afinallyblock ensures the value is cleared after the request, which matters in any environment that reuses worker threads or tasks across requests (otherwise a stale correlation id could leak into the next request handled by the same worker).
Demo
Running the module directly (with uuid.uuid4 replaced by a deterministic counter-based generator, uuid.UUID(int=n) for n = 1, 2, 3, ..., purely so this demo's output is reproducible instead of depending on real randomness):
INFO corr=00000000-0000-0000-0000-000000000001 received request path=/orders
INFO corr=00000000-0000-0000-0000-000000000001 order validated
INFO corr=upstream-fixed-id received request path=/orders
INFO corr=upstream-fixed-id order validated
INFO corr=00000000-0000-0000-0000-000000000002 about to fail
demo1 response headers: {'X-Correlation-ID': '00000000-0000-0000-0000-000000000001'}
demo2 response headers: {'X-Correlation-ID': 'upstream-fixed-id'}
demo3: exception propagated as expected
context after all requests (should be default '-'): -
Three cases exercised: (1) no inbound header, a fresh id is generated and both log lines during that request carry it, and it lands on the response; (2) an inbound header (upstream-fixed-id) is present, and it's reused rather than replaced; (3) the handler raises, and the correlation id is still attached to the log line emitted before the exception, and the finally reset means the contextvar is back to its default - afterward, i.e. it didn't leak into whatever runs next on that worker.
Complexity
Per request, the decorator does a constant number of operations: one dict lookup for the inbound header, one ContextVar.set/reset pair, one dict write for the outbound header. That's O(1) time and space overhead per request, independent of how many log lines are emitted inside the handler, since each log call just reads the already-set contextvar rather than redoing any work.
Edge cases
- Missing or empty header: falls through to generating a new UUID, handled by the
if incoming elsecheck (an empty string is falsy, so it's treated the same as missing). - Header already present from upstream: reused as-is rather than overwritten, so a multi-hop call chain keeps one correlation id instead of getting a new one at every service.
- Handler raises an exception: the
try/finallystill resets the contextvar and the correlation id is still attached to any log lines emitted before the exception, so the failure is traceable. - Background work spawned from inside the handler (e.g.
asyncio.create_task(...)or a thread pool submission): a new asyncio Task does inherit the current contextvars snapshot at creation time by default, so logs from a task created directly inside the handler will still carry the id; but work handed off to a separate thread pool or a message queue does not automatically carry it; that has to be passed explicitly (e.g. as a queue message field) and re-set with_correlation_id.set(...)on the receiving side. - Concurrent requests on the same worker: isolated correctly by
ContextVar, unlike a plain module-level variable, which would let concurrent requests clobber each other's id.
What's the difference between structured and unstructured logging? Also, walk through when you'd log at DEBUG versus INFO versus WARN versus ERROR, and how that choice affects an on-call engineer during an incident.
Sample Answer
Direct answer
Unstructured logging is free-form text intended for a human to read line by line; structured logging emits each entry as a consistent, machine-parsable record (typically JSON) with named fields, so a log-query tool can filter, aggregate, and correlate across millions of lines the way you'd query a database. Log levels (DEBUG, INFO, WARN, ERROR) are an orthogonal concept from structure: they control which entries you even keep or surface, so during an incident an on-call engineer isn't wading through routine noise to find the handful of lines that actually explain what broke.
Structured versus unstructured
| Unstructured | Structured | |
|---|---|---|
| Format | Free text, e.g. 2026-07-17 ERROR payment failed for user 123: timeout | Consistent fields, e.g. {"level":"ERROR","service":"payments","user_id":"123","error_type":"timeout"} |
| Querying | Regex/grep, brittle if the message wording ever changes | Filter and aggregate directly on fields (status_code >= 500), stable across wording changes |
| Correlation | Hard to reliably join with traces or other services | A shared trace_id/request_id field lets you pivot straight from a metric spike to the exact request's log lines |
| Best for | Quick local debugging, human-only reading | Production systems at any real scale, automated alerting on log content |
When to log at each level, and why it matters during an incident
| Level | Use for | Effect on an on-call engineer during an incident |
|---|---|---|
| DEBUG | Fine-grained internal state, useful only when actively investigating | Should normally be off in production (or sampled), since it's high-volume noise; if it's flooding the log stream, it drowns out the ERROR line the engineer actually needs |
| INFO | Normal, expected events: a request completed, a job started | Confirms the system is doing what it should; useful for confirming a fix worked, not for finding the problem itself |
| WARN | Something unexpected happened but the system recovered or degraded gracefully (a retry succeeded, a fallback kicked in) | Early signal: a spike in WARN volume right before an incident often shows the system trying to compensate before it actually failed |
| ERROR | An operation failed and did not recover on its own | This is what the on-call engineer searches for first; ERROR entries should carry enough context (request_id, error_type, relevant IDs) to explain what failed without needing to reproduce it |
Worked example: a structured log line for a failed request
{
"timestamp": "2026-07-17T15:04:05.123Z",
"level": "ERROR",
"service": "payments-api",
"request_id": "req-8f21",
"trace_id": "trace-a93c",
"status_code": 502,
"duration_ms": 247,
"error_type": "UpstreamTimeout",
"message": "upstream charge provider timed out"
}
During an incident, an on-call engineer can now do something like find all ERROR entries with error_type: UpstreamTimeout in the last 15 minutes, grouped by service, instead of searching for the word "timeout" across every service's free-text logs and hoping the wording matches. The shared trace_id also lets them jump directly from this log line to the distributed trace for the same request.
Trade-offs and pitfalls
- Logging everything at INFO "just in case" defeats the purpose of levels: if INFO volume is as high as DEBUG would be, the on-call engineer is back to searching through noise. Levels only help if they're used with discipline.
- Structured logging without a shared, enforced schema across services becomes almost as unqueryable as unstructured text, just in JSON clothing; a
status_codefield that's a string in one service and an integer in another breaks cross-service queries. - Sensitive fields (user identifiers, tokens, payment details) need to be redacted or hashed at the point of logging, not cleaned up later; once something sensitive is in a log aggregator, deleting it retroactively is unreliable.
- DEBUG-level logging left on in production is a common, avoidable cost problem: log ingestion is usually billed by volume, and DEBUG noise can dominate that cost without adding proportional value.
What are the three pillars of observability? For each one, explain what kind of question it's best at answering, one blind spot it has on its own, and a concrete example of a production issue it would help you catch.
Sample Answer
Direct answer
Observability rests on three complementary signal types: metrics, logs, and traces. Metrics tell you something is wrong and roughly how bad; logs tell you what specifically happened in a given event; traces tell you where in a multi-service request the time or failure occurred. None of the three alone gives a complete picture: a strong incident response usually starts with one pillar to detect and scope the problem, then pivots to another to find root cause.
The three pillars
| Pillar | Best at answering | Blind spot alone | Production issue it would catch |
|---|---|---|---|
| Metrics | "Is something wrong right now, and how widespread?" (aggregated time series: rates, latencies, saturation) | No per-request context, can't tell you which specific request or user was affected | A slow memory leak: heap usage climbing steadily over days trips a capacity alert before an out-of-memory crash |
| Logs | "What exactly happened for this one request or event?" (discrete, timestamped records) | Expensive to query in aggregate at scale; no built-in sense of "normal," so you need to already suspect something to search for it | A payment failing with a specific exception, e.g. a null card-token field surfaced in the stack trace, that a dashboard would only show as "errors up" |
| Traces | "Where in the call chain did the time or failure happen?" (request-scoped, spans across services) | Usually sampled, so rare failures can be missed entirely; requires instrumentation and consistent context propagation to be useful | A checkout endpoint's p99 latency doubles; a trace shows 900ms of the 1000ms total sitting in a single downstream inventory-service span, isolating exactly which hop got slow |
Instrumentation example, one flow
For a checkout endpoint, an on-call engineer might instrument it like this: a checkout_requests_total counter metric with labels {status, payment_provider}, plus a checkout_latency_seconds histogram with the same labels for percentiles; a structured log line at the point of failure with fields {request_id, trace_id, error_type, payment_provider}; and a trace with spans named checkout.validate, checkout.charge, checkout.persist, each carrying the same trace_id that appears in the log line. The shared trace_id and request_id are what let you jump from "the metric moved" to "here is the specific failing request" to "here is the exact log line explaining why."
Trade-offs and pitfalls
- Treating one pillar as sufficient is the most common mistake: teams that only have logs end up searching blind during an incident because they have no aggregated signal telling them where to look first; teams that only have metrics can detect a problem but can't explain it.
- High-cardinality labels (like an unbounded user_id on a metric) turn cheap metrics into an expensive, slow-to-query mess; that data belongs in logs or traces instead.
- Trace sampling is a real trade-off: full sampling captures every rare failure but is expensive at scale; low sampling rates are cheap but can miss the exact failing request you need. Tail-based sampling (keep traces for slow or error requests) is a common middle ground.
- Retention windows differ by pillar in practice (metrics are cheap to keep for months, verbose logs and full traces are usually much more expensive to retain), which shapes how far back a postmortem can actually look.
What's the difference between an SLI, an SLO, and an SLA? Walk through how you'd define each one concretely for a service you've worked on, including how you'd measure the indicator and what time window you'd use.
Sample Answer
Direct answer
An SLI is the measured signal (e.g. the percentage of requests that were fast and successful), an SLO is the internal target for that signal over a time window (e.g. 99.9% of requests succeed under 300ms over a rolling 30 days), and an SLA is the external, often contractual promise built on top of an SLO, usually with margin, and real consequences (service credits, penalties) if it's breached. In short: the SLI measures, the SLO targets, the SLA promises with a business wrapper attached.
Defining each one concretely, for an e-commerce checkout service
| Term | Definition here | Concrete value |
|---|---|---|
| SLI | Percentage of checkout requests that return successfully (2xx) within 300ms | Measured every minute from load-balancer access logs |
| SLO | Internal reliability target for that SLI | 99.9% of checkout requests succeed under 300ms, measured over a rolling 30-day window |
| SLA | External commitment to customers, usually looser than the SLO to leave margin | 99.5% monthly checkout availability, with service credits below that threshold |
Explained for a non-technical stakeholder: the SLO is the bar the engineering team holds itself to internally so problems get caught and fixed before they become customer-visible; the SLA is the, usually more forgiving, bar the business is willing to be held to externally, on paper, with money attached if it's missed. The gap between the two is deliberate headroom, not sloppiness.
Worked example: computing the error budget
The error budget is the amount of allowed failure baked into the SLO. It's what lets a team ship changes at all instead of freezing forever.
error budget=(1−SLO target)For the 99.9% SLO above:
error budget=1−0.999=0.001=0.1%Over the 30-day measurement window (30 days equals 2,592,000 seconds):
0.001×2,592,000s=2,592s≈43.2 minutesSo this service is allowed about 43 minutes of budget-consuming failure (however that failure is defined: downtime, over-300ms responses, and so on) across a 30-day window before the SLO itself is breached. If checkout handles, say, 1,000,000 requests over that same window, the equivalent request-based budget is:
0.001×1,000,000=1,000 requests allowed to violate the SLIBoth framings, time-based and request-based, describe the same budget; which one is more useful depends on whether the SLI is availability-style (was the service up) or ratio-style (what fraction of requests succeeded).
Trade-offs and pitfalls
- Confusing SLO and SLA in conversation causes real problems: teams sometimes design their alerting and release-gating around the SLA (the looser, contractual number) instead of the SLO (the tighter, internal number), which means by the time anyone notices, the team is already close to breaching the external promise with no margin left to react.
- An SLO with no error-budget policy attached is just a number on a dashboard; the value comes from what happens when the budget is nearly exhausted (freeze risky launches, redirect engineering time to reliability work), not from the target itself.
- Picking an SLI that doesn't reflect real user experience (e.g. "server process is running" instead of "requests succeed within an acceptable time") gives a green dashboard while users are still unhappy. The SLI has to be as close to what the user actually experiences as the team can measure.
What does observability mean to you? Explain how metrics, logs, and traces each contribute to understanding a running system, and when you'd reach for one over the others.
Sample Answer
Direct answer
Observability is the ability to ask new questions about a running system's internal state from its external outputs, metrics, logs, and traces, without shipping new code or adding new instrumentation for every question you didn't anticipate in advance. It's a broader goal than monitoring: monitoring is built around known failure modes you already predicted and set up dashboards and alerts for; observability is what lets you investigate the failure modes nobody predicted.
How metrics, logs, and traces each contribute
- Metrics: aggregated numeric signals over time (request rate, error rate, latency). Best for answering "is something wrong, and roughly how bad," cheaply and at a glance, which is why they're the backbone of dashboards and alerting.
- Logs: discrete, timestamped records of specific events. Best for answering "what exactly happened," with the detail (stack traces, payloads, IDs) that a metric can't carry.
- Traces: the path and timing of one request across every service it touched. Best for answering "where in this specific request did the time or failure occur," especially in a distributed system where the slow part could be any of several services.
When I'd reach for one over the others
I'd reach for metrics first, almost always, because they're the cheapest to check and tell me whether I even have a real problem worth digging into. From there, if the problem spans multiple services, I'd reach for traces to localize which hop is responsible; if I need to understand the specific reason a request failed, not just where it slowed down, I'd reach for logs, ideally pivoting there using the same trace or request ID that traces and metrics already pointed me to.
Worked example: an intermittent error spike
Metrics show an elevated error rate and p99 latency for service A over the last 20 minutes. Pulling a trace for a recent failing request shows most of the time sitting in a call to service B, with retries visible as repeated spans. Pulling the logs for that same request (using the shared trace ID) shows the actual failure: a JSON parsing error caused by a malformed field in service B's response, which is what's triggering the retries and driving up latency. The fix is a quick input-validation patch at the edge to reject the malformed field before it reaches service B, plus a cap on retries so a bad request doesn't compound into a latency spike.
Trade-offs and pitfalls
- Observability isn't a tool you buy, it's a property of the system, shaped by how well it's instrumented. Buying a vendor platform without consistent, correlated instrumentation (shared IDs across metrics, logs, and traces) gives you three separate data sources, not observability.
- Monitoring only covers what you thought to watch for in advance; a system can pass every dashboard and alert check while still failing in a way nobody predicted, which is exactly the gap observability is meant to close, by letting you investigate the unknown-unknowns after the fact instead of only the known-knowns.
- Over-instrumenting everything at maximum detail (full trace sampling, verbose logs everywhere) is expensive and can itself become a performance and cost problem; the goal is enough signal to investigate confidently, not maximum data volume.
Unlock Full Question Bank
Get access to all 13 Monitoring, Logging, and Observability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.