Production Incident Diagnosis and Distributed Systems Troubleshooting Questions
Debugging distributed systems under fire: diagnosing latency and reliability regressions, root-causing across service boundaries, reading traces and metrics during an incident, and reasoning about complex production failures. Covers the investigative method for hard-to-reproduce, multi-service problems. The operational counterpart to resilient design.
You're in the on-call rotation and receive alerts that API p95 latency has increased 3x and error rates have risen across several services. Lay out a step-by-step failure-mode analysis using metrics, logs, and distributed traces to isolate the root cause, including which experiments you would run to narrow the search (for example isolating individual downstreams or replaying traffic) and how you would validate a proposed fix safely in production. Include quick mitigation steps you might take while you're still investigating.
Sample Answer
Direct answer. A 3x jump in p95 with errors rising across several services is a signal that something shared or upstream degraded, not that every service independently broke at once, so the investigation should start by finding what those services have in common rather than debugging them one at a time.
Structured elaboration.
- Find the common dependency. Check whether the affected services share a downstream call: a database, a cache, an authentication service, a service mesh sidecar, or a shared piece of infrastructure like DNS or a load balancer. Distributed tracing is the fastest tool for this: pull a handful of slow traces from different affected services and see where the time is actually going. If every trace bottoms out in the same downstream span, you've found your candidate cause in minutes instead of hours. Logs from the affected services for the same window are worth pulling alongside the traces, specifically scoped to the shared downstream call: an explicit error, timeout, or connection-refused message in the logs directly confirms what the trace can only imply, and often narrows 'the auth service is slow' down to something as concrete as 'the auth service is throwing a connection-pool-exhausted exception,' which is a more actionable finding than latency alone.
- Correlate with metrics. Once you have a candidate, check that dependency's own metrics (its error rate, its latency, its saturation) for the same time window. If its p95 or error rate jumped at the same moment your services' did, that's strong corroborating evidence.
- Run targeted experiments to confirm, not just infer. If you suspect a specific downstream, isolate a request to that path alone (bypassing the rest of the flow if you can) and see whether it reproduces the slowdown. Replaying a sample of real traffic against a canary (a small slice of live production instances running the candidate change) or a shadow environment (a copy of the system that receives a duplicate stream of real traffic but whose responses are discarded, used purely to observe behavior safely) is another way to confirm the same hypothesis without risking more production traffic.
- Mitigate while you confirm. Depending on what you find, options include failing over to a healthy replica, shedding non-critical load off the shared dependency, or temporarily degrading a feature that depends on the slow path (serving cached or stale data instead of failing).
- Validate the proposed fix the same way you validated the hypothesis: with a canary, not a blanket rollout. Deploy the candidate fix (a scale-out, a rollback, a config change) to a small slice of the affected service's instances first, and compare that slice's error rate and latency against the untreated rest over the next several minutes before widening it. If the canary slice doesn't recover, the root-cause hypothesis was wrong or incomplete, and that's better discovered on a small slice of traffic than on all of it.
Worked example. Say traces from three of the affected services all show a 300 to 400ms span on a call to the shared authentication service, while every other span in those traces looks normal. You check the auth service's own dashboard and see its p99 latency also jumped 3x at the same timestamp, and its CPU is pegged. That converges on a specific, testable hypothesis (the auth service itself is overloaded) rather than three separate mysteries. A quick check of the auth service's own recent changes and traffic volume tells you whether it's a capacity problem (traffic grew faster than provisioned capacity) or a regression (a recent deploy made it slower per request) so you know whether the mitigation is scaling it out or rolling it back.
Trade-offs and pitfalls. The main trap is debugging each affected service in isolation, which multiplies your investigation time by the number of services and can lead different responders to different, contradictory hypotheses. The other trap is stopping at 'auth service is slow' without confirming WHY, since 'scale it out' and 'roll back the last deploy' are very different mitigations that fix different root causes; guessing wrong burns time you don't have during an active, multi-service incident.
Explain the role of logs, metrics, and distributed traces in troubleshooting a distributed system. For each of the three, cover what information it provides, an example symptom that's best diagnosed with it, and one limitation. Then, given this scenario: an API service shows a sustained increase in its 5xx rate and p95 latency over the last 10 minutes, describe the order in which you would consult logs, metrics, and traces, and why that order.
Sample Answer
Direct answer. Logs, metrics, and traces answer different questions: metrics tell you something is wrong and roughly how bad, traces tell you WHERE in a multi-service request the time or error actually happened, and logs tell you the specific WHY once you know where to look; for the given scenario, the efficient order is metrics first, then traces, then logs.
Structured elaboration.
- Metrics. What they provide: aggregated, numerical signals over time (request rate, error rate, latency percentiles) that are cheap to query and great for detecting that something changed and roughly when. Example symptom best diagnosed with metrics: a gradual latency creep over days, which is hard to spot from individual traces or logs but obvious on a time-series graph. Limitation: metrics are aggregates, so they can't tell you WHICH specific request failed or why; a metric can tell you error rate is 5% but not which five requests out of a hundred, or what made those five different from the other ninety-five.
- Traces. What they provide: the path of a single request across multiple services, showing where time was spent and where in the chain something failed, which metrics alone can't show for multi-service systems. Example symptom best diagnosed with traces: a specific endpoint is slow, and you need to know whether the time is in your own service, a downstream call, or a database query, the way a span breakdown showed in an earlier question here. Limitation: traces are sampled in most systems at any real scale, so a rare intermittent bug may simply not appear in your trace sample, and traces don't easily show you patterns ACROSS many requests, only individual ones.
- Logs. What they provide: detailed, often free-text or structured information about exactly what a piece of code was doing at a specific moment, including values, error messages, and stack traces. Example symptom best diagnosed with logs: understanding the EXACT error message or exception behind a failure that a trace has already told you is happening in a specific service. Limitation: logs are voluminous and expensive to search broadly, so they're a poor starting point when you don't yet know where to look; searching all logs for 'something is slow' is far less efficient than searching the specific service's logs once a trace has pointed you there.
- Applying the order to the given scenario. For 'API service A shows a sustained increase in 5xx rate and p95 latency for the last 10 minutes': start with metrics, since you already have them and they confirm scope, timing, and severity (is it every endpoint or one, is it every region or one, is 5xx rising with latency or independently). Then pull traces for a sample of the actual slow or failing requests during that window, since that narrows WHERE in the request path the problem lives. Only once traces point to a specific service or call would you go read that service's logs for the exact error or exception.
Trade-offs and pitfalls. Starting with logs when you don't yet know where to look is the most common inefficient order: broad log search across many services, with no scope yet established, tends to waste the early minutes of an incident on noise. The reverse mistake, staying in metrics dashboards indefinitely without ever pulling a trace or a log, can leave you with a confirmed symptom and no explanation, since metrics alone rarely tell you the specific mechanism behind a failure in a multi-service system.
Root-cause analysis case: given the signals below, produce a short RCA describing the causal chain, immediate remediation steps, and long-term fixes.
Metrics: service-a p95=800ms, error-rate=10% for the last 15 minutes
Traces (sample): gateway -> service-a (span 600ms) -> service-b (span 580ms) -> redis (span 560ms)
Logs from service-b: repeated 'ERR connect timeout to redis host:6379'
Write the RCA summary: the causal chain, short-term mitigations, and long-term engineering or operational fixes.
Sample Answer
Direct answer. The causal chain here is: service-a is calling service-b, which is calling redis, and redis connections are timing out, which stalls service-b, which stalls service-a; the error message ('connect timeout to redis') is the concrete evidence that redis connectivity, not application logic, is the proximate cause.
Structured elaboration.
- Build the chain from the evidence given, not assumption. The trace shows gateway calling service-a (600ms), which calls service-b (580ms), which calls redis (560ms). Almost all of service-a's 600ms is inside its call to service-b, and almost all of service-b's 580ms is inside its call to redis. Combined with the log line 'ERR connect timeout to redis host:6379' repeated in service-b, the chain reads: redis is not responding to new connections in time, so every request through service-b (and therefore through service-a) pays that timeout cost, and 10% of them fail outright.
- Distinguish plausible causes of a redis connect timeout. This symptom is consistent with several different underlying problems: redis itself is overloaded or down, the network path to redis is degraded, redis's connection limit has been reached so new connections queue or get refused, or service-b's own connection pool to redis is exhausted or misconfigured (for example too few connections for its current traffic, or connections not being released properly).
- Use additional signals to narrow it down. Redis's own metrics (CPU, memory, connected-clients count, command latency) would show whether redis itself is unhealthy. Service-b's connection-pool metrics would show whether it's a client-side exhaustion problem rather than a server-side one. A
redis-cli PINGor a synthetic connection test from the same network path as service-b would show whether the network path itself is the issue, separate from whether redis is overloaded. - Immediate remediation. Options depend on what step 3 finds: if redis is overloaded, shedding load or failing over to a replica; if it's a client-side connection-pool exhaustion, increasing the pool size or adding a circuit breaker so service-b fails fast instead of piling up timeouts; if it's the network path, routing around the degraded path if there's an alternative.
- Long-term fixes. Add a circuit breaker around the redis call so a redis outage degrades service-b gracefully (serving stale or default data, or failing fast) instead of cascading its own latency upstream into service-a and the gateway. Add explicit alerting on redis connection-timeout rate specifically, since it's a distinctive and actionable signal that this trace shows was already present before the wider error rate and latency alerts fired.
Worked example. If redis's own dashboard shows connected_clients pinned at its configured maximum throughout the incident window while CPU and memory on the redis host stay normal, that specifically points to connection exhaustion rather than redis being overloaded: either too many clients are opening connections without closing them (a leak somewhere, possibly in service-b or another caller), or redis's maxclients setting is simply too low for current legitimate traffic. Checking service-b's own connection-pool configuration would tell you which: if service-b's pool size looks reasonable and connections are being released properly, the fix is raising redis's maxclients; if service-b is opening far more connections than its configured pool size suggests it should, the fix is finding and closing the leak in its client code.
Trade-offs and pitfalls. The trap in this kind of RCA is stopping at 'redis connect timeout' and treating that as the root cause, when it's really a symptom with at least four different plausible underlying causes; each one implies a different fix, and applying the wrong one (for example scaling up redis when the real problem is a client-side connection leak) won't resolve the incident. It's also worth being explicit in the writeup about WHY the fix you propose follows from the evidence you gathered, not just from a guess, since that's what lets someone else validate or challenge the conclusion.
Given this simplified trace for a single request, identify where the latency spike originates and why:
TraceID: abc123
Spans:
- gateway (api-gw): duration 50ms
- auth (service-b): duration 5ms
- payments (service-c): duration 400ms
- db-proxy (service-d): duration 380ms
- db-query: duration 370ms
- db-proxy (service-d): duration 380ms
Explain the steps you would take to confirm the database is the true root cause and what further data you'd collect before concluding.
Sample Answer
Direct answer. Reading the spans, the request makes three sequential hops (gateway, then auth, then payments), and within the payments hop the 400ms is almost entirely accounted for by the 380ms db-proxy span beneath it, which is itself almost entirely the 370ms db-query span, so the database query is the dominant contributor to this request's latency; the gateway and auth hops (50ms and 5ms) are not meaningfully part of the problem.
Structured elaboration.
- First, read the SHAPE of the tree correctly, before computing anything.
gateway,auth, andpaymentsare three separate, sibling spans in this trace: one request flowing sequentially through a gateway hop, an auth check, and a payments call, one after another, not one nested inside another. Onlypaymentshas a child of its own (db-proxy, which in turn has its own child,db-query). That distinction matters because it decides which spans you even need to subtract anything from. - Compute each span's OWN time (its duration minus its children's durations), only where a child exists to subtract.
gatewayandauthhave no children shown in this trace, so their own time is simply their full duration: 50ms and 5ms.paymentshas one child (db-proxyat 380ms), so its own time is 400 minus 380 = 20ms: the payments SERVICE's own code cost only 20ms, the rest of its 400ms is time spent waiting on its call todb-proxy.db-proxyhas one child (db-queryat 370ms), so its own time is 380 minus 370 = 10ms.db-queryhas no child shown, so its own time is its full 370ms.span total duration child duration own time gateway 50ms none 50ms auth 5ms none 5ms payments 400ms 380ms (db-proxy) 20ms db-proxy 380ms 370ms (db-query) 10ms db-query 370ms none 370ms - Add up the own-times and see where the total request time actually went. 50 + 5 + 20 + 10 + 370 = 455ms, which is the full sequential request time (50 + 5 + 400). db-query's 370ms alone is roughly 370/455 ≈ 81% of that total, by a wide margin the single largest contributor; gateway and auth together are only about 12%.
- Confirm it's the query itself, not something around it. Before concluding 'the database is slow', check three distinct possibilities that all show up as a slow
db-queryspan: the query itself is slow (missing index, bad query plan, lock contention), the connection pool made the caller wait before the query even started (which would usually show as extra time BEFORE the query span begins, not inside it), or the database host itself is resource-constrained (high CPU, disk I/O, or replication lag if this is a replica). - Pull further data to distinguish those. Check the database's own slow-query log or
EXPLAINplan for this query around the incident window; check whether the query's duration is consistently ~370ms or spiking intermittently (consistent points to a plan or index problem, spiky points to contention or a noisy neighbor on the host); check the database host's own CPU, I/O, and lock-wait metrics for the same window. - Confirm before you commit to a fix. If
EXPLAINshows a sequential scan where an index should be used, or the slow-query log shows this exact query is consistently the slowest one running, that confirms the query itself and points to an index or query-shape fix. If instead the database host's CPU or I/O is saturated across many queries at once, the fix is capacity or isolating this workload, not this one query.
Worked example. Suppose the slow-query log shows this exact query pattern (a lookup by a non-indexed customer_email column) running consistently between 350 and 390ms throughout the incident window, while other indexed queries against the same table stay under 10ms. That consistent, query-specific slowness (not database-wide) points squarely at a missing index on customer_email: adding it would be expected to bring this query down to roughly the same low-single-digit-millisecond range as the other indexed lookups on that table. Since db-query is nested three levels deep (inside db-proxy, inside payments), fixing it collapses the whole chain: db-proxy's own 10ms plus a near-zero query time would bring payments down from 400ms to roughly its own 20ms plus a few milliseconds of query time, and the full request (gateway + auth + payments) would drop from about 455ms to well under 100ms.
Trade-offs and pitfalls. The most common mistake reading a trace like this is to see 'payments took 400ms' and start investigating the payments service's own code, when the trace is telling you the payments service's code only cost about 20ms and the other 380ms is entirely its call to the database, three levels down. The second most common mistake is the one this answer corrects itself on above: subtracting a span's duration from a SIBLING's duration because they happen to be listed near each other, rather than checking which spans are actually nested inside which. Always attribute time to the span whose OWN duration (after subtracting only its ACTUAL children) is largest, and follow the tree down until a span has no slower child left. It's also worth checking for cache misses: if this query is normally served from a cache and the cache was cold or evicted, the fix might be restoring the cache rather than touching the database at all.
You're on-call for a service and see increased 500 errors concentrated in one endpoint minutes after a deploy went out. Walk through the immediate steps you take in the first 15 minutes: how you determine whether the deploy actually caused the regression versus a coincidental correlation, what dashboards and logs you check first, your mitigation options (rollback, canary rollback, throttling), and how you communicate status to stakeholders.
Sample Answer
Direct answer. In the first 15 minutes the goal is to stabilize user-facing traffic and preserve evidence, in that order; you are not trying to find the root cause yet, you are trying to stop the bleeding and make sure you (or whoever picks this up) can still find the root cause afterward.
Structured elaboration.
- Confirm the deploy is actually correlated before assuming it's the cause. Pull up the deploy timeline and compare it against when the 500s started. If a deploy to this endpoint landed in the same window, that is your leading hypothesis, but check it, don't assume it: some deploys are unrelated and the timing is coincidence.
- Look at the error itself. Open a sample of the failing requests in logs or your APM (application performance monitoring tool, e.g. Datadog or New Relic). A stack trace or specific error code will usually tell you whether this is a code bug in the new release (null pointer, unhandled exception, a broken dependency call) versus something environmental (a config value that didn't get set, a database migration that didn't finish).
- Check dashboards for blast radius. Is this endpoint's error rate the only thing affected, or is latency also up, are other endpoints degrading too, is a specific region or instance pool worse than others? This tells you whether containment can be narrow (this endpoint only) or needs to be broader.
- Mitigate, choosing among your real options. A full rollback is usually the fastest, safest action, since patching forward live under pressure is a common source of a SECOND incident. If a full rollback isn't immediately available (say, other changes have shipped on top of it), a canary rollback, reverting just the newest batch of instances back to the prior version while leaving the rest alone, narrows the blast radius while you confirm the fix works. Throttling the affected endpoint (accepting some added latency or rejecting a fraction of requests) is a third lever worth having ready if neither rollback option is immediately deployable, buying time without needing a code or deploy change at all.
- Preserve evidence as you go. Before you roll back, capture a handful of the actual failing request logs, the relevant dashboard screenshots, and the exact timestamps. Once you roll back, the failure signal disappears and you lose your best evidence for the eventual root-cause writeup.
Worked example. You see the 500s are concentrated in one endpoint, a deploy to that service's code landed 6 minutes before the alert, and a sample of the failing requests shows a NullPointerException in a new code path. That is enough correlation plus evidence to roll back immediately: you trigger the rollback, watch the error rate over the next 2 to 3 minutes, and confirm it returns to baseline. You now have a clean signal (rollback fixed it) and preserved logs, which together make the eventual root-cause writeup straightforward: the new code path didn't null-check a field that's optional in a subset of real traffic.
- Communicate status to stakeholders throughout, not just at resolution. Post a short update to your incident channel or status page as soon as you've confirmed the correlation in step 1: what's affected, the rough blast radius from step 3, and what mitigation you're about to try. Once you roll back, a second short update (what you did, whether it worked, and what's still open) closes the loop. This matters even inside a 15-minute window, because stakeholders (support, other engineers, sometimes leadership) making decisions off silence or guesswork is a common secondary problem during an incident.
Trade-offs and pitfalls. The instinct to 'fix it properly' under pressure, by patching the bug live rather than rolling back, is usually the wrong call: a patch written and shipped in the middle of an active incident has not been reviewed or tested the way the original deploy was, and a second bad deploy during an active incident is a genuinely common failure pattern. The other pitfall is rolling back so fast that you never capture the failing-request evidence, which leaves the eventual root-cause analysis guessing instead of grounded in what you actually saw. A third pitfall is treating stakeholder communication as a nice-to-have that happens after the fact: going quiet during an active regression, even a short one, tends to generate more anxious pings and duplicated investigation than a two-line status update would have cost.
Unlock Full Question Bank
Get access to all 7 Production Incident Diagnosis and Distributed Systems Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.