Production Incident Diagnosis and Distributed Systems Troubleshooting Questions
Debugging distributed systems under fire: diagnosing latency and reliability regressions, root-causing across service boundaries, reading traces and metrics during an incident, and reasoning about complex production failures. Covers the investigative method for hard-to-reproduce, multi-service problems. The operational counterpart to resilient design.
A production service shows sporadically high CPU time in the kernel (sys time). Propose how you would use eBPF, bpftrace, or bcc tools to profile syscalls, sample stack traces, and determine whether the cause is kernel-level (for example futex contention, epoll_wait, or network interrupts) or genuinely user-space CPU. Give example bpftrace one-liners or bcc tools you would actually run.
Sample Answer
Direct answer. Sporadic high kernel (sys) time, as opposed to user-space CPU time, means the process is spending real CPU cycles inside the operating system itself, most commonly from syscalls, interrupt handling, or contention the application code doesn't directly control, and eBPF-based tools let you see exactly which kernel functions are consuming that time without modifying or restarting the running process.
Structured elaboration.
- Confirm it really is kernel time, and get a first-pass breakdown. A tool like
bpftrace's built-in profiling (orperf) can sample stack traces system-wide during the sporadic spike and show you a flame-graph-style breakdown of which kernel functions are hottest; this is the fastest way to go from 'sys time is high' to 'here are the specific functions responsible' without guessing. - If the profile points at network interrupts, a high volume of small packets, an interrupt storm, or the network card's interrupts landing disproportionately on one CPU core, rather than being spread across cores via receive-side scaling (a network-card feature that distributes incoming network interrupts across multiple CPU cores instead of pinning them all to one), are common causes;
bpftracecan tracesoftirq(a deferred, lower-priority form of interrupt handling the Linux kernel uses to finish processing a network packet after the initial hardware interrupt) and hardware-interrupt events specifically to confirm. - If the profile points at futex (a common Linux syscall behind most user-space lock implementations), that usually means the APPLICATION is contending on a lock heavily enough that the kernel-level futex wait/wake path itself becomes a meaningful cost, which links this investigation back to the same kind of lock-contention pattern as an earlier CPU-spike question here, just visible at a lower level.
- If the profile points at
epoll_waitor similar, that's often actually benign (a process efficiently waiting for I/O readiness looks like kernel time but isn't a problem on its own), so distinguishing 'a lot of time in epoll_wait because the process is idle and waiting, which is fine' from 'a lot of time in epoll_wait because of excessive wakeups from a misbehaving event loop' matters, and tracing the RATE of wakeups (not just time spent) helps tell them apart. - Example commands. A one-liner like
bpftrace -e 'profile:hz:99 { @[kstack] = count(); }'samples kernel stacks system-wide at 99Hz and counts occurrences, giving you a ranked list of the hottest kernel call paths over your sampling window; a targeted one likebpftrace -e 'kprobe:futex_wait { @[comm] = count(); }'counts futex-wait entries specifically, broken down by process name, which would confirm or rule out lock contention as the mechanism directly.bcc'sfunclatencyorprofiletools provide similar capability with less hand-written tracing code, if available on the host.
Worked example. Suppose the bpftrace kernel-stack profile during a spike shows roughly 70% of sampled kernel time inside futex_wait and related paths, concentrated on this service's own process. That points squarely at application-level lock contention manifesting as kernel time (since the underlying futex mechanism involves a kernel-level wait/wake handshake once contention is high enough), not a network or I/O issue. Cross-checking with a user-space thread dump taken at the same moment, showing many threads blocked waiting on the same lock object, corroborates it independently. The fix, at this point, is the same as any lock-contention problem: reduce the critical section, shard the lock, or reduce concurrency contending on it, not anything at the kernel or network level, even though the SYMPTOM (high sys time) initially pointed at the kernel.
Trade-offs and pitfalls. It's a common and understandable mistake to see 'high kernel time' and assume the problem is infrastructure-level (network, disk, the OS itself) when, as this example shows, it's frequently a downstream SIGNATURE of an application-level problem like lock contention; the kernel is just where that contention becomes visible as CPU time. eBPF tracing is low-overhead and safe to run on production without restarting anything, which is exactly why it's the right first tool here rather than something riskier like attaching a debugger, but sampling frequency (like the 99Hz above) is a real trade-off between profiling resolution and the (small but nonzero) overhead of tracing itself.
A load-balancer health check marks instances unhealthy too aggressively, causing cascading restarts. Given the pseudocode below, identify the problems and propose an improved health-check and backoff strategy.
if cpu_percent > 90:
consecutive_failures += 1
else:
consecutive_failures = 0
if consecutive_failures >= 3:
mark_unhealthy()
Explain your improvements and the reasoning behind them.
Sample Answer
Direct answer. The given health check is a ticking time bomb: it counts consecutive failures using only CPU percentage, with no recovery signal and no distinction between a genuinely dead instance and one that's merely busy, so it will mark healthy-but-loaded instances unhealthy and can create a feedback loop where removing 'unhealthy' instances increases load on the survivors, pushing them past the same threshold too.
Structured elaboration.
- The core bug: no positive signal, only a negative one. The pseudocode only ever asks 'is CPU high', never 'can this instance actually serve a request correctly'; a genuinely responsive instance under a legitimate, temporary CPU spike (a garbage-collection pause, a burst of traffic) gets treated identically to an instance that's truly wedged. A real health check should probe actual request-serving capability (a lightweight endpoint that exercises the real request path, or at least confirms the process can respond at all), not infer health indirectly from one resource metric.
- No hysteresis (requiring a condition to hold steadily for a while before flipping state, so a value bouncing around the threshold doesn't cause rapid back-and-forth decisions) or recovery path shown. The pseudocode increments on failure and resets to 0 on any single success, which actually makes it MORE trigger-happy than a stable threshold: an instance oscillating right around 90% CPU could bounce between counted and reset in a way that never quite reaches 3 to mark unhealthy, or conversely could flap in and out of the unhealthy state repeatedly, causing exactly the cascading-restart pattern described.
- No backoff on the resulting action.
mark_unhealthy()fires immediately at the threshold with no apparent cooldown or gradual response; a health-check system that immediately and simultaneously restarts (or removes from rotation) EVERY instance that crosses the threshold at once, during a real, shared traffic-driven CPU spike, can remove capacity exactly when it's most needed, worsening the load on whatever instances remain. - No visibility into the DECISION. There's no logging or metric emitted at each step, so if this logic behaves unexpectedly in production, there's nothing to look at afterward except the fact that a restart happened; any health-check logic making a consequential decision should emit why.
- Improved design. Check actual responsiveness (a lightweight but real request), not just CPU, as the PRIMARY signal, with CPU as at most a secondary contributing factor. Require sustained failure across BOTH consecutive checks and a minimum wall-clock duration (not just a raw consecutive count, which can be gamed by a fast check interval), and add a distinct, harder-to-hit threshold for consecutive SUCCESSES needed to recover, avoiding a single lucky success resetting a genuinely struggling instance's count to zero. Stagger or rate-limit how many instances can be marked unhealthy within a short window, so a shared, correlated CPU spike across the whole fleet doesn't trigger a mass, simultaneous removal.
Worked example. Say five instances all cross 90% CPU within the same 10-second window because of a genuine, if temporary, traffic surge, each running the given logic with a 10-second check interval; all five reach consecutive_failures = 3 (30 seconds) at roughly the same time and all get marked unhealthy simultaneously, removing 5 instances' worth of capacity from a pool that was already under real load, which very plausibly pushes the REMAINING instances' CPU even higher, triggering the same logic on them next, a classic self-reinforcing cascade. With a stagger rule (say, no more than 1 instance per 30-second window gets marked unhealthy from this signal, forcing the rest to wait and re-check) combined with an actual request-responsiveness check as the primary signal (which would likely show all five instances ARE still successfully serving requests, just slowly, under the genuine load spike), none of the five would have been marked unhealthy at all in this scenario, since high CPU under real, successfully-served load is not the same failure as an unresponsive instance.
Trade-offs and pitfalls. A more conservative health check (requiring longer sustained failure, checking actual responsiveness) trades faster removal of genuinely broken instances for fewer false positives; that trade is almost always worth it, since the cost of a truly dead instance staying in rotation a bit longer is usually much smaller than the cost of a self-reinforcing mass-removal cascade. Staggering removals adds real complexity (some shared state or coordination across instances, rather than each instance deciding independently), which is a genuine engineering cost worth taking seriously rather than hand-waving away.
Given this simplified trace for a single request, identify where the latency spike originates and why:
TraceID: abc123
Spans:
- gateway (api-gw): duration 50ms
- auth (service-b): duration 5ms
- payments (service-c): duration 400ms
- db-proxy (service-d): duration 380ms
- db-query: duration 370ms
- db-proxy (service-d): duration 380ms
Explain the steps you would take to confirm the database is the true root cause and what further data you'd collect before concluding.
Sample Answer
Direct answer. Reading the spans, the request makes three sequential hops (gateway, then auth, then payments), and within the payments hop the 400ms is almost entirely accounted for by the 380ms db-proxy span beneath it, which is itself almost entirely the 370ms db-query span, so the database query is the dominant contributor to this request's latency; the gateway and auth hops (50ms and 5ms) are not meaningfully part of the problem.
Structured elaboration.
- First, read the SHAPE of the tree correctly, before computing anything.
gateway,auth, andpaymentsare three separate, sibling spans in this trace: one request flowing sequentially through a gateway hop, an auth check, and a payments call, one after another, not one nested inside another. Onlypaymentshas a child of its own (db-proxy, which in turn has its own child,db-query). That distinction matters because it decides which spans you even need to subtract anything from. - Compute each span's OWN time (its duration minus its children's durations), only where a child exists to subtract.
gatewayandauthhave no children shown in this trace, so their own time is simply their full duration: 50ms and 5ms.paymentshas one child (db-proxyat 380ms), so its own time is 400 minus 380 = 20ms: the payments SERVICE's own code cost only 20ms, the rest of its 400ms is time spent waiting on its call todb-proxy.db-proxyhas one child (db-queryat 370ms), so its own time is 380 minus 370 = 10ms.db-queryhas no child shown, so its own time is its full 370ms.span total duration child duration own time gateway 50ms none 50ms auth 5ms none 5ms payments 400ms 380ms (db-proxy) 20ms db-proxy 380ms 370ms (db-query) 10ms db-query 370ms none 370ms - Add up the own-times and see where the total request time actually went. 50 + 5 + 20 + 10 + 370 = 455ms, which is the full sequential request time (50 + 5 + 400). db-query's 370ms alone is roughly 370/455 ≈ 81% of that total, by a wide margin the single largest contributor; gateway and auth together are only about 12%.
- Confirm it's the query itself, not something around it. Before concluding 'the database is slow', check three distinct possibilities that all show up as a slow
db-queryspan: the query itself is slow (missing index, bad query plan, lock contention), the connection pool made the caller wait before the query even started (which would usually show as extra time BEFORE the query span begins, not inside it), or the database host itself is resource-constrained (high CPU, disk I/O, or replication lag if this is a replica). - Pull further data to distinguish those. Check the database's own slow-query log or
EXPLAINplan for this query around the incident window; check whether the query's duration is consistently ~370ms or spiking intermittently (consistent points to a plan or index problem, spiky points to contention or a noisy neighbor on the host); check the database host's own CPU, I/O, and lock-wait metrics for the same window. - Confirm before you commit to a fix. If
EXPLAINshows a sequential scan where an index should be used, or the slow-query log shows this exact query is consistently the slowest one running, that confirms the query itself and points to an index or query-shape fix. If instead the database host's CPU or I/O is saturated across many queries at once, the fix is capacity or isolating this workload, not this one query.
Worked example. Suppose the slow-query log shows this exact query pattern (a lookup by a non-indexed customer_email column) running consistently between 350 and 390ms throughout the incident window, while other indexed queries against the same table stay under 10ms. That consistent, query-specific slowness (not database-wide) points squarely at a missing index on customer_email: adding it would be expected to bring this query down to roughly the same low-single-digit-millisecond range as the other indexed lookups on that table. Since db-query is nested three levels deep (inside db-proxy, inside payments), fixing it collapses the whole chain: db-proxy's own 10ms plus a near-zero query time would bring payments down from 400ms to roughly its own 20ms plus a few milliseconds of query time, and the full request (gateway + auth + payments) would drop from about 455ms to well under 100ms.
Trade-offs and pitfalls. The most common mistake reading a trace like this is to see 'payments took 400ms' and start investigating the payments service's own code, when the trace is telling you the payments service's code only cost about 20ms and the other 380ms is entirely its call to the database, three levels down. The second most common mistake is the one this answer corrects itself on above: subtracting a span's duration from a SIBLING's duration because they happen to be listed near each other, rather than checking which spans are actually nested inside which. Always attribute time to the span whose OWN duration (after subtracting only its ACTUAL children) is largest, and follow the tree down until a span has no slower child left. It's also worth checking for cache misses: if this query is normally served from a cache and the cache was cold or evicted, the fix might be restoring the cache rather than touching the database at all.
Your stack runs microservices on Kubernetes with Prometheus and Jaeger. Describe three concrete debugging scenarios you've resolved in a similar environment. For each one, include the data sources you checked (logs, traces, metrics), your root-cause-analysis steps, the immediate mitigation, and the long-term fix.
Sample Answer
Direct answer. Three concrete scenarios that show the range of what 'debugging on Kubernetes with Prometheus and Jaeger' actually looks like: a service degrading from connection-pool exhaustion visible only in traces, a metric-only regression that traces alone wouldn't have explained, and a log-driven discovery of a misconfigured retry loop.
Structured elaboration.
- Scenario one: connection-pool exhaustion visible in traces. Symptom: intermittent slow requests to one service, with no obvious pattern in aggregate metrics. Data checked: Jaeger traces for the slow requests specifically (not a random sample), which showed a consistent gap of 200 to 400ms BEFORE the actual downstream call span even started, meaning the time was spent waiting to acquire a connection from the pool, not in the call itself. Root cause: the connection pool size hadn't been tuned when the service's traffic roughly doubled over the prior month. Immediate mitigation: temporarily raised the pool size via a config change. Long-term fix: added a Prometheus metric for pool-wait-time specifically (not just pool size), so this class of problem shows up on a dashboard instead of requiring someone to notice a gap in a trace.
- Scenario two: a metrics-only regression traces wouldn't explain. Symptom: gradual memory growth in a service over several days, eventually triggering OOM (out-of-memory) restarts. Data checked: Prometheus's memory-usage-over-time graph, which showed a steady linear climb rather than a step change, ruling out a single bad deploy as the trigger; correlating the growth rate against request volume showed memory grew roughly proportional to total requests served since the last restart, not proportional to time. Root cause: a cache inside the service that was never evicting entries, so it grew with cumulative traffic. Immediate mitigation: scheduled restarts as a stopgap. Long-term fix: added a bounded eviction policy to the cache.
- Scenario three: a log-driven discovery. Symptom: a downstream service's request rate was roughly 4x higher than the calling service's actual user-facing request rate, discovered from a routine metrics review rather than an active incident. Data checked: logs on the calling service, which showed a retry loop that, due to a bug in its backoff logic, was retrying every failed call 3 times with effectively no delay between attempts, and treating a specific class of expected 4xx response (a validation error that would never succeed on retry) as retryable. Root cause: overly broad retry-on-error logic. Immediate mitigation: cut the retry count from 3 down to 1 and added a temporary rate limit on the calling service's outbound calls to the downstream service, which brought the amplification down substantially within minutes while the retry logic itself was still being properly fixed. Long-term fix: scoped retries to only the error types that are actually likely to succeed on a retry (timeouts and 5xx), and added exponential backoff.
Trade-offs and pitfalls. These three scenarios deliberately use different primary evidence (traces, metrics, logs) because in practice you rarely know in advance which one will hold the answer; a systematic responder checks whichever source matches the symptom's shape (a specific slow request needs a trace, a trend over time needs a metric graph, an unexpected volume or behavior needs logs) rather than always reaching for the same tool out of habit.
After adopting a service mesh, your telemetry shows increased latency and CPU usage. Describe how you would diagnose whether the mesh itself is the root cause, the steps you'd take to mitigate the regression quickly (configuration changes, bypassing the mesh for specific paths), and the criteria you'd use to decide between continuing to optimize the mesh configuration or rolling back the adoption entirely.
Sample Answer
Direct answer. Because the mesh sits in the request path for every service that has it, the fastest way to confirm or rule it out as the cause is comparing behavior WITH and WITHOUT it in the loop for the same traffic, rather than reasoning about it in the abstract.
A service mesh works by running a small proxy process, called a sidecar, alongside every instance of every service; all of that service's network traffic is routed through its own sidecar first, which is what lets the mesh add security, observability, and traffic-control features without any changes to the application's own code. mTLS (mutual TLS) is one of those features: it's a handshake where BOTH sides of a connection present a certificate to authenticate each other, not just the server proving its identity to the client the way ordinary TLS does, and it's the sidecars, not the application, that do this handshake work on the service's behalf.
Structured elaboration.
- Diagnose whether the mesh is the cause. If your tracing captures spans for the mesh sidecar specifically (many service meshes inject their own spans), check how much of the added latency is attributable to the sidecar hop versus your application code; a sidecar span consistently adding meaningful time on every call is direct evidence. If sidecar-specific spans aren't available, a controlled A/B comparison (a subset of traffic routed through the mesh, a subset bypassing it, if your infrastructure allows a temporary bypass for testing) isolates the effect cleanly.
- Check CPU usage attribution specifically, since the question calls it out: is the increased CPU usage in your application's own process, or in the sidecar's process? If the sidecar itself is consuming meaningfully more CPU than expected, that points at the mesh's own overhead (its proxying, its TLS handling, its policy evaluation) rather than anything about your application.
- Quick mitigation options if the mesh is implicated. Configuration changes are usually the fastest lever: check whether mTLS, detailed telemetry collection, or a specific policy feature is more expensive than expected and could be tuned down for now without abandoning the mesh entirely. A bypass for the highest-traffic or most latency-sensitive specific paths (routing those specific calls around the mesh while leaving it in place elsewhere) is a more targeted mitigation than reverting everything.
- Criteria for continuing to optimize versus rolling back entirely. If the overhead is attributable to a SPECIFIC, tunable configuration (not an inherent cost of the mesh architecture itself), continuing to optimize is usually the right call, since the mesh's benefits (consistent mTLS, observability, traffic policy) are still available once tuned. If the overhead persists even after reasonable tuning and appears to be an inherent cost of the mesh's architecture for your specific traffic pattern (very high-volume, latency-sensitive calls where even a well-tuned sidecar hop's overhead is unacceptable), rolling back, at least for those specific highest-sensitivity paths, is the more honest conclusion. A concrete decision criterion: define an acceptable added-latency budget for the mesh up front (for example, no more than 5 to 10% added latency for calls in this service's critical path) and measure your best-tuned configuration against that budget rather than deciding based on general impression.
Worked example. Suppose sidecar-specific trace spans show the mesh consistently adds about 8ms to each call, and separately, sidecar CPU usage is elevated specifically because mTLS is configured to do a full handshake on every connection rather than reusing established connections. Switching to persistent, reused connections between the sidecar and its peers (a common, well-supported mesh configuration option) drops the added latency to roughly 2 to 3ms in testing, comfortably inside a 5 to 10% budget for most of this service's calls. That's a case for continuing to optimize (a specific, fixable configuration issue was the actual cause) rather than rolling back the mesh entirely, since the underlying architecture wasn't inherently too expensive, one specific setting was.
Trade-offs and pitfalls. Rolling back the mesh entirely at the first sign of overhead throws away real benefits (consistent security posture, observability, traffic management) that were presumably worth adopting it for in the first place; exhausting the tuning options first, with a clear latency budget to measure against, avoids an overreaction to what might be a fixable configuration issue. Conversely, sticking with an under-optimized mesh indefinitely because rolling back feels like admitting a mistake is its own trap; a pre-defined, objective latency budget removes some of that emotional weight from the decision by making it a measurement question rather than a judgment call made under pressure.
Unlock Full Question Bank
Get access to all Production Incident Diagnosis and Distributed Systems Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.