Production Incident Diagnosis and Distributed Systems Troubleshooting Questions

Debugging distributed systems under fire: diagnosing latency and reliability regressions, root-causing across service boundaries, reading traces and metrics during an incident, and reasoning about complex production failures. Covers the investigative method for hard-to-reproduce, multi-service problems. The operational counterpart to resilient design.

HardTechnical
64 practiced

A production service shows sporadically high CPU time in the kernel (sys time). Propose how you would use eBPF, bpftrace, or bcc tools to profile syscalls, sample stack traces, and determine whether the cause is kernel-level (for example futex contention, epoll_wait, or network interrupts) or genuinely user-space CPU. Give example bpftrace one-liners or bcc tools you would actually run.

MediumTechnical
93 practiced

A load-balancer health check marks instances unhealthy too aggressively, causing cascading restarts. Given the pseudocode below, identify the problems and propose an improved health-check and backoff strategy.

if cpu_percent > 90:
  consecutive_failures += 1
else:
  consecutive_failures = 0
if consecutive_failures >= 3:
  mark_unhealthy()

Explain your improvements and the reasoning behind them.

MediumTechnical
73 practiced

Given this simplified trace for a single request, identify where the latency spike originates and why:

TraceID: abc123
Spans:

  • gateway (api-gw): duration 50ms
  • auth (service-b): duration 5ms
  • payments (service-c): duration 400ms
    • db-proxy (service-d): duration 380ms
      • db-query: duration 370ms

Explain the steps you would take to confirm the database is the true root cause and what further data you'd collect before concluding.

MediumTechnical
100 practiced

Your stack runs microservices on Kubernetes with Prometheus and Jaeger. Describe three concrete debugging scenarios you've resolved in a similar environment. For each one, include the data sources you checked (logs, traces, metrics), your root-cause-analysis steps, the immediate mitigation, and the long-term fix.

HardTechnical
76 practiced

After adopting a service mesh, your telemetry shows increased latency and CPU usage. Describe how you would diagnose whether the mesh itself is the root cause, the steps you'd take to mitigate the regression quickly (configuration changes, bypassing the mesh for specific paths), and the criteria you'd use to decide between continuing to optimize the mesh configuration or rolling back the adoption entirely.

Unlock Full Question Bank

Get access to all Production Incident Diagnosis and Distributed Systems Troubleshooting interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.