Production Incident Diagnosis and Distributed Systems Troubleshooting Questions

Debugging distributed systems under fire: diagnosing latency and reliability regressions, root-causing across service boundaries, reading traces and metrics during an incident, and reasoning about complex production failures. Covers the investigative method for hard-to-reproduce, multi-service problems. The operational counterpart to resilient design.

HardTechnical
72 practiced

Cross-region asynchronous replication sometimes lags substantially during traffic peaks, causing stale reads. Propose monitoring and alerting for replication lag, compensation logic to avoid serving wrong data (reading from primary, session affinity, or degrading a feature), and architectural changes to reduce the lag while balancing cost and latency.

MediumTechnical
73 practiced

Given this simplified trace for a single request, identify where the latency spike originates and why:

TraceID: abc123
Spans:

  • gateway (api-gw): duration 50ms
  • auth (service-b): duration 5ms
  • payments (service-c): duration 400ms
    • db-proxy (service-d): duration 380ms
      • db-query: duration 370ms

Explain the steps you would take to confirm the database is the true root cause and what further data you'd collect before concluding.

MediumTechnical
116 practiced

You see conflicting observability signals: tail latency (p95) is up, but the overall error rate is down and throughput is steady. Walk through how you would investigate this, what quick experiments or probes you would run, and how you'd make an operational decision while minimizing customer impact.

HardSystem Design
63 practiced

Design a comprehensive debugging and mitigation strategy for an intermittent production outage that affects about 1% of users across multiple regions in a microservices architecture. Cover the instrumentation you'd add, how controlled rollouts (canaries or feature flags) help isolate the cause without widening the blast radius, the distributed tracing you'd rely on, and how you'd check for cross-region consistency and data-replication issues as a possible cause.

EasyTechnical
57 practiced

Write a runbook fragment for an on-call engineer to follow when a region-wide network partition causes partial failures. Include immediate mitigation steps, prioritized checks, escalation paths, and recovery-validation steps that confirm there is no data loss or inconsistent state across services once the partition heals.

Unlock Full Question Bank

Get access to all 34 Production Incident Diagnosis and Distributed Systems Troubleshooting interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.