Production Incident Diagnosis and Distributed Systems Troubleshooting Questions
Debugging distributed systems under fire: diagnosing latency and reliability regressions, root-causing across service boundaries, reading traces and metrics during an incident, and reasoning about complex production failures. Covers the investigative method for hard-to-reproduce, multi-service problems. The operational counterpart to resilient design.
Cross-region asynchronous replication sometimes lags substantially during traffic peaks, causing stale reads. Propose monitoring and alerting for replication lag, compensation logic to avoid serving wrong data (reading from primary, session affinity, or degrading a feature), and architectural changes to reduce the lag while balancing cost and latency.
Sample Answer
Direct answer. Because the lag is specifically PEAK-dependent, the investigation should focus on what's different about peak traffic (volume, a specific workload pattern, contention with other jobs) rather than treating the replication pipeline as uniformly broken.
Structured elaboration.
- Quantify the relationship between load and lag precisely. Plot replication lag against write throughput (or total traffic) over enough days to see the pattern clearly: does lag start climbing at a specific, identifiable throughput threshold, or does it scale gradually with load the whole time? A clear threshold points at a specific resource hitting a ceiling (network bandwidth between regions, replication-thread capacity); a gradual scaling suggests replication is simply always somewhat throughput-bound and peaks just push it further.
- Check what specifically is the bottleneck during a lag spike. Network bandwidth between regions (replication competing with other cross-region traffic for the same link), the replication mechanism's own throughput ceiling (a single-threaded or limited-parallelism replication stream can't keep up regardless of network capacity), or contention on the target region's write path (the replica applying changes is itself resource-constrained) are the most common candidates, and each has a different fix.
- Design monitoring and alerting for replication lag specifically, not just for downstream symptoms like stale reads: an alert that fires when lag crosses a threshold BEFORE it causes visible staleness gives you lead time to react (shed load, or trigger compensation logic) before users are affected.
- Compensation logic for the read side, while lag exists. Reading from the primary for known-sensitive or known-recent data, session affinity (route a given user consistently to the same region so they at least see a self-consistent view even if it's slightly stale relative to global truth), and explicitly degrading a feature that depends on fresh cross-region data (rather than silently serving wrong data) are the standard toolkit; which one fits depends on the specific feature and its tolerance for staleness.
- Architectural changes to reduce lag itself, balancing cost and latency. Options include increasing replication parallelism or bandwidth (a cost trade-off), switching specific critical paths from fully async to semi-synchronous replication (a latency trade-off, since the write now waits for at least partial cross-region acknowledgment), or partitioning data so less of it needs cross-region replication at all (an architecture trade-off that may not fit every data model).
Worked example. Suppose the lag-versus-throughput plot shows a clear knee: lag stays under 200ms up to about 8,000 writes per second, then climbs steeply, reaching several seconds above roughly 12,000 writes per second, right around your typical peak. Checking the replication mechanism's own metrics shows the replication stream is single-threaded and its own throughput caps out right around 8,000 to 9,000 writes per second, independent of available network bandwidth (which has plenty of headroom). That converges on replication PARALLELISM as the bottleneck, not network capacity: increasing the number of parallel replication streams (partitioned by key range, for example) would be expected to raise that throughput ceiling roughly in proportion to the added parallelism, whereas adding network bandwidth alone would not have helped, since the bottleneck was never bandwidth.
Trade-offs and pitfalls. It's easy to assume 'network' is the bottleneck for cross-region issues by default, but as this example shows, the replication mechanism's own throughput ceiling is at least as common a cause and needs a completely different fix (parallelism, not bandwidth). Session affinity as a compensation strategy has its own trade-off: it improves per-user consistency but can create uneven load if a disproportionate share of active users happen to be affinity-routed to the same region during a regional traffic imbalance.
Given this simplified trace for a single request, identify where the latency spike originates and why:
TraceID: abc123
Spans:
- gateway (api-gw): duration 50ms
- auth (service-b): duration 5ms
- payments (service-c): duration 400ms
- db-proxy (service-d): duration 380ms
- db-query: duration 370ms
- db-proxy (service-d): duration 380ms
Explain the steps you would take to confirm the database is the true root cause and what further data you'd collect before concluding.
Sample Answer
Direct answer. Reading the spans, the request makes three sequential hops (gateway, then auth, then payments), and within the payments hop the 400ms is almost entirely accounted for by the 380ms db-proxy span beneath it, which is itself almost entirely the 370ms db-query span, so the database query is the dominant contributor to this request's latency; the gateway and auth hops (50ms and 5ms) are not meaningfully part of the problem.
Structured elaboration.
- First, read the SHAPE of the tree correctly, before computing anything.
gateway,auth, andpaymentsare three separate, sibling spans in this trace: one request flowing sequentially through a gateway hop, an auth check, and a payments call, one after another, not one nested inside another. Onlypaymentshas a child of its own (db-proxy, which in turn has its own child,db-query). That distinction matters because it decides which spans you even need to subtract anything from. - Compute each span's OWN time (its duration minus its children's durations), only where a child exists to subtract.
gatewayandauthhave no children shown in this trace, so their own time is simply their full duration: 50ms and 5ms.paymentshas one child (db-proxyat 380ms), so its own time is 400 minus 380 = 20ms: the payments SERVICE's own code cost only 20ms, the rest of its 400ms is time spent waiting on its call todb-proxy.db-proxyhas one child (db-queryat 370ms), so its own time is 380 minus 370 = 10ms.db-queryhas no child shown, so its own time is its full 370ms.span total duration child duration own time gateway 50ms none 50ms auth 5ms none 5ms payments 400ms 380ms (db-proxy) 20ms db-proxy 380ms 370ms (db-query) 10ms db-query 370ms none 370ms - Add up the own-times and see where the total request time actually went. 50 + 5 + 20 + 10 + 370 = 455ms, which is the full sequential request time (50 + 5 + 400). db-query's 370ms alone is roughly 370/455 ≈ 81% of that total, by a wide margin the single largest contributor; gateway and auth together are only about 12%.
- Confirm it's the query itself, not something around it. Before concluding 'the database is slow', check three distinct possibilities that all show up as a slow
db-queryspan: the query itself is slow (missing index, bad query plan, lock contention), the connection pool made the caller wait before the query even started (which would usually show as extra time BEFORE the query span begins, not inside it), or the database host itself is resource-constrained (high CPU, disk I/O, or replication lag if this is a replica). - Pull further data to distinguish those. Check the database's own slow-query log or
EXPLAINplan for this query around the incident window; check whether the query's duration is consistently ~370ms or spiking intermittently (consistent points to a plan or index problem, spiky points to contention or a noisy neighbor on the host); check the database host's own CPU, I/O, and lock-wait metrics for the same window. - Confirm before you commit to a fix. If
EXPLAINshows a sequential scan where an index should be used, or the slow-query log shows this exact query is consistently the slowest one running, that confirms the query itself and points to an index or query-shape fix. If instead the database host's CPU or I/O is saturated across many queries at once, the fix is capacity or isolating this workload, not this one query.
Worked example. Suppose the slow-query log shows this exact query pattern (a lookup by a non-indexed customer_email column) running consistently between 350 and 390ms throughout the incident window, while other indexed queries against the same table stay under 10ms. That consistent, query-specific slowness (not database-wide) points squarely at a missing index on customer_email: adding it would be expected to bring this query down to roughly the same low-single-digit-millisecond range as the other indexed lookups on that table. Since db-query is nested three levels deep (inside db-proxy, inside payments), fixing it collapses the whole chain: db-proxy's own 10ms plus a near-zero query time would bring payments down from 400ms to roughly its own 20ms plus a few milliseconds of query time, and the full request (gateway + auth + payments) would drop from about 455ms to well under 100ms.
Trade-offs and pitfalls. The most common mistake reading a trace like this is to see 'payments took 400ms' and start investigating the payments service's own code, when the trace is telling you the payments service's code only cost about 20ms and the other 380ms is entirely its call to the database, three levels down. The second most common mistake is the one this answer corrects itself on above: subtracting a span's duration from a SIBLING's duration because they happen to be listed near each other, rather than checking which spans are actually nested inside which. Always attribute time to the span whose OWN duration (after subtracting only its ACTUAL children) is largest, and follow the tree down until a span has no slower child left. It's also worth checking for cache misses: if this query is normally served from a cache and the cache was cold or evicted, the fix might be restoring the cache rather than touching the database at all.
You see conflicting observability signals: tail latency (p95) is up, but the overall error rate is down and throughput is steady. Walk through how you would investigate this, what quick experiments or probes you would run, and how you'd make an operational decision while minimizing customer impact.
Sample Answer
Direct answer. Since throughput and overall error rate look fine, the tail-latency increase is most likely confined to a specific subset of requests rather than a system-wide degradation, so the investigation should look for what's different about that slice before considering any broad, system-wide action.
Structured elaboration.
- Isolate what's actually in the slow tail. Filter to the specific requests contributing to the p95 increase and look for a shared trait: a specific endpoint, customer, payload characteristic, or time-of-day pattern, the same filter-to-the-outliers-and-look-for-a-shared-trait method that works for any tail-only latency regression.
- Run quick, low-risk probes rather than broad changes. A synthetic request against the suspected slow path, or comparing a handful of real slow requests' traces against a handful of fast ones for the same endpoint, can confirm or rule out a hypothesis in minutes without touching production configuration.
- Consider that a healthy aggregate can mask a real, contained problem. Overall error rate being down and throughput being steady are reassuring but don't rule out a genuine issue affecting a minority of requests badly enough to move p95 without moving the headline numbers; don't let the healthy aggregate talk you out of investigating.
- Make an operational decision under uncertainty, prioritizing customer impact. If the affected slice is small and its impact on real users is mild (a moderate latency increase, not outright failures), it may be reasonable to keep investigating without an urgent mitigation. If the affected slice, though small in volume, represents a high-value case (a specific large customer, a critical workflow), a more urgent, even if narrowly-scoped, mitigation is warranted despite the healthy aggregate.
- Avoid overcorrecting based on an aggregate metric alone. Since error rate is actually down, a broad rollback or a wide mitigation aimed at 'fixing everything' risks solving a problem that doesn't exist system-wide while potentially disrupting the parts that are currently working fine.
Worked example. Suppose filtering the slow requests contributing to the p95 increase shows they're concentrated on one endpoint that recently started making an additional downstream call for a subset of requests (say, an enrichment lookup that only fires when a specific optional field is present in the request). That call adds real latency only for the requests that trigger it, which is small enough as a fraction of total traffic to leave overall throughput and error rate essentially unchanged, while still being large enough within that specific slice to move the overall p95. The operational decision, since the enrichment call isn't failing (which is why error rate is unaffected) and only some requests need it, might be to make that specific call async or cached rather than synchronous and blocking, targeting the fix at the actual mechanism rather than anything system-wide.
Trade-offs and pitfalls. The main risk here is either extreme: dismissing a p95 anomaly because the aggregate looks healthy, or overreacting with a broad, disruptive mitigation because ANY anomaly feels alarming. The middle path, quickly isolating the affected slice and matching the response's urgency and scope to what you actually find, protects users without unnecessary disruption to the parts of the system that are working fine.
Design a comprehensive debugging and mitigation strategy for an intermittent production outage that affects about 1% of users across multiple regions in a microservices architecture. Cover the instrumentation you'd add, how controlled rollouts (canaries or feature flags) help isolate the cause without widening the blast radius, the distributed tracing you'd rely on, and how you'd check for cross-region consistency and data-replication issues as a possible cause.
Sample Answer
Direct answer. Because this outage is intermittent, low-volume (about 1% of users), and spans multiple regions, the core challenge is generating enough signal to actually see the pattern, so the strategy centers on instrumentation and safe, incremental investigation rather than a single decisive test.
Structured elaboration.
- Instrumentation first. Before you can find an intermittent, low-volume problem, you need enough detail captured on EVERY request (or a high enough sample rate) to distinguish the roughly 1% of failing requests from the 99% that succeed: request-level tracing with enough span detail to see which service and which call is implicated when a failure does occur, plus structured logging that includes enough context (region, a request or trace ID, relevant feature flags) to group failures once you have several examples.
- Look for what the ~1% have in common. Once you can reliably capture failing requests, check whether they cluster on a specific region, a specific data-replication path, a specific user segment, or a specific downstream dependency; a 1% failure rate that's actually 100% of requests hitting one specific, rarely-used code path looks very different from a genuinely random 1%.
- Check cross-region consistency and data-replication specifically, since those were called out as plausible for a reason: if this system relies on replicated state across regions, verify whether the failures correlate with replication lag or a conflict-resolution edge case, which would explain both the low rate (only requests landing during a lag window are affected) and the multi-region spread (it's a property of the replication mechanism, not any one region's infrastructure).
- Use controlled, incremental rollouts as an investigative tool, not just a deployment safety net. Canaries (a small, live slice of production instances running a candidate fix or extra instrumentation) and feature flags (toggling a specific code path for a small percentage of traffic without a full deploy) are the two practical mechanisms for this: both let you compare a treated slice of real traffic against the untreated rest, testing a hypothesis or a fix safely without committing to a global change based on a guess.
- Communicate and close the loop. Because the impact, while real, is low-volume and intermittent, this is a case where a clear internal update on current understanding and next steps (even before root cause is confirmed) helps other engineers avoid duplicating investigation, and a documented resolution once found closes out the incident properly.
Worked example. Say enhanced tracing on a sample of requests reveals that the roughly 1% of affected requests all touch a specific data path that reads a value which was just written in a DIFFERENT region within the prior second, a classic cross-region replication-lag window. The fix, reading from the primary region for that specific data path when a write is recent (a bounded staleness check) rather than always reading from the nearest replica, directly targets the replication-lag mechanism rather than treating the symptom generically. Rolling that fix out to 5% of traffic first and comparing the affected-request rate against the 95% control group would confirm the fix works before a full rollout.
Trade-offs and pitfalls. The main risk with a low-volume, intermittent problem is under-instrumenting and never generating enough signal to see the pattern at all, in which case you're stuck reasoning from a handful of anecdotal reports; investing in better capture BEFORE you're confident in a hypothesis is often the highest-leverage first step, even though it doesn't feel like 'real' progress. It's also worth being honest that 1% affecting real users over enough volume is still a real number of people, so the investigation deserves genuine urgency even though it wouldn't show up as a dramatic spike on a top-line dashboard.
Write a runbook fragment for an on-call engineer to follow when a region-wide network partition causes partial failures. Include immediate mitigation steps, prioritized checks, escalation paths, and recovery-validation steps that confirm there is no data loss or inconsistent state across services once the partition heals.
Sample Answer
Direct answer. A region-wide network partition means the on-call engineer is dealing with genuine uncertainty about which side of the partition owns truth, so the runbook has to prioritize NOT making the split worse before it prioritizes restoring full service.
Structured elaboration.
- Immediate mitigation steps. Confirm the partition's scope first (is it truly the whole region, or a subset of services within it) using an independent monitoring path that doesn't itself depend on the partitioned region, since monitoring that routes through the affected region can give a false picture. If there's a designated failover region or a documented authoritative side for this partition scenario, follow that designation rather than deciding ad hoc under pressure; if there isn't, that's itself a gap this incident should surface.
- Prioritized checks. Which services have region-local state that could diverge if both sides keep accepting writes (the split-brain risk); which services are stateless or read-only and can safely continue serving from either side without correctness risk; and whether any cross-region dependency (a shared queue, a shared database) is itself affected by the partition or still reachable from one or both sides.
- Escalation paths. Who owns the decision to formally fail over the region (this is usually a decision above a single on-call engineer for anything with real data-correctness stakes), and what's the communication chain to notify affected teams and, if customer-facing impact is significant, support or communications teams.
- Recovery validation steps once the partition heals. Before declaring the incident over, explicitly check for divergence: did both sides of the partition accept writes to the same data during the split, and if so, has that been reconciled (following the same logic as any split-brain reconciliation) before resuming normal, unrestricted operation. Confirm no data loss by comparing a checksum or count of critical data against what was expected, not just by observing that services report healthy again.
Worked example. A concrete instantiation: region A and region B lose connectivity between them for 12 minutes. The runbook's first check (independent, cross-region monitoring) confirms it's a true regional partition affecting all inter-region traffic, not a single service issue. Following the documented authoritative-side designation, region A continues accepting writes for the shared data store while region B is placed into a read-only, degraded mode for anything requiring cross-region coordination. Stateless read services in both regions continue serving local traffic normally throughout. Once connectivity restores, before lifting region B's read-only mode, a reconciliation check confirms no writes were accepted on B's side during the partition (since it was correctly held read-only) and no data loss occurred; if B HAD somehow accepted writes despite the read-only mode (a bug in the mode's enforcement, for example), that would be a second, more serious finding requiring the split-brain-style reconciliation before recovery is complete.
Trade-offs and pitfalls. The instinct to restore full service everywhere as fast as possible has to be weighed against the risk of both sides having independently accepted writes, since resuming full bidirectional operation before confirming no divergence occurred can permanently and silently corrupt data; a slightly slower, verified recovery is almost always the better trade for anything with real correctness stakes. It's also worth pre-deciding the authoritative-side designation and the stateless-vs-stateful service list BEFORE an incident, in the runbook itself, since deciding either under active pressure is slower and more error-prone than following a pre-made call.
Unlock Full Question Bank
Get access to all 34 Production Incident Diagnosis and Distributed Systems Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.