Production Incident Diagnosis and Distributed Systems Troubleshooting Questions
Debugging distributed systems under fire: diagnosing latency and reliability regressions, root-causing across service boundaries, reading traces and metrics during an incident, and reasoning about complex production failures. Covers the investigative method for hard-to-reproduce, multi-service problems. The operational counterpart to resilient design.
A vendor integration suddenly changes its response contract without notice, breaking your clients. Propose an emergency incident response and a longer-term strategy to prevent future vendor-induced breakages, covering contract enforcement, integration testing against the vendor's real API, and the commercial terms you'd push for.
Sample Answer
Direct answer. An unannounced breaking change from a vendor needs an emergency response that treats the vendor as an unreliable dependency for the moment, followed by a longer-term relationship and architecture change that makes the NEXT surprise less damaging.
Structured elaboration.
- Emergency incident response. Confirm the scope of what's actually broken for your clients (which of your features depend on the changed response contract), and apply the fastest safe mitigation: if you can adapt to the new contract quickly (a small parsing or mapping change on your side), that's often faster than waiting for the vendor to revert. If you can't adapt quickly, consider whether you can temporarily fall back to a cached or last-known-good version of whatever data the vendor provides, or gracefully degrade the dependent feature rather than let it fail outright for your own users.
- Communicate with the vendor immediately, both to report the issue (they may not know it's breaking integrators) and to understand their intent (was this deliberate, will they revert, is there a timeline), since that materially changes whether your best move is adapting to their new contract or waiting them out.
- Longer-term: contract enforcement. Push for a formal API contract or SLA with the vendor that includes advance notice for breaking changes, ideally with a defined deprecation window; without this, you're structurally exposed to a repeat of the same surprise.
- Integration testing against the vendor's real API on a schedule, not just at initial integration time: a contract test that runs periodically against the vendor's actual API (not just a mock) would have caught this specific change close to when it happened, rather than only when it caused a live production failure for real users.
- Commercial terms. If this vendor is significant enough to your business, this incident is legitimate leverage to negotiate stronger contractual protections (advance notice requirements, an SLA with remedies for breaking changes without notice) as part of the relationship going forward, not just a technical fix.
Worked example. Suppose the vendor's response contract change turns out to be a field that used to always be present and is now sometimes omitted for a subset of records; if your integration can tolerate treating that field as optional with a sensible default (say, treating a missing status field as 'unknown' rather than crashing), that's a same-day mitigation that doesn't require waiting on the vendor at all. In parallel, adding a scheduled contract test that specifically asserts the shape of the vendor's response (including that this field is present) against their live API on a recurring basis would have caught this exact change within, at most, one test cycle after it shipped, rather than after it caused a customer-facing failure.
Trade-offs and pitfalls. Adapting quickly to a vendor's new contract fixes the immediate problem but can create an awkward situation if the vendor later reverts the change, since your adapted code might then need to handle BOTH the old and new shapes during the transition; keeping the adaptation defensive (tolerant of either shape) rather than assuming the new shape is permanent is a safer default until the vendor confirms their intent. It's also worth being realistic that not every vendor relationship has enough leverage to negotiate a formal advance-notice SLA; where that's not achievable, investing more heavily in your own scheduled contract testing is the fallback that doesn't depend on the vendor's cooperation.
You're on-call for a service and see increased 500 errors concentrated in one endpoint minutes after a deploy went out. Walk through the immediate steps you take in the first 15 minutes: how you determine whether the deploy actually caused the regression versus a coincidental correlation, what dashboards and logs you check first, your mitigation options (rollback, canary rollback, throttling), and how you communicate status to stakeholders.
Sample Answer
Direct answer. In the first 15 minutes the goal is to stabilize user-facing traffic and preserve evidence, in that order; you are not trying to find the root cause yet, you are trying to stop the bleeding and make sure you (or whoever picks this up) can still find the root cause afterward.
Structured elaboration.
- Confirm the deploy is actually correlated before assuming it's the cause. Pull up the deploy timeline and compare it against when the 500s started. If a deploy to this endpoint landed in the same window, that is your leading hypothesis, but check it, don't assume it: some deploys are unrelated and the timing is coincidence.
- Look at the error itself. Open a sample of the failing requests in logs or your APM (application performance monitoring tool, e.g. Datadog or New Relic). A stack trace or specific error code will usually tell you whether this is a code bug in the new release (null pointer, unhandled exception, a broken dependency call) versus something environmental (a config value that didn't get set, a database migration that didn't finish).
- Check dashboards for blast radius. Is this endpoint's error rate the only thing affected, or is latency also up, are other endpoints degrading too, is a specific region or instance pool worse than others? This tells you whether containment can be narrow (this endpoint only) or needs to be broader.
- Mitigate, choosing among your real options. A full rollback is usually the fastest, safest action, since patching forward live under pressure is a common source of a SECOND incident. If a full rollback isn't immediately available (say, other changes have shipped on top of it), a canary rollback, reverting just the newest batch of instances back to the prior version while leaving the rest alone, narrows the blast radius while you confirm the fix works. Throttling the affected endpoint (accepting some added latency or rejecting a fraction of requests) is a third lever worth having ready if neither rollback option is immediately deployable, buying time without needing a code or deploy change at all.
- Preserve evidence as you go. Before you roll back, capture a handful of the actual failing request logs, the relevant dashboard screenshots, and the exact timestamps. Once you roll back, the failure signal disappears and you lose your best evidence for the eventual root-cause writeup.
Worked example. You see the 500s are concentrated in one endpoint, a deploy to that service's code landed 6 minutes before the alert, and a sample of the failing requests shows a NullPointerException in a new code path. That is enough correlation plus evidence to roll back immediately: you trigger the rollback, watch the error rate over the next 2 to 3 minutes, and confirm it returns to baseline. You now have a clean signal (rollback fixed it) and preserved logs, which together make the eventual root-cause writeup straightforward: the new code path didn't null-check a field that's optional in a subset of real traffic.
- Communicate status to stakeholders throughout, not just at resolution. Post a short update to your incident channel or status page as soon as you've confirmed the correlation in step 1: what's affected, the rough blast radius from step 3, and what mitigation you're about to try. Once you roll back, a second short update (what you did, whether it worked, and what's still open) closes the loop. This matters even inside a 15-minute window, because stakeholders (support, other engineers, sometimes leadership) making decisions off silence or guesswork is a common secondary problem during an incident.
Trade-offs and pitfalls. The instinct to 'fix it properly' under pressure, by patching the bug live rather than rolling back, is usually the wrong call: a patch written and shipped in the middle of an active incident has not been reviewed or tested the way the original deploy was, and a second bad deploy during an active incident is a genuinely common failure pattern. The other pitfall is rolling back so fast that you never capture the failing-request evidence, which leaves the eventual root-cause analysis guessing instead of grounded in what you actually saw. A third pitfall is treating stakeholder communication as a nice-to-have that happens after the fact: going quiet during an active regression, even a short one, tends to generate more anxious pings and duplicated investigation than a two-line status update would have cost.
Monitoring shows that adding more instances of a microservice increased average and p95 latency instead of reducing it. Walk through a debugging checklist explaining possible causes (for example shared-resource contention, DNS or iptables issues, connection-pool exhaustion, or leader-election thrashing), how you'd gather evidence for each, and the remedial action for each cause.
Sample Answer
Direct answer. This is a genuinely counter-intuitive result, since adding instances should spread load and reduce latency, so the debugging checklist has to specifically look for mechanisms where MORE instances create MORE overhead or contention rather than simply assuming the scaling itself is broken.
Structured elaboration.
- Shared-resource contention. If the new instances all compete for the same downstream resource (a database, a cache, a shared connection pool with a fixed total size), adding instances doesn't add capacity to that shared resource, it just adds more competitors for the same fixed pie; check whether a shared downstream's own load or connection count grew proportionally to the new instance count, and whether ITS latency (not just this service's) got worse at the same time.
- DNS or iptables-level issues. A load balancer or service-discovery mechanism that takes non-trivial time to register or fully propagate new instances can cause uneven traffic distribution during scale-up (some instances overloaded while others are still ramping up); at the OS level, an inefficient iptables ruleset (iptables is the Linux kernel's packet-filtering and routing system, commonly used under the hood to implement service load-balancing rules in container networking) can scale poorly with a growing number of backend targets, adding real per-packet overhead as the instance count grows.
- Connection-pool exhaustion, from a different angle than shared-resource contention: if EACH instance opens its own pool of connections to a downstream, more instances can mean MORE total connections than the downstream can handle, even if each instance's own pool looks reasonably sized in isolation.
- Leader-election thrashing, if this service participates in any kind of coordination or leader election: more instances competing for leadership, or more instances triggering more frequent rebalancing in whatever coordination mechanism is in use, can itself consume real resources and add latency, especially if the coordination overhead scales poorly with instance count.
- Gather evidence for each candidate rather than guessing. For each of the above, there's a specific, checkable signal: shared-resource metrics correlated with instance count and latency; load-balancer target-registration timing and per-instance traffic distribution; per-instance versus aggregate connection counts against the shared downstream's own limits; and coordination-service metrics (rebalance frequency, election frequency) if relevant.
- Remedial actions per cause. Shared-resource contention needs the shared resource itself scaled or partitioned, not just the calling service. Load-balancer or iptables issues need investigation of the specific networking layer causing the overhead, which may mean a different load-balancing algorithm or container-networking mode. Connection-pool exhaustion needs a TOTAL connection budget considered across all instances, not just per-instance. Leader-election thrashing needs either fewer participants in the coordination (a subset act as candidates, not every instance) or a coordination mechanism that scales better.
Worked example. Suppose per-instance connection-pool size is a fixed 20 connections to a shared database, and the database's own max-connections limit is 200; at 8 instances, that's 160 total possible connections, comfortably under the limit. Scaling to 15 instances pushes the theoretical maximum to 300, which exceeds the database's 200-connection limit; once actual usage approaches that ceiling, connections start queueing or getting rejected, and EVERY instance (not just the newest ones) experiences worse latency waiting for a connection slot, which is exactly the paradox described: more instances made the shared resource, not any individual instance, the bottleneck. The fix is either reducing per-instance pool size as instance count grows (keeping the total bounded) or increasing the database's connection limit and capacity to match, with the total connection budget explicitly tracked as a fleet-wide constraint rather than a per-instance setting nobody re-evaluates as the fleet grows.
Trade-offs and pitfalls. The core lesson worth internalizing here is that per-instance settings (pool sizes, timeouts, cache sizes) that look reasonable in isolation can become a fleet-wide problem purely from being multiplied across a growing instance count; any config that's 'per instance' should be evaluated against what happens at your MAXIMUM realistic instance count, not just today's count. It's also worth being skeptical of your own instinct here: 'add more instances' is such a common, usually-correct scaling response that it's easy to reach for it again as the FIX when it was actually the trigger.
A load-balancer health check marks instances unhealthy too aggressively, causing cascading restarts. Given the pseudocode below, identify the problems and propose an improved health-check and backoff strategy.
if cpu_percent > 90:
consecutive_failures += 1
else:
consecutive_failures = 0
if consecutive_failures >= 3:
mark_unhealthy()
Explain your improvements and the reasoning behind them.
Sample Answer
Direct answer. The given health check is a ticking time bomb: it counts consecutive failures using only CPU percentage, with no recovery signal and no distinction between a genuinely dead instance and one that's merely busy, so it will mark healthy-but-loaded instances unhealthy and can create a feedback loop where removing 'unhealthy' instances increases load on the survivors, pushing them past the same threshold too.
Structured elaboration.
- The core bug: no positive signal, only a negative one. The pseudocode only ever asks 'is CPU high', never 'can this instance actually serve a request correctly'; a genuinely responsive instance under a legitimate, temporary CPU spike (a garbage-collection pause, a burst of traffic) gets treated identically to an instance that's truly wedged. A real health check should probe actual request-serving capability (a lightweight endpoint that exercises the real request path, or at least confirms the process can respond at all), not infer health indirectly from one resource metric.
- No hysteresis (requiring a condition to hold steadily for a while before flipping state, so a value bouncing around the threshold doesn't cause rapid back-and-forth decisions) or recovery path shown. The pseudocode increments on failure and resets to 0 on any single success, which actually makes it MORE trigger-happy than a stable threshold: an instance oscillating right around 90% CPU could bounce between counted and reset in a way that never quite reaches 3 to mark unhealthy, or conversely could flap in and out of the unhealthy state repeatedly, causing exactly the cascading-restart pattern described.
- No backoff on the resulting action.
mark_unhealthy()fires immediately at the threshold with no apparent cooldown or gradual response; a health-check system that immediately and simultaneously restarts (or removes from rotation) EVERY instance that crosses the threshold at once, during a real, shared traffic-driven CPU spike, can remove capacity exactly when it's most needed, worsening the load on whatever instances remain. - No visibility into the DECISION. There's no logging or metric emitted at each step, so if this logic behaves unexpectedly in production, there's nothing to look at afterward except the fact that a restart happened; any health-check logic making a consequential decision should emit why.
- Improved design. Check actual responsiveness (a lightweight but real request), not just CPU, as the PRIMARY signal, with CPU as at most a secondary contributing factor. Require sustained failure across BOTH consecutive checks and a minimum wall-clock duration (not just a raw consecutive count, which can be gamed by a fast check interval), and add a distinct, harder-to-hit threshold for consecutive SUCCESSES needed to recover, avoiding a single lucky success resetting a genuinely struggling instance's count to zero. Stagger or rate-limit how many instances can be marked unhealthy within a short window, so a shared, correlated CPU spike across the whole fleet doesn't trigger a mass, simultaneous removal.
Worked example. Say five instances all cross 90% CPU within the same 10-second window because of a genuine, if temporary, traffic surge, each running the given logic with a 10-second check interval; all five reach consecutive_failures = 3 (30 seconds) at roughly the same time and all get marked unhealthy simultaneously, removing 5 instances' worth of capacity from a pool that was already under real load, which very plausibly pushes the REMAINING instances' CPU even higher, triggering the same logic on them next, a classic self-reinforcing cascade. With a stagger rule (say, no more than 1 instance per 30-second window gets marked unhealthy from this signal, forcing the rest to wait and re-check) combined with an actual request-responsiveness check as the primary signal (which would likely show all five instances ARE still successfully serving requests, just slowly, under the genuine load spike), none of the five would have been marked unhealthy at all in this scenario, since high CPU under real, successfully-served load is not the same failure as an unresponsive instance.
Trade-offs and pitfalls. A more conservative health check (requiring longer sustained failure, checking actual responsiveness) trades faster removal of genuinely broken instances for fewer false positives; that trade is almost always worth it, since the cost of a truly dead instance staying in rotation a bit longer is usually much smaller than the cost of a self-reinforcing mass-removal cascade. Staggering removals adds real complexity (some shared state or coordination across instances, rather than each instance deciding independently), which is a genuine engineering cost worth taking seriously rather than hand-waving away.
A microservice has become noisy and occasionally causes cascading failures in upstream services. Outline immediate mitigation steps (configuration and network-level), medium-term fixes (code or architecture changes), and long-term remediation to prevent recurrence. Specify the instrumentation you'd add to verify the improvements actually worked and the governance you'd put in place to limit future regressions.
Sample Answer
Direct answer. A noisy microservice causing intermittent cascading failures calls for a layered response: contain the blast radius immediately, fix the noisy behavior at a medium-term horizon, and put guardrails in place so a regression like this can't silently recur.
Structured elaboration.
- Immediate: configuration and network-level. Rate-limit or circuit-break calls TO the noisy service from its callers, so its bad behavior can't consume unlimited resources in the services that depend on it. If the noisy service is the one making excessive outbound calls (rather than being slow to respond), rate-limit or throttle IT at the network or gateway level. These are fast to apply and don't require a code change, which matters when you need to stop the bleeding now.
- Medium-term: code and architecture changes. Find and fix the actual cause of the noisiness: this is often a retry policy with no backoff or cap, a missing timeout that lets slow calls pile up, or a bug that occasionally sends a burst of redundant requests. Add proper backoff, timeouts, and request coalescing where the investigation points.
- Long-term: remediation to prevent recurrence. Add resource isolation (bulkheads) between this service and its callers so a future regression is contained structurally, not just by a runtime rate limit someone has to remember exists. Add automated tests or canary checks that specifically catch retry-storm or excessive-call-volume patterns before they reach full production traffic.
- Instrumentation to verify improvements. Add or confirm metrics for the specific behavior you're fixing (outbound call rate per instance, retry rate, circuit-breaker trip frequency) so you can see directly whether the fix worked, rather than inferring it indirectly from the absence of incidents.
- Governance to limit future regressions. This might mean a checklist or review step for any new retry logic (since that's the recurring theme across steps 1 to 3), or an automated linter/test that flags retry code with no backoff or cap during code review, so this class of bug is caught before it ships rather than after it causes an incident.
Worked example. Suppose the investigation finds the noisy service occasionally enters a state (triggered by a specific downstream timeout) where it retries a failed call up to 10 times with only a 50ms fixed delay between attempts and no jitter, so a brief downstream hiccup turns into a 10x request-volume burst from this one service. The immediate fix is capping the retry count and adding exponential backoff (waiting progressively longer between each retry attempt instead of retrying immediately) with jitter (a small random delay added so many clients retrying at once don't all hit the service at the exact same instant) in the code (medium-term), while a rate limit at the gateway (immediate) protects callers in the meantime. The long-term guardrail is a code-review checklist item (or an automated check) requiring any new retry logic to specify a max attempt count and a backoff strategy explicitly, since a fixed-delay, uncapped retry is exactly the pattern that caused this.
Trade-offs and pitfalls. Immediate rate-limiting protects the system but can also mean legitimate requests get throttled during the window before the real fix ships, which is a real cost that's usually still worth paying to avoid a wider cascade. The governance step is easy to skip once the immediate fire is out, but it's the piece that actually prevents the NEXT service from shipping the same uncapped-retry pattern; without it, this remediation only fixes one instance of a repeatable class of bug.
Unlock Full Question Bank
Get access to all 49 Production Incident Diagnosis and Distributed Systems Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.