Systematic Debugging and Root Cause Analysis Questions
Methodically diagnosing failures and identifying their true cause. Covers hypothesis-driven debugging, bisection and instrumentation, full-stack and production diagnosis, debugging under pressure, and root-cause analysis that prevents recurrence. Emphasizes a repeatable process over guesswork.
Design a Python helper function that uses git CLI to perform a bisect over a list of recent commits to find the commit that introduced a failing test. Provide the function interface, describe how it runs tests, handles flaky tests, and what assumptions you make about the environment.
Sample Answer
git bisect performs binary search over commit history: you mark one commit known-good and one known-bad, it checks out the midpoint, you (or a script) report good/bad, and it repeats, isolating the first bad commit in O(log n) tests. A Python helper can drive this directly through the git CLI via subprocess, one commit at a time, instead of leaving the whole loop to a shell one-liner.
Function interface
import subprocess
def bisect_find_regression(
good_commit: str,
bad_commit: str,
test_command: list,
repo_path: str = ".",
retries: int = 5,
failure_threshold: int = 1,
) -> str:
"""
Uses `git bisect` (via the git CLI) to find the first commit between
good_commit and bad_commit that introduced a failure of `test_command`.
Returns the SHA of the first bad commit.
Assumptions:
- repo_path is a clean git checkout with no uncommitted changes: bisect
checks out different commits in place, and this function does not
stash or restore working-tree edits.
- test_command is a list suitable for subprocess.run, e.g.
["pytest", "-k", "regression_case"].
- a commit whose build/setup itself fails (not the test under
investigation) must be treated as untestable so bisect skips it
instead of misattributing the regression to an unrelated break;
this function does that by checking for a distinguished exit code.
- failure_threshold / retries handle test flakiness: the test is run
up to `retries` times on a commit and only counted "bad" if it
fails at least `failure_threshold` times, so one flaky failure on
an actually-good commit does not derail the search.
"""
subprocess.run(["git", "bisect", "start"], cwd=repo_path, check=True)
subprocess.run(["git", "bisect", "bad", bad_commit], cwd=repo_path, check=True)
subprocess.run(["git", "bisect", "good", good_commit], cwd=repo_path, check=True)
def _verdict_for_current_checkout() -> str:
failures = 0
for _ in range(retries):
proc = subprocess.run(test_command, cwd=repo_path)
if proc.returncode == 125:
return "skip"
if proc.returncode != 0:
failures += 1
return "bad" if failures >= failure_threshold else "good"
try:
while True:
verdict = _verdict_for_current_checkout()
out = subprocess.run(
["git", "bisect", verdict],
cwd=repo_path,
capture_output=True,
text=True,
check=True,
)
if "is the first bad commit" in out.stdout:
return out.stdout.splitlines()[0].split()[0]
finally:
subprocess.run(["git", "bisect", "reset"], cwd=repo_path, check=True)
How it runs tests
Each iteration runs test_command against whatever commit git bisect has currently checked out, using its process exit code as the verdict: 0 means good, a nonzero code (other than 125) means bad, and 125 means "cannot be tested, skip this commit" (the convention git bisect run itself uses, so a wrapper script that already follows it plugs straight into this function's _verdict_for_current_checkout). This was verified end to end against a small real git repository: five commits, a passing test on the first two, a bug introduced on commit four (test starts failing) with an unrelated no-op commit after it, and bisect_find_regression correctly returned the exact commit that introduced the failure.
Handling flaky tests
A test that is flaky will make bisect converge on the wrong commit, because a single failing run at a good commit looks identical to a real regression. _verdict_for_current_checkout runs the test up to retries times per commit and only reports "bad" once at least failure_threshold of those runs failed, so an occasional flake on a good commit does not get reported as bad, while a commit that fails consistently still gets reported bad quickly. Tune retries/failure_threshold to the test's known flake rate: a test that fails 1 in 20 runs when truly "good" needs enough retries that a false-bad verdict is unlikely (e.g. requiring 2+ failures out of 5 runs), while a very stable test can use retries=1.
Assumptions about the environment
- The repo has a linear, bisectable history between
good_commitandbad_commit(a single monotonic transition, not multiple unrelated interleaved regressions in the range). test_commandis runnable from a fresh checkout with no manual setup steps the script doesn't perform (dependencies are installed as part oftest_commanditself, or already present).- The caller has already established that
good_committruly passes andbad_committruly fails before starting; the function does not re-verify the endpoints.
Trade-offs and pitfalls
Bisect assumes a single monotonic transition from good to bad; if the bug is intermittent even in its "bad" state, or if multiple unrelated changes landed in the range, the repeat-N-times guard above is what keeps the search from being derailed, and batched-commit CI setups should bisect at the smallest unit of change actually available, not the batch. Because git bisect reset runs in a finally block, the repo is always left back on its original branch even if the loop is interrupted or a commit is genuinely untestable end to end.
You are on-call and receive alerts that a production web application is returning 500 errors and experiencing high latency from multiple regions. Describe, step‑by‑step, a systematic troubleshooting process you would follow to identify the root cause. Include: what data and artifacts you would collect first, which commands/tools you would run on affected hosts, how you'd triage service vs network vs DB vs infrastructure, and how you'd prioritize actions under time pressure.
Sample Answer
A strong candidate starts from a fixed loop, not a guess: gather signal, form a hypothesis, run the cheapest test that could disprove it, and only then act.
The loop
- Collect first, don't touch anything. Pull the alert/report, recent deploys and config changes, and the three signal types: metrics (what changed and when), logs (what the service says happened), traces (where time went across components).
- Triage the layer before the cause, using a fixed set of checks per layer (see below): does it correlate with one host, one region, one dependency, or all of them? A failure isolated to one instance points at that instance; a failure across all instances after a deploy points at the deploy; a failure correlated with a dependency's own error rate points downstream.
- Form one falsifiable hypothesis at a time ("the new deploy is the cause") and pick a test that would disprove it cheaply (check whether the errors started at the exact deploy timestamp, or whether rolling back one canary host clears them). Avoid changing five things at once.
- Prioritize under time pressure: mitigate first (rollback, scale out, fail over) if user impact is ongoing, investigate the true cause in parallel or after.
Commands and tools per layer, and how to triage between them
Run these roughly in parallel across a couple of affected hosts, not sequentially one host at a time:
- Service/application layer:
kubectl get pods -o wideandkubectl describe pod <pod>(orsystemctl status <service>on VMs) to check restart counts and recent events;kubectl logs -f <pod> --previousorjournalctl -u <service> -ffor the exact error at the moment of the alert;top/htoporps aux --sort=-%cpu,-%memfor CPU/memory pressure on the process itself; for managed runtimes,jstack <pid>orjcmd <pid> Thread.printto check for stuck threads. Signal that this is the layer: errors and restarts correlate with specific pods/hosts or with the deploy timestamp, not with a single dependency or the network path. - Network layer:
curl -vagainst the exact endpoint from an affected host to separate DNS/TLS/connect time from application response time;dig/nslookupfor DNS resolution issues;traceroute/mtrfor path/latency between hops;ss -sornetstat -antpfor socket/connection-state saturation (too manyTIME_WAIT, exhausted ephemeral ports);tcpdumpon a specific host if a particular hop is suspected. Signal:curl -v's connect/TLS phase is slow while the app's own processing time (visible in traces) is normal, or the problem tracks a specific region/CDN edge rather than a specific service version. - Database layer: active/slow query lists (
SHOW PROCESSLISTon MySQL,pg_stat_activityon Postgres), lock waits (SHOW ENGINE INNODB STATUS,pg_locks), and the app's own connection-pool metrics (checked-out connections near the pool limit). Signal: request latency traces show most of the time inside the DB span, and DB-side query/lock metrics show a corresponding spike at the same timestamp. - Infrastructure layer:
kubectl describe node/kubectl top nodefor node-level CPU/memory/disk pressure,dmesgorjournalctl -kfor OOM-killer or kernel-level events, and the cloud provider's status page or recent autoscaling/capacity events. Signal: the failure correlates with a specific node, availability zone, or a capacity/autoscaling event rather than with a code deploy or a single dependency.
The layer whose checks show a signal exactly aligned with the alert's onset time is the one to dig into first; the others should still be glanced at briefly to rule out a compounding factor, but don't get equal depth until the primary layer is ruled out.
Worked example
A service starts returning 500s at 14:02. A deploy went out at 14:00. Metrics show error rate flat on hosts still running the old build and elevated only on hosts running the new one. That single comparison (same traffic, different build, different outcome) is strong evidence for the deploy as root cause, tested in under a minute using data you already have, before touching any code.
Trade-offs and pitfalls
The most common mistake is skipping straight to "it's probably the database" because that's where the last incident was, without checking whether this failure actually correlates with DB latency. A hypothesis not tested against data is a guess wearing an RCA costume. The other common failure is fixing the first plausible thing that appears in the logs, when it's a symptom of an earlier upstream cause; correlating the alert time against the deploy/change timeline first avoids that trap.
You're in a live incident and have two straightforward options: rollback the latest deploy or toggle a feature flag that should revert behavior. Walk through your decision-making process: how you assess risk, preserve data and logs, perform the rollback or toggle, and verify that service behavior is restored. Include communication steps and how you would avoid causing more disruption.
Sample Answer
When the two live-incident options are specifically a rollback of the latest deploy or toggling a feature flag that should revert the new behavior, the choice comes down to which one restores the previous behavior more completely and more verifiably, not simply which is faster to execute.
Comparing the two options
Feature flag toggle: fast (a config change, often seconds to propagate) and narrowly scoped if the flag genuinely gates all of the new behavior; the risk is that the flag may not cover every code path touched by the change (a partial flag, or a change that also altered a shared library or schema outside the flag's reach), in which case toggling it off looks like it worked but leaves some of the regression in place.
Rollback of the latest deploy: restores a fully known-good binary, so it does not depend on the flag actually covering all of the change; the cost is it also reverts any other, unrelated changes bundled in the same deploy, and a rollback typically takes longer to execute and verify than flipping a flag.
Assessing risk
The deciding question is confidence that the flag fully covers the regression: if the change was built and tested specifically to be flag-gated (and QA verified the flag-off path matches the prior release), toggle the flag first since it is faster and reversible with less collateral. If there is any doubt the flag is complete (schema changes, shared state, a partial rollout of the flag itself), prefer the rollback, since a flag toggle that only partially fixes the incident quietly wastes time while customers keep seeing errors.
Preserving data and logs
Before acting, capture the current error logs, traces, and the exact deploy/flag state (which version is live, which flag values are set, for which cohorts), since a rollback in particular can make the failing state harder to inspect afterward; snapshot dashboards and pull a sample of failing requests first if it costs only seconds.
Performing the action and verifying restoration
Toggle the flag (or execute the rollback) for a narrow slice first if that is safely possible (one region, one host group) to confirm the fix actually restores correct behavior before flipping it globally; then verify against the original symptom directly (the specific error rate, endpoint, or user cohort that was affected) rather than a generic drop in overall errors, since a coincidental dip can be mistaken for a fix.
Communication and avoiding further disruption
Announce which action was taken and its known side effects before taking it where possible (for example: "we are toggling flag X off, which also reverts behavior Y that shipped with it") so stakeholders are not surprised, and avoid stacking a second untested change on top while the first action's effect is still being verified, since two simultaneous changes make it impossible to attribute the outcome to either one.
Trade-offs and pitfalls
Defaulting to the flag toggle purely because it is faster, without first confirming it actually covers the full blast radius of the change, is the most common mistake in this scenario; when in doubt about coverage, the rollback's completeness is worth its slower execution time.
Explain how you decide which events to log at each log level (debug/info/warn/error/fatal) and what to instrument with metrics versus traces versus logs. Provide guidelines for structured logging, correlation IDs, sampling decisions, and approaches to avoid log spam while retaining useful diagnostic signal.
Sample Answer
Logs, metrics, and traces answer different questions, and picking the wrong one first wastes the minutes that matter most in an investigation.
What each answers
- Metrics: aggregate, cheap-to-store numeric signals over time ("is error rate elevated, since when, how much"). Best for detecting that something is wrong and roughly when.
- Traces: the path and timing of one request across services ("where did this specific request spend its time, which downstream call failed"). Best for localizing where in a distributed call graph a problem lives.
- Logs: detailed, often unstructured records of what a specific component did ("what exact error, what input, what stack trace"). Best for the why once metrics and traces have pointed at a component.
Log levels and instrumentation guidelines
- debug: verbose, developer-facing detail, off by default in production.
- info: normal operational milestones (request received, job completed).
- warn: something recoverable happened that a human should notice trends in.
- error: an operation failed and needs attention.
- fatal: the process cannot continue.
Attach a correlation ID to every log line and trace span for a given request so the three signal types can be joined together later, and sample verbose logs (e.g., 1% of requests, or 100% only when an anomaly is already flagged) rather than logging everything at full volume, which both costs money and buries the useful signal.
Worked triage order
For a latency spike: check the metric first (confirm it's real and scope it: one endpoint, one region, all traffic), then the trace for a slow request in that scope (find which span is slow), then the logs for that specific span (find the exact error or slow query). Going log-first on a fleet-wide problem means grepping millions of lines before you even know what you're looking for.
Trade-offs and pitfalls
Over-logging at debug level in production is a common self-inflicted problem: it raises cost, and paradoxically makes finding the one relevant line harder. The fix is dynamic log levels (raise verbosity temporarily and narrowly, not permanently and broadly) plus structured fields so a query can filter precisely instead of relying on human eyeballing.
A service intermittently times out trying to reach a dependency that lives in a different subnet. How would you use VPC Flow Logs to figure out whether it's routing, security groups, or something else?
Sample Answer
Direct answer
Pull Flow Log records for the source and destination ENIs (Elastic Network Interfaces, the virtual network cards attached to each instance) and read the action field. A REJECT for that exact tuple means a security group or NACL (Network Access Control List, a stateless, subnet-level firewall, separate from the per-instance security group) is blocking it, while ACCEPT records with the app still timing out mean the problem is above the network layer entirely.
Structured elaboration
- Query Flow Logs (Athena or CloudWatch Insights) filtered to the incident window and the ENIs/ports involved.
- On
REJECT, check both the security group and the NACL, since NACLs are stateless and can block the return leg even when the security group allows the request. - On
ACCEPTwith no timely response, look at DNS resolution, the TLS handshake, or the destination process itself, none of which Flow Logs show. - No records at all suggests routing, a missing route table entry or peering/Transit Gateway (a managed hub that routes traffic between multiple VPCs and on-premises networks over VPN or dedicated connections) misconfiguration, rather than a security rule.
Worked example
action=REJECT for 10.0.1.15:443 -> 10.0.2.20:5432 conclusively points at SG/NACL rules; action=ACCEPT for the same tuple with a client-side timeout redirects the investigation entirely toward the destination service instead.
Trade-offs and pitfalls
Flow Logs sample and aggregate rather than log every packet, so very brief issues can be underrepresented. They carry no payload detail, so ACCEPT doesn't mean the request was handled correctly.
What the interviewer probes next
Why NACLs being stateless matters for return traffic, and how you'd alert on a REJECT spike for a given path.
Unlock Full Question Bank
Get access to all 22 Systematic Debugging and Root Cause Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.