Systematic Debugging and Root Cause Analysis Questions
Methodically diagnosing failures and identifying their true cause. Covers hypothesis-driven debugging, bisection and instrumentation, full-stack and production diagnosis, debugging under pressure, and root-cause analysis that prevents recurrence. Emphasizes a repeatable process over guesswork.
Design a Python helper function that uses git CLI to perform a bisect over a list of recent commits to find the commit that introduced a failing test. Provide the function interface, describe how it runs tests, handles flaky tests, and what assumptions you make about the environment.
Sample Answer
git bisect performs binary search over commit history: you mark one commit known-good and one known-bad, it checks out the midpoint, you (or a script) report good/bad, and it repeats, isolating the first bad commit in O(log n) tests. A Python helper can drive this directly through the git CLI via subprocess, one commit at a time, instead of leaving the whole loop to a shell one-liner.
Function interface
import subprocess
def bisect_find_regression(
good_commit: str,
bad_commit: str,
test_command: list,
repo_path: str = ".",
retries: int = 5,
failure_threshold: int = 1,
) -> str:
"""
Uses `git bisect` (via the git CLI) to find the first commit between
good_commit and bad_commit that introduced a failure of `test_command`.
Returns the SHA of the first bad commit.
Assumptions:
- repo_path is a clean git checkout with no uncommitted changes: bisect
checks out different commits in place, and this function does not
stash or restore working-tree edits.
- test_command is a list suitable for subprocess.run, e.g.
["pytest", "-k", "regression_case"].
- a commit whose build/setup itself fails (not the test under
investigation) must be treated as untestable so bisect skips it
instead of misattributing the regression to an unrelated break;
this function does that by checking for a distinguished exit code.
- failure_threshold / retries handle test flakiness: the test is run
up to `retries` times on a commit and only counted "bad" if it
fails at least `failure_threshold` times, so one flaky failure on
an actually-good commit does not derail the search.
"""
subprocess.run(["git", "bisect", "start"], cwd=repo_path, check=True)
subprocess.run(["git", "bisect", "bad", bad_commit], cwd=repo_path, check=True)
subprocess.run(["git", "bisect", "good", good_commit], cwd=repo_path, check=True)
def _verdict_for_current_checkout() -> str:
failures = 0
for _ in range(retries):
proc = subprocess.run(test_command, cwd=repo_path)
if proc.returncode == 125:
return "skip"
if proc.returncode != 0:
failures += 1
return "bad" if failures >= failure_threshold else "good"
try:
while True:
verdict = _verdict_for_current_checkout()
out = subprocess.run(
["git", "bisect", verdict],
cwd=repo_path,
capture_output=True,
text=True,
check=True,
)
if "is the first bad commit" in out.stdout:
return out.stdout.splitlines()[0].split()[0]
finally:
subprocess.run(["git", "bisect", "reset"], cwd=repo_path, check=True)
How it runs tests
Each iteration runs test_command against whatever commit git bisect has currently checked out, using its process exit code as the verdict: 0 means good, a nonzero code (other than 125) means bad, and 125 means "cannot be tested, skip this commit" (the convention git bisect run itself uses, so a wrapper script that already follows it plugs straight into this function's _verdict_for_current_checkout). This was verified end to end against a small real git repository: five commits, a passing test on the first two, a bug introduced on commit four (test starts failing) with an unrelated no-op commit after it, and bisect_find_regression correctly returned the exact commit that introduced the failure.
Handling flaky tests
A test that is flaky will make bisect converge on the wrong commit, because a single failing run at a good commit looks identical to a real regression. _verdict_for_current_checkout runs the test up to retries times per commit and only reports "bad" once at least failure_threshold of those runs failed, so an occasional flake on a good commit does not get reported as bad, while a commit that fails consistently still gets reported bad quickly. Tune retries/failure_threshold to the test's known flake rate: a test that fails 1 in 20 runs when truly "good" needs enough retries that a false-bad verdict is unlikely (e.g. requiring 2+ failures out of 5 runs), while a very stable test can use retries=1.
Assumptions about the environment
- The repo has a linear, bisectable history between
good_commitandbad_commit(a single monotonic transition, not multiple unrelated interleaved regressions in the range). test_commandis runnable from a fresh checkout with no manual setup steps the script doesn't perform (dependencies are installed as part oftest_commanditself, or already present).- The caller has already established that
good_committruly passes andbad_committruly fails before starting; the function does not re-verify the endpoints.
Trade-offs and pitfalls
Bisect assumes a single monotonic transition from good to bad; if the bug is intermittent even in its "bad" state, or if multiple unrelated changes landed in the range, the repeat-N-times guard above is what keeps the search from being derailed, and batched-commit CI setups should bisect at the smallest unit of change actually available, not the batch. Because git bisect reset runs in a finally block, the repo is always left back on its original branch even if the loop is interrupted or a commit is genuinely untestable end to end.
Explain how you decide which events to log at each log level (debug/info/warn/error/fatal) and what to instrument with metrics versus traces versus logs. Provide guidelines for structured logging, correlation IDs, sampling decisions, and approaches to avoid log spam while retaining useful diagnostic signal.
Sample Answer
Logs, metrics, and traces answer different questions, and picking the wrong one first wastes the minutes that matter most in an investigation.
What each answers
- Metrics: aggregate, cheap-to-store numeric signals over time ("is error rate elevated, since when, how much"). Best for detecting that something is wrong and roughly when.
- Traces: the path and timing of one request across services ("where did this specific request spend its time, which downstream call failed"). Best for localizing where in a distributed call graph a problem lives.
- Logs: detailed, often unstructured records of what a specific component did ("what exact error, what input, what stack trace"). Best for the why once metrics and traces have pointed at a component.
Log levels and instrumentation guidelines
- debug: verbose, developer-facing detail, off by default in production.
- info: normal operational milestones (request received, job completed).
- warn: something recoverable happened that a human should notice trends in.
- error: an operation failed and needs attention.
- fatal: the process cannot continue.
Attach a correlation ID to every log line and trace span for a given request so the three signal types can be joined together later, and sample verbose logs (e.g., 1% of requests, or 100% only when an anomaly is already flagged) rather than logging everything at full volume, which both costs money and buries the useful signal.
Worked triage order
For a latency spike: check the metric first (confirm it's real and scope it: one endpoint, one region, all traffic), then the trace for a slow request in that scope (find which span is slow), then the logs for that specific span (find the exact error or slow query). Going log-first on a fleet-wide problem means grepping millions of lines before you even know what you're looking for.
Trade-offs and pitfalls
Over-logging at debug level in production is a common self-inflicted problem: it raises cost, and paradoxically makes finding the one relevant line harder. The fix is dynamic log levels (raise verbosity temporarily and narrowly, not permanently and broadly) plus structured fields so a query can filter precisely instead of relying on human eyeballing.
A release introduces errors for a small subset of users. Describe how you'd use canary deployments and feature flags together to detect and bisect the failing change across multiple services, minimizing blast radius. Explain metrics to look at, rollback strategy, and how to automate or manually drive the bisect process.
Sample Answer
Canaries and feature flags combine to localize a regression across services without exposing all users at once: a flag lets you turn a change on for a small, targeted slice while a canary population lets you observe its effect before promoting it further.
Bisecting across services with this toolset
- Ship the suspected change behind a flag, defaulted off everywhere.
- Enable it for a small canary slice (a percentage of traffic, or one region/tenant) in one service at a time, holding the others at their previous behavior.
- Compare the canary's key metrics (error rate, latency, business metric) against a control group receiving identical traffic with the flag off.
- If the canary shows the regression, that service is implicated; if not, flip the flag off there and enable it in the next candidate service, repeating until the regression appears, which is the same halving-the-search-space idea as a code bisect, applied to a set of services instead of a commit range.
Metrics, rollback, and automation
Watch both the target metric and 2-3 guardrail metrics (so a fix in one metric isn't masking regressions in another), and set an automatic rollback trigger (e.g. error-rate delta beyond a threshold sustained for N minutes) rather than relying on someone watching a dashboard. Automating the bisect-across-services loop requires each service's flag state and its canary metrics to be queryable programmatically so a controller can walk the search space itself.
Trade-offs and pitfalls
Canary-based isolation is slower than a code bisect (each step needs a real observation window, typically minutes, not seconds) and assumes the regression is visible in the canary's traffic mix; a regression that only appears for a rare user segment may need a canary sized or targeted specifically at that segment rather than a generic percentage rollout.
Write a short plan (steps and checks) to reproducibly investigate an intermittent production bug where inference results for a particular user occasionally differ from expected despite using the same model version.
Sample Answer
Inference results differing for the same user on the same model version usually trace to a source of nondeterminism the system didn't intend: sampling, non-deterministic GPU ops, model warm-up state, or input preprocessing that isn't as stable as assumed.
Investigation plan
- Rule out sampling first: if the model uses stochastic decoding (temperature/top-k sampling) rather than greedy/deterministic decoding, differing outputs are expected, not a bug; confirm the decoding config actually used at inference time.
- Check for non-deterministic ops: certain GPU kernels (some cuDNN (NVIDIA's GPU library that implements the low-level convolution math deep-learning frameworks call into) convolution algorithms, atomic-based reductions) are not bit-exact deterministic by default; verify whether determinism flags were set (e.g. forcing deterministic algorithms) and whether that setting differs across the serving fleet.
- Model warm-up state: confirm whether the model or its runtime has any stateful warm-up (cached graph optimizations, JIT compilation choices) that could differ between a cold and warm instance serving the same request.
- Input preprocessing: verify the exact preprocessing pipeline is byte-for-byte identical for the same logical input across requests (floating-point order-of-operation differences, locale-dependent formatting, or non-deterministic tokenization can all silently vary the input tensor).
Verified: seeding gaps produce exactly this class of bug
The same root cause class shows up, and can be directly reproduced, in the RNG layer around a model rather than the model itself. Two concrete, executed examples:
1) Cross-library seeding gap (seeding random but forgetting numpy, or vice versa, so only part of the pipeline is actually reproducible):
import random
import numpy as np
def naive_seed_everything(seed):
random.seed(seed) # BUG: forgets numpy (or any other RNG source)
def full_seed_everything(seed):
random.seed(seed)
np.random.seed(seed) # FIX: seed every RNG source the pipeline touches
# naive: random.random() is reproducible run-to-run, np.random.rand() is NOT,
# because something else in the process advances numpy's global RNG state
# between runs. Executed result:
# naive : random() reproducible=True numpy reproducible=False
# fixed : random() reproducible=True numpy reproducible=True
2) Worker-seeding bug in a multi-worker data loader (seeding every worker with the same fixed value at init makes every worker's "shuffle" produce the identical, not independently random, sequence):
import random
def buggy_worker_shuffle(worker_id, base_seed, dataset):
rng = random.Random(base_seed) # BUG: ignores worker_id
order = dataset[:]
rng.shuffle(order)
return order
def fixed_worker_shuffle(worker_id, base_seed, dataset, epoch=0):
rng = random.Random(base_seed + worker_id * 1000 + epoch) # FIX: per-worker, per-epoch
order = dataset[:]
rng.shuffle(order)
return order
# Executed with 4 workers, dataset of 20 items, base_seed=7:
# buggy_worker_shuffle -> all 4 workers produce the IDENTICAL order
# fixed_worker_shuffle -> all 4 workers produce DISTINCT orders, and the
# same base_seed across two separate runs reproduces
# the identical set of 4 orders (still deterministic,
# just no longer duplicated across workers)
Both snippets were executed directly (Python 3, stdlib random + numpy); the "buggy" variants reproduce exactly the failure mode described (unintended duplication / partial irreproducibility), and the "fixed" variants confirm the corrected behavior (independent-but-reproducible per-worker sequences; fully reproducible cross-library output).
Trade-offs and pitfalls
Forcing full determinism (deterministic GPU algorithms, single-threaded data loading) is a useful diagnostic mode but often costs real throughput; it's typically enabled to investigate, then a narrower fix (proper independent-but-reproducible seeding, confirming decoding config) is shipped rather than leaving full determinism mode on in production permanently.
A timing-related race condition affects a distributed lock acquisition algorithm. Design a test harness that deterministically reproduces the race using process scheduling control or record-and-replay techniques. Describe tools and OS facilities you'd use and how to assert the race occurred.
Sample Answer
Deterministically reproducing a timing-sensitive distributed-lock race means controlling scheduling explicitly rather than hoping the natural interleaving happens, and then asserting the race actually occurred rather than only observing a bad final state.
Executable harness (Python, forced interleaving via events)
import threading
class NaiveDistributedLock:
"""A racy lock: check-then-set against a shared store is NOT atomic,
simulating e.g. two separate GET then SET calls against a remote store
instead of a real atomic SET-if-not-exists."""
def __init__(self, store):
self.store = store
def acquire(self, owner, key="lock", pause_after_check=None):
holder = self.store.get(key) # CHECK
if pause_after_check is not None:
pause_after_check() # deterministic injection point
if holder is None: # ACT (not atomic with CHECK)
self.store[key] = owner
return True
return False
def run_forced_interleaving():
store = {}
lock = NaiveDistributedLock(store)
results = {}
a_checked = threading.Event()
b_done = threading.Event()
def thread_a():
def pause():
a_checked.set()
b_done.wait(timeout=2)
results["A"] = lock.acquire("A", pause_after_check=pause)
def thread_b():
a_checked.wait(timeout=2) # only act once A has checked: forces the window
results["B"] = lock.acquire("B")
b_done.set()
ta, tb = threading.Thread(target=thread_a), threading.Thread(target=thread_b)
ta.start(); tb.start(); ta.join(); tb.join()
return results, store
results, store = run_forced_interleaving()
assert results["A"] and results["B"], "expected both to acquire under the forced interleaving"
print("RACE CONFIRMED:", results, "final holder:", store["lock"])
Running this reliably prints RACE CONFIRMED on every trial, since threading.Event pins the exact order (A checks, A pauses, B checks-and-sets, A resumes and also sets) instead of leaving it to the scheduler.
Asserting the race, not just the symptom
The assertion above proves the mechanism (two participants both passed the check before either completed the act) rather than only checking a downstream symptom like a corrupted counter, which could have other explanations.
Confirming the fix
def run_fixed_with_mutex():
store = {}
guard = threading.Lock()
results = {}
a_checked = threading.Event()
def acquire_atomic(owner, key="lock"):
with guard:
if store.get(key) is None:
store[key] = owner
return True
return False
def thread_a():
results["A"] = acquire_atomic("A")
a_checked.set()
def thread_b():
a_checked.wait(timeout=2)
results["B"] = acquire_atomic("B")
ta, tb = threading.Thread(target=thread_a), threading.Thread(target=thread_b)
ta.start(); tb.start(); ta.join(); tb.join()
return results, store
results, store = run_fixed_with_mutex()
assert not (results["A"] and results["B"]), "fixed version must never let both acquire"
print("Fixed: only one acquired ->", results)
Wrapping the check-and-act in a single critical section (threading.Lock) removes the window entirely; run against the exact same forced ordering, only one of A or B ever acquires.
Tools and OS facilities for the real (non-toy) system
For an actual distributed lock service, the same idea applies with heavier tools: a debugger breakpoint (gdb/lldb) to pause one process right after its check step and before its act step, deliberate SIGSTOP/SIGCONT on a process to freeze it mid-operation, or a record-and-replay tool (rr) to capture a real production interleaving once and replay it deterministically afterward. Each logs a monotonic sequence/version number at every critical step (check, act, release) from every participant so the assertion can be made on the logged order, exactly as done above with the in-process events.
Trade-offs and pitfalls
Forcing an interleaving with events/breakpoints proves the mechanism exists and gives a permanent regression test, but the exact synchronization primitives used to force the race are test-only scaffolding; they prove causality, not that production will hit this interleaving at any particular rate, so the test's value is as a regression guard, not a production-frequency estimate.
Unlock Full Question Bank
Get access to all Systematic Debugging and Root Cause Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.