Debugging and Systematic Troubleshooting Questions
Diagnosing defects methodically: reproducing failures, forming and testing hypotheses, reading stack traces and logs, bisecting changes, and reasoning about error handling and edge cases. Covers a disciplined root-cause approach that applies from local bugs to production issues, distinct from embedded hardware-level debugging. A universally probed engineering-craft skill.
Find and fix the bug in this JavaScript async function, where a missing await causes unpredictable ordering and errors:
async function processItems(items, processor) {
items.forEach(async item => {
await processor(item);
});
console.log('done');
}
Explain the root cause and provide a corrected version that guarantees 'done' prints only after every item has been processed.
Sample Answer
Direct answer
Array.prototype.forEach does not await its callback: it fires each async callback and moves on immediately, so console.log('done') runs before any of the await processor(item) calls have resolved. The fix is to replace forEach with a construct that actually waits, either a for...of loop with await inside it (sequential) or Promise.all over items.map(...) (concurrent).
Structured elaboration
forEach was designed before async/await existed and its callback's return value, including a Promise, is simply discarded. Passing an async function as the callback doesn't change this: each invocation still returns a Promise that forEach never looks at, so forEach itself completes synchronously (having merely started every item, not finished any of them) the instant it has called the callback once per array element. The console.log('done') line after the forEach call then runs immediately, while the async work is still in flight in the background.
Worked example
async function sleep(ms) { return new Promise(r => setTimeout(r, ms)); }
// Fixed, distinct per-item delays (not Math.random()) so the interleaving below
// is deterministic and reproduces identically on every run, not just illustrative.
const delays = { 1: 20, 2: 10, 3: 30 };
// BUGGY
async function processItemBuggy(x) { await sleep(delays[x]); console.log(' processed', x, '(buggy)'); }
async function processBuggy(items) {
items.forEach(async item => { await processItemBuggy(item); });
console.log('done (buggy)');
}
// FIXED: sequential
async function processItemFixed(x) { await sleep(delays[x]); console.log(' processed', x, '(fixed)'); }
async function processFixed(items) {
for (const item of items) { await processItemFixed(item); }
console.log('done (fixed)');
}
(async () => {
console.log('--- buggy run ---');
await processBuggy([1, 2, 3]);
await sleep(100); // let the buggy run's background work finish before the next section starts
console.log('--- fixed run (sequential for-of) ---');
await processFixed([1, 2, 3]);
})();
Executed output (buggy items finish asynchronously in the background; with the fixed per-item delays above, item 2 always resolves first, then item 1, then item 3, after done has already printed):
--- buggy run ---
done (buggy)
processed 2 (buggy)
processed 1 (buggy)
processed 3 (buggy)
--- fixed run (sequential for-of) ---
processed 1 (fixed)
processed 2 (fixed)
processed 3 (fixed)
done (fixed)
done (buggy) prints immediately, before any processed line: forEach fired all three async callbacks and moved on without waiting for any of them. The processed (buggy) lines do eventually print, once their fixed delays resolve, in delay order (2, then 1, then 3, not the input order), rather than before done. The for...of version correctly prints all three processed lines, strictly in input order, before done.
Trade-offs and pitfalls
There are two valid fixes with different semantics, and picking the wrong one is itself a common mistake: a for...of loop processes items strictly one at a time (useful when order matters or you must not overwhelm a downstream dependency with concurrent calls), while await Promise.all(items.map(item => processor(item))) processes all items concurrently and finishes as soon as the slowest one does (faster, but only safe if the operations are independent and the downstream system can handle concurrent load). Silently reaching for forEach out of habit, rather than deliberately choosing sequential versus concurrent semantics, is exactly how this bug class gets reintroduced even by developers who already know the rule.
Find and explain the bug in this Python function, then provide a corrected implementation:
def append_item(item, lst=[]):
lst.append(item)
return lst
# append_item(1) -> [1]
# append_item(2) -> [1, 2] # unexpected: shared across calls
Explain why the sharing happens and show the safe fix for the default argument.
Sample Answer
Direct answer
The bug is that lst=[] creates the default list once, at function-definition time, not once per call. Every call that omits the lst argument reuses the exact same list object, so items accumulate across calls that were never meant to share state. The fix is to use None as the sentinel default and create a fresh list inside the function body when no list was passed.
Structured elaboration
In Python, default argument values are evaluated exactly once, when the def statement executes, and the resulting object is bound to the function itself. A mutable default (a list, dict, or set) is therefore the SAME object across every call that relies on the default, so mutating it in one call is visible in the next. This is a well-known Python pitfall precisely because the code reads as if a fresh list is created per call, which is the intuitive (and correct) behavior for immutable defaults like None, 0, or "", but not for mutable ones.
Worked example
def append_item_buggy(item, lst=[]):
lst.append(item)
return lst
def append_item_fixed(item, lst=None):
if lst is None:
lst = []
lst.append(item)
return lst
print("BUGGY:")
print(append_item_buggy(1))
print(append_item_buggy(2))
print("FIXED:")
print(append_item_fixed(1))
print(append_item_fixed(2))
Executed output:
BUGGY:
[1]
[1, 2]
FIXED:
[1]
[2]
The buggy version's second call returns [1, 2] instead of the expected [2], because it's still appending to the exact same list object created back when the function was defined. The fixed version creates a brand-new list inside the function body on every call where the caller didn't explicitly pass one, so each call starts clean.
Trade-offs and pitfalls
This bug is especially dangerous because it's silent and cumulative: a function like this can appear to work correctly in isolated tests (each test creates its own scenario and might not notice the shared state) and only manifest as data mysteriously growing or leaking between unrelated calls once the function is used repeatedly in a long-running process, which is exactly the kind of "works in tests, fails in production" gap this class of bug produces. The general rule: never use a mutable object (list, dict, set, or a custom mutable class instance) as a default argument value; use None and construct the mutable object inside the function body instead.
Tell the story of a concrete bug or production failure you found. Explain how you detected it, how you reproduced it if that was possible, the debugging tools and techniques you used, the root cause, and the permanent fix you implemented.
Sample Answer
Direct answer
A concrete story: a service occasionally returned stale pricing data to a subset of users, detected via a customer complaint rather than any internal alert (since the values were plausible-looking, just wrong, not obviously broken); the root cause traced to a caching layer that keyed its cache entries incorrectly, causing two logically-distinct pricing contexts to collide and overwrite each other's cached value, and the permanent fix corrected the cache key's uniqueness rather than just adjusting the cache's expiry time.
Structured elaboration
How it was detected: a customer support ticket reported seeing a price that didn't match what should have applied to their account tier, with no corresponding error or alert on the engineering side, since the returned value was a real, validly-formatted price, just the WRONG one; this is a useful detail because it illustrates a class of bug (returning plausible-but-wrong data) that's structurally invisible to error-rate-based monitoring, and only surfaces via a downstream consumer noticing a substantive discrepancy.
How it was reproduced: confirming the report wasn't a one-off required identifying the PATTERN, not just the single instance; checking whether other users on the same account tier around the same time window also received an unexpected price showed a small but real cluster, ruling out "one weird one-off" and confirming a systemic, reproducible mechanism worth a full investigation.
Debugging tools and techniques used: traced the pricing-lookup code path for the affected requests, and found it flows through an in-memory cache keyed, it turned out, on account tier ALONE rather than on the combination of account tier AND region (pricing legitimately varies by both); when two users on the same tier but different regions made requests close together in time, the second request's result could overwrite the first's cache entry under the shared, insufficiently-specific key, and a THIRD user (same tier, either region) arriving shortly after could then receive whichever region's price happened to be cached most recently, regardless of their own actual region.
The root cause: a cache key that didn't include every dimension the underlying value actually varied by, a classic caching-correctness bug: the cache was implicitly promising "this value is valid for anyone with this tier," when the real invariant needed was "this value is valid for anyone with this tier AND this region."
The permanent fix implemented: updated the cache key to include region alongside tier, restoring the correct invariant; also added a specific integration test that exercises exactly this scenario (two regions, same tier, interleaved requests) to catch a regression of this specific mechanism in the future, since the original bug had shipped without any test covering this particular combination of dimensions.
What you learned that helps you avoid similar bugs: whenever introducing a cache, explicitly enumerate every dimension the cached value can legitimately vary by, and verify the cache key includes ALL of them, not just the ones that happen to be obvious or top-of-mind at implementation time; a caching bug of this shape is especially dangerous specifically because it fails SILENTLY (no error, no crash, just occasionally-wrong data) and is invisible to typical error-rate monitoring, which argues for treating "does this cache key capture every dimension of variation" as a deliberate design-review question on any future caching work, not something to verify only after a bug report arrives.
Trade-offs and pitfalls
The tempting quick fix, once the symptom (stale/wrong cached price) was understood, would have been to simply reduce the cache's TTL (expiry time), which would have reduced the WINDOW during which a collision could produce visibly wrong data without fixing the actual collision mechanism at all; the permanent fix specifically addressed the cache KEY's correctness, not the expiry duration, since a shorter TTL would have masked the bug's visible frequency without removing its root cause.
What techniques do you use to prioritize multiple concurrent bugs or incidents affecting ML systems? Describe a decision rubric considering severity, user impact, reproducibility, rollback cost, and business KPIs, and explain how you would apply it during a busy incident window.
Sample Answer
Direct answer
Prioritize concurrent bugs and incidents using an explicit rubric weighing severity, user impact, reproducibility, rollback cost, and business KPIs together, not any single factor alone, since a high-severity-sounding bug with low actual user impact and an expensive rollback can rank BELOW a "smaller" bug that's actively costing revenue and has a trivial fix available.
Structured elaboration
The rubric, and how each factor is weighed:
- Severity: how bad is the failure mode itself (data loss/corruption ranks above a cosmetic issue, a security-relevant bug ranks above a performance blip), independent of how many users are currently affected.
- User impact: how many users/requests are affected right now, and is that number growing, stable, or shrinking; a severe bug affecting a handful of users may rank below a moderate bug affecting a large fraction of traffic.
- Reproducibility: a reliably reproducible bug can be diagnosed and fixed faster (lower time-to-resolution for the same engineering effort) than an intermittent one, which affects how quickly EACH candidate bug can actually be resolved if picked next, not just how bad it is.
- Rollback cost: if a bug traces to a specific recent change, how cheap and safe is reverting that change right now; a bug with a trivial, low-risk rollback available should often be handled immediately regardless of its rank on other factors, since the fix is nearly free.
- Business KPIs: which specific business metric is being affected (revenue, a contractual SLA, a compliance requirement) and how directly; a bug affecting a metric with hard, immediate business consequences (a broken payment flow) generally outranks one affecting a softer, longer-horizon metric even at similar technical severity.
Applying the rubric during a busy incident window: first, quickly triage EVERY open issue against the rubric (a few minutes per issue, not a deep investigation), to get a relative ranking rather than working issues in the order they arrived; second, look specifically for any issue with BOTH meaningful impact AND a cheap available mitigation (like a trivial rollback), since these should jump the queue regardless of their raw severity ranking, because the cost of addressing them is so low relative to the benefit; third, re-triage periodically as the window continues, since impact and severity can both change (a bug's blast radius growing, or a rollback becoming available partway through investigation of one issue), rather than treating the initial ranking as fixed for the whole incident window.
Worked example
Three concurrent issues during a busy window: (A) a high-severity-sounding data-consistency bug affecting a small, specific edge case (low current user impact, no clear quick fix, moderate rollback risk since the change is tangled with other recent work); (B) a moderate-severity bug causing a checkout-flow error for roughly 5% of transactions (clear, growing user impact, directly hitting a revenue KPI, and traced quickly to a specific recent config change with a trivial, low-risk rollback available); (C) a low-severity cosmetic UI bug reported by a few users (minimal impact, no urgency). Applying the rubric: (B) is prioritized FIRST despite being technically "less severe" than (A) in the abstract, specifically because it combines real, growing, revenue-affecting impact with an almost-free rollback fix, making it both the highest-leverage and fastest issue to resolve. (A) is prioritized second, staffed for a proper investigation given its rollback isn't cheap and its actual mechanism needs to be understood before a safe fix can be applied. (C) is deprioritized entirely for the duration of the busy window, revisited once the window calms down.
Trade-offs and pitfalls
A common mistake is prioritizing purely by SEVERITY LABEL (treating "critical" as automatically first regardless of current impact or fix cost), which can leave a rapidly-growing, revenue-affecting, cheaply-fixable issue waiting behind a technically-severe-but-currently-narrow, expensive-to-fix one; the rubric's explicit multi-factor weighing exists specifically to avoid that trap, and re-triaging periodically (rather than committing to the initial ranking for the whole window) accounts for the fact that these factors genuinely change as an incident window progresses.
Describe the role of instrumentation (logs, metrics, traces) in effective debugging. Give a concise checklist of five things you would verify are in place before handing a service off to operations for production use.
Sample Answer
Direct answer
Before handing a service to operations, verify: (1) every meaningful failure path logs enough context to diagnose it without a redeploy, (2) the four golden signals (latency, traffic, errors, saturation) are exposed as metrics, not just logs, (3) a request can be traced end to end when it crosses more than one service, (4) alert thresholds exist and point to an actual runbook, not just a page with no next step, and (5) someone other than the author has actually looked at the dashboards and confirmed they answer "is this healthy right now" at a glance.
Structured elaboration
Each item on the checklist exists because of a specific failure mode it prevents:
- Actionable error logs. A log line that says "operation failed" with no request ID, no input summary, and no stack trace forces on-call to redeploy with more logging just to understand a 2am page. The bar: could someone who has never read this code diagnose the failure category from the log line alone?
- The four golden signals as metrics, not just logs. Logs answer "what happened in this one case"; metrics answer "is this normal right now." Without dashboards for latency, traffic, error rate, and saturation (CPU/memory/queue depth/connection pool usage), operations has no way to distinguish a healthy blip from a developing outage without grepping logs under pressure.
- Distributed tracing or correlation IDs. The moment a request crosses a service boundary, "check the logs" stops being a single grep and becomes "which of these five services' logs, and how do I know they're the same request?" A propagated request ID is the cheapest fix and the most commonly missing piece.
- Alert thresholds tied to a runbook. An alert that fires with no documented first step trains on-call to snooze it, which is worse than no alert at all: it becomes noise that hides the next real incident.
- A second set of eyes on the dashboards. The author of a service is the worst-positioned person to judge whether their own dashboard is readable to someone unfamiliar with the code, the same blind spot that makes self-review of documentation unreliable in general.
Worked example
A concrete pre-handoff review of a new payment-retry service: logs include payment_id, attempt_number, and the specific failure reason on every retry (satisfies #1); a Grafana panel shows retry rate, success rate, and queue depth (satisfies #2); the payment_id is propagated as a header to the downstream charge service so both services' logs can be joined (satisfies #3); the "retry queue depth > 500" alert links directly to a runbook section titled "Retry queue backing up" with three ranked likely causes (satisfies #4); a teammate who did not write the service opened the dashboard cold and correctly identified within 30 seconds whether the system was healthy (satisfies #5).
Trade-offs and pitfalls
The most common gap isn't missing instrumentation entirely, it's instrumentation that only makes sense to the person who wrote it: log lines with internal variable names instead of business-meaningful fields, dashboards with no annotations explaining what "normal" looks like, or alerts that reference a metric name with no context. The checklist is deliberately about READINESS for someone else to operate the system, not about whether instrumentation exists in principle.
Unlock Full Question Bank
Get access to all 25 Debugging and Systematic Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.