Debugging and Systematic Troubleshooting Questions
Diagnosing defects methodically: reproducing failures, forming and testing hypotheses, reading stack traces and logs, bisecting changes, and reasoning about error handling and edge cases. Covers a disciplined root-cause approach that applies from local bugs to production issues, distinct from embedded hardware-level debugging. A universally probed engineering-craft skill.
Given these simplified application logs:
2025-05-01T10:12:01Z [INFO] Start request id=abc123 user=42
2025-05-01T10:12:01Z [ERROR] Database timeout after 30s query='SELECT * FROM orders WHERE id=?' id=42
2025-05-01T10:12:02Z [INFO] End request id=abc123 status=504
What is the most likely cause of the error, and what is your first investigative step? Explain what you would check in the database and the application to reproduce and isolate the timeout.
Sample Answer
Direct answer
The database timeout for id=42 is the direct, proximate cause: the query SELECT * FROM orders WHERE id=? took longer than 30 seconds and was aborted, which propagated up as a 504 (gateway timeout) to the caller. The first investigative step is checking whether that timeout is isolated to this one request/row (a lock, a corrupted index, a full table scan on a specific value) or part of a broader pattern (all queries against this table are slow right now, pointing at load, a missing index, or a database-wide issue).
Structured elaboration
- Read the log sequence in order. The request starts (
Start request), 30 seconds later the database call itself reports a timeout on a specific query with a specific parameter, and the request ends with a 504. This tells you the 504 isn't a generic gateway problem; it has a specific, identified cause one layer down: the query never completed. - Check whether this is isolated or systemic. Query the database's own slow-query log or active-query view for the same time window: are OTHER queries against the
orderstable also running long, or is this specific toid=42? If it's isolated to specific IDs, suspect row-level lock contention (another transaction holding a lock on that row) or a data anomaly specific to that row. If it's systemic, suspect load, a missing/degraded index, or a resource constraint on the database host. - Check the query plan. Run
EXPLAINon the exact query to confirm it's using an index onidas expected; a missing or invalidated index (from a recent migration, a statistics update, or a query planner regression) would turn what should be an instant primary-key lookup into a full table scan, and that would degrade specifically under load or as the table grows, matching an intermittent-timeout symptom. - Check for lock contention specifically. Query the database's lock/transaction views for anything holding a lock on the
ordersrow withid=42at that timestamp; a long-running transaction elsewhere in the system (a batch job, a stuck transaction that never committed) can block an otherwise-instant query indefinitely. - Reproduce and isolate. Once you have a hypothesis (say, a missing index), reproduce the slow query directly against a copy of the database outside the application, confirm the fix (adding the index, or fixing the transaction that was holding a lock) actually resolves the timing, before deploying it.
- Check the application side too, not just the database. Pull the application's own metrics for that window: was there a concurrent spike in request volume or a burst of retries hitting this same endpoint that could itself be driving DB load, independent of any query-plan or locking issue? Check the application's connection-pool usage and configured statement timeout: if the pool was saturated, requests could queue on the application side before ever reaching the database, which looks identical to a slow query from the caller's point of view. Also check whether the application (or a client library's default) is retrying failed calls in a way that amplifies load on the exact query that's already struggling.
Worked example
Checking the database's active-queries view for that exact timestamp shows several OTHER queries against orders also running unusually long, all doing primary-key lookups that should be near-instant. EXPLAIN on the query confirms it's doing a full table scan, not an index seek. Checking recent schema changes shows an index on orders.id was dropped as part of an unrelated migration two days earlier, under the mistaken assumption it was redundant with the primary key, when in fact the table used a different clustering key (the column the database physically sorts and stores rows by on disk; looking a row up by a different column can't reuse that fast physical ordering, so the database falls back to a full table scan). Re-adding the index resolves the timeouts across the board, confirming it as a systemic (not row-specific) cause.
Trade-offs and pitfalls
The trap is fixating on the SPECIFIC request (id=42) and trying to reproduce a timeout on exactly that row, when the actual cause is systemic and would show up on any query hitting the same missing index; checking for a broader pattern FIRST, before deep-diving one instance, avoids wasting time chasing a red herring that happens to be specific to the one log line you were handed.
What signals in your logs and metrics would push you to roll back a deployment versus continue debugging the currently deployed version of a service? Give a short rubric you use in production, plus a brief example of applying it.
Sample Answer
Direct answer
Roll back when the signals show the blast radius is growing or user-facing impact is severe and the fix isn't yet proven; keep debugging in place when the signals are stable or narrowing and you have a low-risk path to more evidence. The rubric is about trend and confidence, not just the current error rate: a flat 2% error rate you understand is safer to sit with briefly than a 0.5% rate that's climbing and whose cause is still unknown.
Structured elaboration
Concrete signals that push toward rollback:
- Error rate or latency is still climbing, not plateaued, meaning the blast radius is actively growing while you investigate.
- The failure touches a critical path (payments, auth, data writes) where continued exposure risks data integrity or revenue, not just degraded UX.
- You cannot yet state a specific hypothesis for the cause. If you don't know what's wrong, you can't bound how bad it might get, and every minute of uncertainty is a minute of unbounded risk.
- The most recent deploy is a plausible cause and rolling it back is cheap (a well-tested, low-risk rollback path exists). When rollback is cheap and the cost of being wrong is high, the asymmetry favors rolling back even before the cause is confirmed.
Signals that support continuing to debug in place:
- The metric is flat or already recovering on its own (e.g., a transient dependency blip that's clearing), meaning the risk is bounded and shrinking.
- You have a specific, testable hypothesis and the evidence to confirm or reject it is minutes away, not hours.
- The failure is isolated to a non-critical path or a small, known user segment, so continued exposure has a low, well-understood ceiling.
- Rolling back has its own real cost (loses a needed migration, reintroduces a different known bug, or the rollback path itself is unproven), making rollback riskier than the current state.
Two more factors that belong in the rubric alongside the raw signals: your technical confidence in a proposed quick patch, and your monitoring capability while you wait. A patch you are highly confident in, that you can deploy and verify within minutes, tips the balance toward fixing forward even at a moderate error rate; a patch you're guessing at does not, no matter how appealing "just ship the fix" feels under pressure. Separately, if your monitoring can only tell you the aggregate error rate but not WHO is affected or WHY, you are debugging half-blind, and that itself is a reason to prefer the safer, well-understood state (rollback) over continuing to poke at a system you can't fully observe.
Worked example
A rubric applied in production: after a deploy, error rate on the checkout endpoint rises from 0.1% to 1.5% and holds flat for four minutes while you check logs. The rate is not climbing further, the failures are concentrated in one specific edge case (carts with a discount code applied), and a log line points directly at a null-handling bug in the new discount logic. Decision: continue debugging, because the blast radius is bounded, isolated, and you have a specific hypothesis you can confirm in the next two minutes. Contrast: if that same 1.5% had climbed to 6% over those four minutes with no clear pattern in the failing requests, the correct call flips to immediate rollback, because the trend is the dominant signal, not the absolute number.
Trade-offs and pitfalls
The rubric fails when applied to a single snapshot instead of a trend: a rate that looks acceptable in isolation can be five minutes from becoming a major incident, or a rate that looks alarming can already be resolving itself. The discipline is to always ask "is this getting better, worse, or staying the same" before deciding, and to treat "I don't have a hypothesis yet" as itself a strong vote toward rollback, since it means you cannot bound the risk of waiting.
Tell the story of a concrete bug or production failure you found. Explain how you detected it, how you reproduced it if that was possible, the debugging tools and techniques you used, the root cause, and the permanent fix you implemented.
Sample Answer
Direct answer
A concrete story: a service occasionally returned stale pricing data to a subset of users, detected via a customer complaint rather than any internal alert (since the values were plausible-looking, just wrong, not obviously broken); the root cause traced to a caching layer that keyed its cache entries incorrectly, causing two logically-distinct pricing contexts to collide and overwrite each other's cached value, and the permanent fix corrected the cache key's uniqueness rather than just adjusting the cache's expiry time.
Structured elaboration
How it was detected: a customer support ticket reported seeing a price that didn't match what should have applied to their account tier, with no corresponding error or alert on the engineering side, since the returned value was a real, validly-formatted price, just the WRONG one; this is a useful detail because it illustrates a class of bug (returning plausible-but-wrong data) that's structurally invisible to error-rate-based monitoring, and only surfaces via a downstream consumer noticing a substantive discrepancy.
How it was reproduced: confirming the report wasn't a one-off required identifying the PATTERN, not just the single instance; checking whether other users on the same account tier around the same time window also received an unexpected price showed a small but real cluster, ruling out "one weird one-off" and confirming a systemic, reproducible mechanism worth a full investigation.
Debugging tools and techniques used: traced the pricing-lookup code path for the affected requests, and found it flows through an in-memory cache keyed, it turned out, on account tier ALONE rather than on the combination of account tier AND region (pricing legitimately varies by both); when two users on the same tier but different regions made requests close together in time, the second request's result could overwrite the first's cache entry under the shared, insufficiently-specific key, and a THIRD user (same tier, either region) arriving shortly after could then receive whichever region's price happened to be cached most recently, regardless of their own actual region.
The root cause: a cache key that didn't include every dimension the underlying value actually varied by, a classic caching-correctness bug: the cache was implicitly promising "this value is valid for anyone with this tier," when the real invariant needed was "this value is valid for anyone with this tier AND this region."
The permanent fix implemented: updated the cache key to include region alongside tier, restoring the correct invariant; also added a specific integration test that exercises exactly this scenario (two regions, same tier, interleaved requests) to catch a regression of this specific mechanism in the future, since the original bug had shipped without any test covering this particular combination of dimensions.
What you learned that helps you avoid similar bugs: whenever introducing a cache, explicitly enumerate every dimension the cached value can legitimately vary by, and verify the cache key includes ALL of them, not just the ones that happen to be obvious or top-of-mind at implementation time; a caching bug of this shape is especially dangerous specifically because it fails SILENTLY (no error, no crash, just occasionally-wrong data) and is invisible to typical error-rate monitoring, which argues for treating "does this cache key capture every dimension of variation" as a deliberate design-review question on any future caching work, not something to verify only after a bug report arrives.
Trade-offs and pitfalls
The tempting quick fix, once the symptom (stale/wrong cached price) was understood, would have been to simply reduce the cache's TTL (expiry time), which would have reduced the WINDOW during which a collision could produce visibly wrong data without fixing the actual collision mechanism at all; the permanent fix specifically addressed the cache KEY's correctness, not the expiry duration, since a shorter TTL would have masked the bug's visible frequency without removing its root cause.
A production API sometimes returns elements in an inconsistent order across clients because sets are used internally. You are responsible for triage: how do you investigate, explain the nondeterminism to stakeholders, and implement a stable ordering for the API output while keeping acceptable performance?
Sample Answer
Direct answer
Sets in most languages make no guarantee about iteration order, so building API output directly from set iteration produces order that can legitimately differ across processes, language/runtime versions, or even between runs of the SAME process, depending on internal hash-table implementation details; the investigation confirms this mechanism directly, then implements a stable, explicit ordering rather than relying on incidental set-iteration behavior.
Structured elaboration
How to investigate: confirm the specific code path building the API response iterates over a set (or a dict/map in a language where iteration order isn't guaranteed) rather than a list or an explicitly sorted structure; reproduce by calling the endpoint multiple times, or across multiple server instances/processes, and diffing the returned element order directly, which should show inconsistency if a set is indeed the cause, versus consistent-but-simply-unexpected order from some other source (like a database query with no explicit ORDER BY, a related but distinct cause worth ruling out with the same investigative approach).
Explaining the nondeterminism to stakeholders: frame it precisely: this isn't a random or buggy failure, it's the EXPECTED behavior of an unordered collection, and the API was implicitly promising an ordering guarantee it was never actually designed to provide; different clients (or the same client at different times) can legitimately see different orderings today, which may have gone unnoticed as long as most callers didn't depend on order, until a specific consumer's logic (or a stricter test) started depending on stability that was never actually guaranteed.
Implementing stable ordering while maintaining acceptable performance:
- Sort explicitly at the point of serialization, using whatever ordering makes sense for the API's actual semantics (alphabetical, insertion order if that's meaningful and trackable, or a natural key like an ID or timestamp); for most APIs, the cost of sorting a response-sized collection (typically not enormous) is negligible compared to the request's other costs (network, serialization itself).
- If insertion order specifically needs to be preserved (and the language's default set doesn't track it), switch to an ordered-set-like structure if the language provides one (some languages/standard libraries offer collections that combine set semantics with insertion-order iteration), avoiding a separate sort step while still gaining determinism.
- For very large collections where sorting cost genuinely matters, consider whether the ordering can be established earlier in the pipeline (e.g., if the data already comes from a sorted source like a database query with an explicit
ORDER BY, preserving that order through to the response rather than passing it through an unordered set at any intermediate step) rather than re-sorting a large collection at serialization time on every request.
Worked example
Confirming the mechanism: the endpoint's handler collects results into a set (used originally just to deduplicate, with no awareness that its iteration order would become externally visible), then serializes that set directly to the response. Diffing repeated calls to the same endpoint shows genuinely different orderings across calls, confirming set-iteration nondeterminism as the mechanism (as opposed to, for example, a database query lacking an explicit sort, which would tend to be consistent WITHIN one server/database session but could still differ across sessions or after a schema change, a related but mechanistically distinct possibility worth ruling out explicitly rather than assuming). Fix: after deduplicating via the set (keeping that step, since dedup itself is still correct and desired), explicitly convert to a list and sort it by a natural, stable key (the item's own ID) before serializing, at negligible added cost relative to the rest of the request, resolving the nondeterminism while preserving the original deduplication behavior.
Trade-offs and pitfalls
The temptation to "fix" this by simply switching the internal data structure to something that happens to iterate in insertion order today, without an EXPLICIT sort, risks re-introducing the same class of bug if the underlying collection or its implementation ever changes in a future language/runtime version; an explicit, intentional sort at the serialization boundary is more robust than relying on an implementation detail of whatever collection happens to be used internally, even if that detail is currently observed to be stable.
Find and fix the bug in this JavaScript async function, where a missing await causes unpredictable ordering and errors:
async function processItems(items, processor) {
items.forEach(async item => {
await processor(item);
});
console.log('done');
}
Explain the root cause and provide a corrected version that guarantees 'done' prints only after every item has been processed.
Sample Answer
Direct answer
Array.prototype.forEach does not await its callback: it fires each async callback and moves on immediately, so console.log('done') runs before any of the await processor(item) calls have resolved. The fix is to replace forEach with a construct that actually waits, either a for...of loop with await inside it (sequential) or Promise.all over items.map(...) (concurrent).
Structured elaboration
forEach was designed before async/await existed and its callback's return value, including a Promise, is simply discarded. Passing an async function as the callback doesn't change this: each invocation still returns a Promise that forEach never looks at, so forEach itself completes synchronously (having merely started every item, not finished any of them) the instant it has called the callback once per array element. The console.log('done') line after the forEach call then runs immediately, while the async work is still in flight in the background.
Worked example
async function sleep(ms) { return new Promise(r => setTimeout(r, ms)); }
// Fixed, distinct per-item delays (not Math.random()) so the interleaving below
// is deterministic and reproduces identically on every run, not just illustrative.
const delays = { 1: 20, 2: 10, 3: 30 };
// BUGGY
async function processItemBuggy(x) { await sleep(delays[x]); console.log(' processed', x, '(buggy)'); }
async function processBuggy(items) {
items.forEach(async item => { await processItemBuggy(item); });
console.log('done (buggy)');
}
// FIXED: sequential
async function processItemFixed(x) { await sleep(delays[x]); console.log(' processed', x, '(fixed)'); }
async function processFixed(items) {
for (const item of items) { await processItemFixed(item); }
console.log('done (fixed)');
}
(async () => {
console.log('--- buggy run ---');
await processBuggy([1, 2, 3]);
await sleep(100); // let the buggy run's background work finish before the next section starts
console.log('--- fixed run (sequential for-of) ---');
await processFixed([1, 2, 3]);
})();
Executed output (buggy items finish asynchronously in the background; with the fixed per-item delays above, item 2 always resolves first, then item 1, then item 3, after done has already printed):
--- buggy run ---
done (buggy)
processed 2 (buggy)
processed 1 (buggy)
processed 3 (buggy)
--- fixed run (sequential for-of) ---
processed 1 (fixed)
processed 2 (fixed)
processed 3 (fixed)
done (fixed)
done (buggy) prints immediately, before any processed line: forEach fired all three async callbacks and moved on without waiting for any of them. The processed (buggy) lines do eventually print, once their fixed delays resolve, in delay order (2, then 1, then 3, not the input order), rather than before done. The for...of version correctly prints all three processed lines, strictly in input order, before done.
Trade-offs and pitfalls
There are two valid fixes with different semantics, and picking the wrong one is itself a common mistake: a for...of loop processes items strictly one at a time (useful when order matters or you must not overwhelm a downstream dependency with concurrent calls), while await Promise.all(items.map(item => processor(item))) processes all items concurrently and finishes as soon as the slowest one does (faster, but only safe if the operations are independent and the downstream system can handle concurrent load). Silently reaching for forEach out of habit, rather than deliberately choosing sequential versus concurrent semantics, is exactly how this bug class gets reintroduced even by developers who already know the rule.
Unlock Full Question Bank
Get access to all 25 Debugging and Systematic Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.