Debugging and Systematic Troubleshooting Questions
Diagnosing defects methodically: reproducing failures, forming and testing hypotheses, reading stack traces and logs, bisecting changes, and reasoning about error handling and edge cases. Covers a disciplined root-cause approach that applies from local bugs to production issues, distinct from embedded hardware-level debugging. A universally probed engineering-craft skill.
Provide a checklist and short plan to benchmark and optimize a SQL query that runs slowly in production, including steps to collect explain plans, indexes, and statistics, and how to test schema or query changes safely in staging without affecting production data.
Sample Answer
Direct answer
Benchmark and optimize a slow production query by first capturing its actual execution plan (not a guess at one), checking whether the indexes and table statistics the planner is relying on are current and appropriate, and then testing any proposed change in a safe environment against a REALISTIC copy of the data before touching production.
Structured elaboration
- Capture the current execution plan (
EXPLAIN ANALYZEor the equivalent for your database engine) against the actual slow query with realistic parameters, not a simplified version of it. The plan tells you exactly where time is going: a full table scan where an index seek was expected, a sort spilling to disk, a suboptimal join order. - Check the indexes actually in play. Confirm an index exists on the columns the query filters and joins on, and separately confirm the planner is actually CHOOSING to use it; a query can have a relevant index that the planner ignores because table statistics are stale, making the planner underestimate how selective the index would be.
- Check table statistics freshness. Most database engines rely on periodically-updated statistics (row counts, value distributions) to choose a plan; statistics that haven't been refreshed since a large data change can lead the planner to a badly wrong plan even with perfect indexes in place. Refreshing statistics (
ANALYZEor equivalent) is often a fast, low-risk first thing to try. - Test schema or query changes in staging against representative data, not a small or synthetic dataset. A query that looks fast against a staging database with 1% of production's row count can behave completely differently once real data volume and skew are involved; use a production-like data volume and distribution wherever possible, ideally an actual (anonymized, if needed) production data copy or snapshot.
- Validate the change doesn't harm other queries. Adding an index speeds up reads that use it but adds write overhead and can change the planner's choices for OTHER queries that touch the same table; check the broader query workload against the table, not just the one query being optimized, before deploying.
- Roll out safely. Prefer additive, reversible changes first (adding an index) over changes that are harder to undo (dropping a column, restructuring a table); if an index build is heavy on a large production table, use the database's online/concurrent index-creation mechanism if available, to avoid locking the table during the build.
Worked example
EXPLAIN ANALYZE on the slow query shows a sequential scan on a 40-million-row table where an index seek was expected, despite an index existing on the filtered column. Checking statistics shows they were last updated three weeks ago, before a large batch import roughly doubled the table's size; the planner's row-count estimate is stale enough that it (wrongly) judges a sequential scan cheaper than using the index. Running ANALYZE to refresh statistics, then re-running EXPLAIN ANALYZE, confirms the planner now correctly chooses the index seek, and query time drops from several seconds to under 50ms, with zero schema changes needed.
Trade-offs and pitfalls
Refreshing statistics or adding an index are both LOW-RISK, reversible first moves and should usually be tried and measured before a more invasive schema change; jumping straight to denormalization or a schema redesign without first confirming the planner even has current, accurate information to work with risks solving a problem that a much cheaper fix would have resolved. Testing exclusively against a small staging dataset is the most common way a "fix" turns out not to actually help once it hits production-scale data.
Write a short script, or describe in Bash commands, how you would extract and correlate log lines that share a request-id across three services' log files for a single one-minute window. Mention scaling considerations for large log volumes.
Sample Answer
Direct answer
Grep each service's log file for the shared request_id, merge the results, and sort by timestamp, since the request_id is what ties otherwise-independent log lines across services into one causal sequence. A short, reusable script wraps this into a single command rather than three manual greps and a mental merge.
Structured elaboration
#!/bin/bash
# correlate.sh <request_id> <log_dir>
set -euo pipefail
request_id="$1"
log_dir="$2"
# Pre-filter with a fast fixed-string grep, then require request_id= to match
# as an EXACT whitespace-delimited token, so a shorter id like "req-1" can never
# match a longer one like "req-10" as a substring.
grep -h -F "request_id=${request_id}" "${log_dir}"/*.log \
| awk -v rid="request_id=${request_id}" '{
for (i = 1; i <= NF; i++) { if ($i == rid) { print; next } }
}' \
| sort
The grep -F pass is a fast pre-filter; the awk pass then requires request_id=<id> to match one whitespace-delimited token EXACTLY rather than as a substring, since a plain grep "request_id=${request_id}" would incorrectly also return rows for request_id=req-10 when searching for req-1 (and could also mismatch on any regex metacharacter that happened to appear in a request_id). sort on ISO-8601 timestamps (which sort correctly as plain strings, since the format is lexicographically ordered) then produces a single chronological view across all three services without needing to know which log came from which service ahead of time.
Worked example
Sample logs for three services, each with an entry for the same request_id:
service-a.log: 2025-06-01T10:00:00Z INFO request_id=req-1 msg="received request"
service-b.log: 2025-06-01T10:00:01Z INFO request_id=req-1 msg="forwarded to backend"
service-c.log: 2025-06-01T10:00:02Z INFO request_id=req-1 msg="processed successfully"
Executed: ./correlate.sh req-1 logs
2025-06-01T10:00:00Z INFO request_id=req-1 msg="received request"
2025-06-01T10:00:01Z INFO request_id=req-1 msg="forwarded to backend"
2025-06-01T10:00:02Z INFO request_id=req-1 msg="processed successfully"
A second request that fails midway (req-2), run the same way, correctly isolates just that request's cross-service trail:
2025-06-01T10:00:05Z INFO request_id=req-2 msg="received request"
2025-06-01T10:00:06Z ERROR request_id=req-2 msg="timeout calling downstream"
confirming the timeline: service A received the request, forwarded it to service B, and service B reported a timeout one second later, exactly the sequence needed to know which hop failed.
Scaling considerations for large log volumes
- Plain
grepacross large files is I/O-bound, not CPU-bound; for logs beyond a few GB per service, preferripgrep(rg, which is substantially faster on large files due to better I/O and regex engine choices) or push log correlation into a proper log-aggregation system (Elasticsearch/OpenSearch, Loki, Splunk) where the request_id becomes an indexed field and lookups are near-instant regardless of total volume. - Log rotation means a single request's trail can span multiple rotated files (
app.log,app.log.1,app.log.2.gz); the script would need to glob across rotated files and handle gzipped ones (zgrepinstead ofgrepfor.gzfiles) to stay correct once rotation is in play. - Clock skew across hosts can make timestamp-sorted output subtly wrong at sub-second granularity; for services on different hosts, prefer a monotonic sequence number or a trace-span ordering over raw wall-clock timestamps when precise ordering matters, and treat wall-clock timestamps as approximate once multiple hosts are involved.
Trade-offs and pitfalls
A one-off grep-and-sort script is the right tool for an ad hoc investigation into a handful of requests; it does not scale to "which requests failed across the fleet in the last hour," which needs the aggregation-system approach instead. Reaching for the heavier tool too early wastes setup time on an investigation that a two-line grep would have resolved in under a minute; reaching for grep too late (on a fleet-wide investigation) wastes far more time waiting on slow, unindexed scans across many hosts.
Tell the story of a concrete bug or production failure you found. Explain how you detected it, how you reproduced it if that was possible, the debugging tools and techniques you used, the root cause, and the permanent fix you implemented.
Sample Answer
Direct answer
A concrete story: a service occasionally returned stale pricing data to a subset of users, detected via a customer complaint rather than any internal alert (since the values were plausible-looking, just wrong, not obviously broken); the root cause traced to a caching layer that keyed its cache entries incorrectly, causing two logically-distinct pricing contexts to collide and overwrite each other's cached value, and the permanent fix corrected the cache key's uniqueness rather than just adjusting the cache's expiry time.
Structured elaboration
How it was detected: a customer support ticket reported seeing a price that didn't match what should have applied to their account tier, with no corresponding error or alert on the engineering side, since the returned value was a real, validly-formatted price, just the WRONG one; this is a useful detail because it illustrates a class of bug (returning plausible-but-wrong data) that's structurally invisible to error-rate-based monitoring, and only surfaces via a downstream consumer noticing a substantive discrepancy.
How it was reproduced: confirming the report wasn't a one-off required identifying the PATTERN, not just the single instance; checking whether other users on the same account tier around the same time window also received an unexpected price showed a small but real cluster, ruling out "one weird one-off" and confirming a systemic, reproducible mechanism worth a full investigation.
Debugging tools and techniques used: traced the pricing-lookup code path for the affected requests, and found it flows through an in-memory cache keyed, it turned out, on account tier ALONE rather than on the combination of account tier AND region (pricing legitimately varies by both); when two users on the same tier but different regions made requests close together in time, the second request's result could overwrite the first's cache entry under the shared, insufficiently-specific key, and a THIRD user (same tier, either region) arriving shortly after could then receive whichever region's price happened to be cached most recently, regardless of their own actual region.
The root cause: a cache key that didn't include every dimension the underlying value actually varied by, a classic caching-correctness bug: the cache was implicitly promising "this value is valid for anyone with this tier," when the real invariant needed was "this value is valid for anyone with this tier AND this region."
The permanent fix implemented: updated the cache key to include region alongside tier, restoring the correct invariant; also added a specific integration test that exercises exactly this scenario (two regions, same tier, interleaved requests) to catch a regression of this specific mechanism in the future, since the original bug had shipped without any test covering this particular combination of dimensions.
What you learned that helps you avoid similar bugs: whenever introducing a cache, explicitly enumerate every dimension the cached value can legitimately vary by, and verify the cache key includes ALL of them, not just the ones that happen to be obvious or top-of-mind at implementation time; a caching bug of this shape is especially dangerous specifically because it fails SILENTLY (no error, no crash, just occasionally-wrong data) and is invisible to typical error-rate monitoring, which argues for treating "does this cache key capture every dimension of variation" as a deliberate design-review question on any future caching work, not something to verify only after a bug report arrives.
Trade-offs and pitfalls
The tempting quick fix, once the symptom (stale/wrong cached price) was understood, would have been to simply reduce the cache's TTL (expiry time), which would have reduced the WINDOW during which a collision could produce visibly wrong data without fixing the actual collision mechanism at all; the permanent fix specifically addressed the cache KEY's correctness, not the expiry duration, since a shorter TTL would have masked the bug's visible frequency without removing its root cause.
Describe a systematic, repeatable approach you use to troubleshoot an unfamiliar technical problem end to end. Cover how you observe the symptom, form and prioritize hypotheses, gather and interpret evidence such as logs, metrics, and traces, isolate the root cause, implement and validate a fix, and decide when to escalate, roll back, or write up a postmortem.
Sample Answer
Direct answer
Troubleshooting an unfamiliar problem is a loop, not a single step: observe the symptom precisely, form a small set of testable hypotheses ranked by likelihood and cost to check, gather evidence that discriminates between them, isolate the true cause, implement and validate a fix, and decide whether the incident needs a rollback, an escalation, or a written postmortem. The loop repeats: each piece of evidence should narrow the hypothesis set, not just confirm what you already believed.
Structured elaboration
- Observe the symptom precisely. Write down exactly what is wrong, in falsifiable terms: not "the API is slow" but "p99 latency (the response time slower than 99% of requests, i.e. how bad the worst cases are, not just the average) on
POST /ordersrose from 80ms to 900ms starting at 14:32 UTC, affecting roughly 3% of requests." Vague symptoms produce vague hypotheses. - Form hypotheses before you start digging. List the plausible causes given what changed recently (deploys, config, traffic pattern, dependency versions) and what the symptom rules out. A hypothesis you cannot state is a hypothesis you cannot test.
- Prioritize by expected information gain divided by cost. A five-minute log grep that could confirm or kill three hypotheses at once beats a one-hour deep profiling session that only speaks to one.
- Gather evidence that discriminates. Logs tell you what happened at a point; metrics tell you the shape of the problem over time; traces tell you where time went inside one request. Pick the instrument that actually distinguishes your live hypotheses, not the one you're most comfortable with.
- Isolate the root cause, not just a correlated symptom. A dropped hypothesis should be dropped because evidence contradicts it, not because you got bored of it.
- Implement and validate the fix against the same evidence that revealed the problem. If you diagnosed via a specific metric, watch that metric recover before declaring victory.
- Decide what happens next. If customer impact is ongoing and the fix is unproven, roll back first and diagnose second. If a similar failure could recur, or the incident had real impact, write it up so the org doesn't relearn the same lesson.
Worked example
A "the checkout page is slow" report, applied through the loop: symptom precisely stated as "median load time is normal, but a subset of loads takes 8-12 seconds, starting after this morning's deploy." Hypotheses: (a) the new deploy added a blocking call, (b) a downstream dependency degraded independently, (c) the slow subset shares a common attribute (e.g., a specific region or a large cart). A single log query grouping slow requests by attribute would discriminate between (c) and the other two in minutes, before touching a profiler. Suppose it shows the slow requests all hit a newly added inventory-check call to a dependency with no timeout: that both confirms (a) and rules out (b)/(c) as primary causes. Fix: add a timeout and a fallback; validate by watching the p99 metric drop back to baseline over the next hour, not just the fix compiling.
Trade-offs and pitfalls
The biggest failure mode is skipping hypothesis formation and going straight to your favorite tool (attaching a profiler because you're comfortable with it, even when a five-minute log check would have ruled out two hypotheses first). The second is treating the first correlated signal as the cause without checking whether it's actually causal. Under time pressure it's tempting to fix the first plausible thing you see; that's fine as a mitigation, but the loop isn't complete until you've confirmed the metric recovered and understood why, or you will be back debugging the same symptom next week.
You receive a bug report: a routine that removes duplicates from an array in place is intermittently failing with out-of-bounds writes in production, but it works fine in tests. Describe how you would debug this, what tests you would add, and which language-specific pitfalls, such as signed versus unsigned indices, integer overflow, or aliasing, you would check first.
Sample Answer
Direct answer
The out-of-bounds write is almost certainly a signed/unsigned index mismatch or an off-by-one in the removal logic itself, both classic in-place-array-manipulation pitfalls that "works in tests" precisely because typical test arrays are small and don't happen to exercise the exact boundary condition that triggers the bug, while production data occasionally does.
Structured elaboration
- Reproduce with a boundary-focused test case first, rather than trying to guess from code review alone: specifically test with an array where duplicates occur near the END of the array (the last one or two elements), since in-place removal algorithms that shift elements left as duplicates are removed are especially prone to write-index errors exactly at the tail, where there's less "room" for an off-by-one to go unnoticed.
- Check for a signed/unsigned index mismatch specifically, since the question calls it out directly: if the write index is computed via a subtraction that can go negative in an edge case (e.g., removing duplicates from a very short array, or one that's ENTIRELY duplicates), and that index is stored in an unsigned integer type, a negative result wraps around to a huge positive value, producing a write far outside the array's actual bounds, which is a classic C/C++/embedded-systems bug (also possible in any language with explicit unsigned integer types) and matches the "out-of-bounds write" symptom precisely, as opposed to a simple off-by-one which would typically write just one slot too far, not wildly out of bounds.
- Check for integer overflow in any index arithmetic if the array or index values could be large enough to approach the integer type's limit, though this is less likely to be the specific cause here unless the array sizes involved are unusually large.
- Check for aliasing/in-place-mutation hazards: if the removal logic reads from and writes to the SAME underlying array simultaneously (common in an in-place algorithm), a read that should happen BEFORE a corresponding write, but doesn't due to a loop-ordering bug, can produce corrupted results distinct from, but sometimes confused with, an out-of-bounds write; confirming which specific failure mode is occurring (via the debugger or a sanitizer) avoids fixing the wrong mechanism.
- What tests to add once the specific mechanism is found: boundary cases specifically (duplicates at the very start, very end, an array that's entirely duplicates, an array with no duplicates at all, a single-element array, an empty array), since these are exactly the cases most likely to expose an off-by-one or signed/unsigned bug that a "typical" test case with duplicates scattered comfortably in the middle would never exercise.
Worked example
Reproducing with an array that's ENTIRELY duplicate values (an aggressive boundary case) triggers the crash reliably, where a more typical mixed test case doesn't. Stepping through with a debugger shows the write index computed as write_pos = write_pos - 1 at one point in the loop, intended to "back up" one slot when a duplicate is detected, but under this specific input pattern, write_pos reaches zero and then goes negative on the next iteration; since it's declared as an unsigned type, that negative value wraps to a very large positive number, and the subsequent array write at that "index" is wildly out of bounds, matching the crash symptom exactly. Fix: change the index variable to a signed type (allowing the intermediate negative value to be handled correctly, with an explicit check before it's ever used as an actual array index) or restructure the loop logic to avoid ever needing to decrement below zero in the first place, whichever better fits the surrounding code's conventions.
Trade-offs and pitfalls
A naive fix like clamping the index to zero whenever it "looks" negative, without understanding WHY it went negative, risks silently producing a different, subtler correctness bug (skipping or duplicating an element) instead of a crash, which is arguably worse, since it fails silently rather than loudly; understanding the exact arithmetic that produces the negative value, and fixing the LOGIC rather than just guarding the symptom, is what actually resolves it correctly.
Unlock Full Question Bank
Get access to all 42 Debugging and Systematic Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.