Debugging and Systematic Troubleshooting Questions
Diagnosing defects methodically: reproducing failures, forming and testing hypotheses, reading stack traces and logs, bisecting changes, and reasoning about error handling and edge cases. Covers a disciplined root-cause approach that applies from local bugs to production issues, distinct from embedded hardware-level debugging. A universally probed engineering-craft skill.
Your payment provider intermittently returns 502 errors, causing checkout failures for customers. Describe the debugging steps you would take to determine whether the root cause is your own integration, the provider itself, the network, or a configuration issue, and describe short-term mitigations to reduce customer impact while you investigate.
Sample Answer
Direct answer
First establish WHO owns the fault, before trying to fix anything: check whether the 502s originate from your own service (a bug in how you call the provider), the network path between you and them, your configuration (wrong endpoint, expired credentials, timeout misconfigured), or the provider itself. A 502 specifically means an upstream server returned an invalid response to a gateway, which already narrows the search: it usually means something between you and the provider's actual application server broke, not that your request was malformed (that would more often be a 4xx).
Structured elaboration
- Check your own logs for the exact request/response pair. Capture the full request you sent (headers, body, timing) and the full 502 response, including any body the provider returned. A 502 with a body from the provider's own error page suggests an issue near their edge, not deep in their systems.
- Check timing and correlation. Is the failure rate correlated with your request VOLUME (suggests rate limiting or a capacity issue on their side), with TIME OF DAY (suggests their scheduled maintenance or your own traffic pattern), or with a SPECIFIC request shape (suggests your own payload triggers an edge case in their processing)?
- Rule out your own network and configuration. Confirm DNS resolves to the expected endpoint, TLS handshakes succeed, and you're not accidentally hitting a sandbox/staging URL in production. Check whether a recent config or credential change on your side coincides with when the 502s started.
- Check the provider's status page and support channels. Many payment providers publish real-time incident status; if others report the same symptom at the same time, that's strong evidence the fault is upstream, not yours.
- Determine if it's truly intermittent or has a pattern. A steady low background rate of 502s can be normal for any third-party dependency at scale; a sudden step-change is what actually indicates an incident, on either side.
Worked example
A checkout service seeing 502s from a payment provider: logs show the failures are NOT correlated with request volume (rules out simple rate limiting) but ARE correlated with a specific payment method (Apple Pay tokens specifically), while card payments succeed at the normal rate. That pattern points at the provider's Apple-Pay-specific processing path, not a general outage or a problem in the checkout service's own code, since the same service, same network path, and same general request shape succeed for card payments. Confirmed by checking the provider's status page, which shows a partial incident affecting exactly that payment method.
Trade-offs and pitfalls
Short-term mitigation while you investigate: implement a retry with backoff for 502s specifically (a 502 is often transient), and if you can identify a stable pattern like the Apple-Pay-specific example, temporarily route that payment method to a fallback or clearly surface the failure to the user rather than silently retrying a request that will keep failing the same way. The trap is assuming "third-party" automatically means "not my problem to investigate further": even a genuine provider-side incident is worth root-causing on your end, both to build an accurate mitigation and because sometimes what looks like a provider outage is actually your own malformed request that only fails for a specific payload shape.
Find and fix the bug in this JavaScript async function, where a missing await causes unpredictable ordering and errors:
async function processItems(items, processor) {
items.forEach(async item => {
await processor(item);
});
console.log('done');
}
Explain the root cause and provide a corrected version that guarantees 'done' prints only after every item has been processed.
Sample Answer
Direct answer
Array.prototype.forEach does not await its callback: it fires each async callback and moves on immediately, so console.log('done') runs before any of the await processor(item) calls have resolved. The fix is to replace forEach with a construct that actually waits, either a for...of loop with await inside it (sequential) or Promise.all over items.map(...) (concurrent).
Structured elaboration
forEach was designed before async/await existed and its callback's return value, including a Promise, is simply discarded. Passing an async function as the callback doesn't change this: each invocation still returns a Promise that forEach never looks at, so forEach itself completes synchronously (having merely started every item, not finished any of them) the instant it has called the callback once per array element. The console.log('done') line after the forEach call then runs immediately, while the async work is still in flight in the background.
Worked example
async function sleep(ms) { return new Promise(r => setTimeout(r, ms)); }
// Fixed, distinct per-item delays (not Math.random()) so the interleaving below
// is deterministic and reproduces identically on every run, not just illustrative.
const delays = { 1: 20, 2: 10, 3: 30 };
// BUGGY
async function processItemBuggy(x) { await sleep(delays[x]); console.log(' processed', x, '(buggy)'); }
async function processBuggy(items) {
items.forEach(async item => { await processItemBuggy(item); });
console.log('done (buggy)');
}
// FIXED: sequential
async function processItemFixed(x) { await sleep(delays[x]); console.log(' processed', x, '(fixed)'); }
async function processFixed(items) {
for (const item of items) { await processItemFixed(item); }
console.log('done (fixed)');
}
(async () => {
console.log('--- buggy run ---');
await processBuggy([1, 2, 3]);
await sleep(100); // let the buggy run's background work finish before the next section starts
console.log('--- fixed run (sequential for-of) ---');
await processFixed([1, 2, 3]);
})();
Executed output (buggy items finish asynchronously in the background; with the fixed per-item delays above, item 2 always resolves first, then item 1, then item 3, after done has already printed):
--- buggy run ---
done (buggy)
processed 2 (buggy)
processed 1 (buggy)
processed 3 (buggy)
--- fixed run (sequential for-of) ---
processed 1 (fixed)
processed 2 (fixed)
processed 3 (fixed)
done (fixed)
done (buggy) prints immediately, before any processed line: forEach fired all three async callbacks and moved on without waiting for any of them. The processed (buggy) lines do eventually print, once their fixed delays resolve, in delay order (2, then 1, then 3, not the input order), rather than before done. The for...of version correctly prints all three processed lines, strictly in input order, before done.
Trade-offs and pitfalls
There are two valid fixes with different semantics, and picking the wrong one is itself a common mistake: a for...of loop processes items strictly one at a time (useful when order matters or you must not overwhelm a downstream dependency with concurrent calls), while await Promise.all(items.map(item => processor(item))) processes all items concurrently and finishes as soon as the slowest one does (faster, but only safe if the operations are independent and the downstream system can handle concurrent load). Silently reaching for forEach out of habit, rather than deliberately choosing sequential versus concurrent semantics, is exactly how this bug class gets reintroduced even by developers who already know the rule.
You have this Python function intended to remove duplicates from a list while preserving order:
def dedupe(items):
seen = set()
result = []
for x in items:
if x not in seen:
result.append(x)
return result
Run on [1, 2, 1] it returns [1, 2, 1] instead of [1, 2]. Explain why, provide a corrected implementation, and state its time and space complexity.
Sample Answer
Direct answer
The function checks if x not in seen but never actually adds x to seen, so seen stays empty forever and every element passes the "not seen" check, meaning nothing is ever filtered out. The fix is one missing line: seen.add(x) inside the if block, right after appending to result.
Structured elaboration
The intended algorithm is correct in structure: track which values have already been emitted in a set for O(1) membership checks, and only append a value to the result if it hasn't been seen before. The bug is a single omitted mutation, not a logic error in the surrounding control flow, which is exactly the kind of defect that's easy to miss on a quick read because the code LOOKS complete: it has a seen set, it has a membership check, it just never updates the set it's checking against.
Worked example
def dedupe_buggy(items):
seen = set()
result = []
for x in items:
if x not in seen:
result.append(x)
# BUG: forgot seen.add(x)
return result
def dedupe_fixed(items):
seen = set()
result = []
for x in items:
if x not in seen:
result.append(x)
seen.add(x)
return result
print("buggy:", dedupe_buggy([1, 2, 1]))
print("fixed:", dedupe_fixed([1, 2, 1]))
print("fixed larger:", dedupe_fixed([3, 1, 4, 1, 5, 9, 2, 6, 5, 3, 5]))
Executed with python3:
buggy: [1, 2, 1]
fixed: [1, 2]
fixed larger: [3, 1, 4, 5, 9, 2, 6]
The buggy version returns the input completely unchanged, since seen never gains any members and the check x not in seen is always true. The fixed version correctly keeps only the first occurrence of each value, preserving original order, and the larger example confirms it generalizes correctly beyond the minimal [1, 2, 1] case.
Complexity: O(n) time, since each element does one O(1) average-case set membership check and at most one O(1) set insertion; O(n) space in the worst case (all elements distinct), for both the seen set and the result list.
Trade-offs and pitfalls
This bug is a strong argument for testing with an input specifically designed to exercise the "already seen" branch, since a test using only distinct elements (e.g., [1, 2, 3]) would pass against BOTH the buggy and fixed versions, never exposing the missing line. The general lesson: when a function's entire purpose is to filter based on accumulated state, any test suite for it should include at least one input where that state is actually expected to change the output, not just inputs where the filtering happens to be a no-op.
Write a memory-efficient Python function parse_log_counts(file_path, top_n=5) that scans a large server log file, which may be larger than available memory, and returns the top N error types or HTTP status codes by count without loading the whole file into memory. Example lines: 2025-01-01T12:00:00Z INFO request_id=1 status=200 and 2025-01-01T12:00:01Z ERROR request_id=2 status=500 exception=ValueError. Describe edge cases and how your implementation handles gzipped logs and malformed lines.
Sample Answer
Direct answer
Stream the file line by line instead of loading it into memory, maintaining only a bounded Counter of error-type/status-code totals as you go, which keeps memory usage proportional to the number of DISTINCT keys seen rather than the number of lines in the file. Handling gzipped input just means picking the right file-opening function based on the extension; handling malformed lines means skipping them without crashing the whole scan.
Structured elaboration
import gzip
from collections import Counter
from typing import List, Tuple
def parse_log_counts(file_path: str, top_n: int = 5) -> List[Tuple[str, int]]:
counts = Counter()
opener = gzip.open if file_path.endswith('.gz') else open
with opener(file_path, 'rt', errors='replace') as f:
for line in f:
line = line.rstrip('\n')
if not line:
continue
status, exc = None, None
for tok in line.split():
if tok.startswith('status='):
status = tok.split('=', 1)[1]
elif tok.startswith('exception='):
exc = tok.split('=', 1)[1]
if exc:
counts[f'exception:{exc}'] += 1
elif status:
counts[f'status:{status}'] += 1
# lines with neither field are silently skipped as malformed
return counts.most_common(top_n)
Key design points:
- Iterating
for line in freads one line at a time from the underlying file object rather than materializing the whole file in memory; this is what makes the function safe against a 10GB+ input. gzip.openversusopenare chosen based on the file extension, so the same function transparently handles both compressed and uncompressed logs without the caller needing to know which.errors='replace'on the text-mode open prevents a single malformed byte sequence (a truncated write, a binary artifact mid-file) from raising aUnicodeDecodeErrorand aborting the entire scan.- Lines with neither a
status=norexception=token are simply skipped rather than raising, satisfying "handle malformed lines" without crashing on them. Counter.most_common(top_n)does the top-N selection in O(k log n) where k is the number requested, more efficient than sorting the entire counts dictionary when only a few top entries are needed.
Worked example
A synthetic log with 50 status=200 lines, 12 exception=ValueError lines, 8 exception=TimeoutError lines, 20 status=404 lines, plus two malformed lines (one garbage line, one blank):
plain file: [('status:200', 50), ('status:404', 20), ('exception:ValueError', 12), ('exception:TimeoutError', 8)]
gzipped file: [('status:200', 50), ('status:404', 20), ('exception:ValueError', 12), ('exception:TimeoutError', 8)]
top_n=2: [('status:200', 50), ('status:404', 20)]
The gzipped and plain versions of the identical content produce byte-identical results, confirming the compression branch works correctly, and top_n=2 correctly truncates to the two highest counts without needing a separate code path.
Edge cases
- Empty file: the loop simply never executes,
countsstays empty, andmost_common(top_n)returns an empty list, no special-casing needed. - File with only malformed lines: same result, an empty list, which is the correct behavior (nothing to report) rather than an error.
- A line with BOTH
status=andexception=: the current implementation prioritizesexception:(checked first), which is a deliberate choice since an exception is generally the more specific and actionable signal; this priority should be called out explicitly rather than left as an implicit accident of code order. top_nlarger than the number of distinct keys:most_commonsimply returns all available entries, no error.
Trade-offs and pitfalls
This implementation keeps ALL distinct keys in memory (bounded by the number of distinct error types/status codes, not the number of log lines, which is normally small), but if the log contained a very high-cardinality field mistakenly counted this way (say, counting by full request URL instead of status code), memory could still grow unboundedly with the number of distinct values; the safety guarantee here specifically relies on error types and status codes being a naturally small, bounded set.
Describe the role of instrumentation (logs, metrics, traces) in effective debugging. Give a concise checklist of five things you would verify are in place before handing a service off to operations for production use.
Sample Answer
Direct answer
Before handing a service to operations, verify: (1) every meaningful failure path logs enough context to diagnose it without a redeploy, (2) the four golden signals (latency, traffic, errors, saturation) are exposed as metrics, not just logs, (3) a request can be traced end to end when it crosses more than one service, (4) alert thresholds exist and point to an actual runbook, not just a page with no next step, and (5) someone other than the author has actually looked at the dashboards and confirmed they answer "is this healthy right now" at a glance.
Structured elaboration
Each item on the checklist exists because of a specific failure mode it prevents:
- Actionable error logs. A log line that says "operation failed" with no request ID, no input summary, and no stack trace forces on-call to redeploy with more logging just to understand a 2am page. The bar: could someone who has never read this code diagnose the failure category from the log line alone?
- The four golden signals as metrics, not just logs. Logs answer "what happened in this one case"; metrics answer "is this normal right now." Without dashboards for latency, traffic, error rate, and saturation (CPU/memory/queue depth/connection pool usage), operations has no way to distinguish a healthy blip from a developing outage without grepping logs under pressure.
- Distributed tracing or correlation IDs. The moment a request crosses a service boundary, "check the logs" stops being a single grep and becomes "which of these five services' logs, and how do I know they're the same request?" A propagated request ID is the cheapest fix and the most commonly missing piece.
- Alert thresholds tied to a runbook. An alert that fires with no documented first step trains on-call to snooze it, which is worse than no alert at all: it becomes noise that hides the next real incident.
- A second set of eyes on the dashboards. The author of a service is the worst-positioned person to judge whether their own dashboard is readable to someone unfamiliar with the code, the same blind spot that makes self-review of documentation unreliable in general.
Worked example
A concrete pre-handoff review of a new payment-retry service: logs include payment_id, attempt_number, and the specific failure reason on every retry (satisfies #1); a Grafana panel shows retry rate, success rate, and queue depth (satisfies #2); the payment_id is propagated as a header to the downstream charge service so both services' logs can be joined (satisfies #3); the "retry queue depth > 500" alert links directly to a runbook section titled "Retry queue backing up" with three ranked likely causes (satisfies #4); a teammate who did not write the service opened the dashboard cold and correctly identified within 30 seconds whether the system was healthy (satisfies #5).
Trade-offs and pitfalls
The most common gap isn't missing instrumentation entirely, it's instrumentation that only makes sense to the person who wrote it: log lines with internal variable names instead of business-meaningful fields, dashboards with no annotations explaining what "normal" looks like, or alerts that reference a metric name with no context. The checklist is deliberately about READINESS for someone else to operate the system, not about whether instrumentation exists in principle.
Unlock Full Question Bank
Get access to all 25 Debugging and Systematic Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.