Debugging and Systematic Troubleshooting Questions
Diagnosing defects methodically: reproducing failures, forming and testing hypotheses, reading stack traces and logs, bisecting changes, and reasoning about error handling and edge cases. Covers a disciplined root-cause approach that applies from local bugs to production issues, distinct from embedded hardware-level debugging. A universally probed engineering-craft skill.
A production API sometimes returns elements in an inconsistent order across clients because sets are used internally. You are responsible for triage: how do you investigate, explain the nondeterminism to stakeholders, and implement a stable ordering for the API output while keeping acceptable performance?
Sample Answer
Direct answer
Sets in most languages make no guarantee about iteration order, so building API output directly from set iteration produces order that can legitimately differ across processes, language/runtime versions, or even between runs of the SAME process, depending on internal hash-table implementation details; the investigation confirms this mechanism directly, then implements a stable, explicit ordering rather than relying on incidental set-iteration behavior.
Structured elaboration
How to investigate: confirm the specific code path building the API response iterates over a set (or a dict/map in a language where iteration order isn't guaranteed) rather than a list or an explicitly sorted structure; reproduce by calling the endpoint multiple times, or across multiple server instances/processes, and diffing the returned element order directly, which should show inconsistency if a set is indeed the cause, versus consistent-but-simply-unexpected order from some other source (like a database query with no explicit ORDER BY, a related but distinct cause worth ruling out with the same investigative approach).
Explaining the nondeterminism to stakeholders: frame it precisely: this isn't a random or buggy failure, it's the EXPECTED behavior of an unordered collection, and the API was implicitly promising an ordering guarantee it was never actually designed to provide; different clients (or the same client at different times) can legitimately see different orderings today, which may have gone unnoticed as long as most callers didn't depend on order, until a specific consumer's logic (or a stricter test) started depending on stability that was never actually guaranteed.
Implementing stable ordering while maintaining acceptable performance:
- Sort explicitly at the point of serialization, using whatever ordering makes sense for the API's actual semantics (alphabetical, insertion order if that's meaningful and trackable, or a natural key like an ID or timestamp); for most APIs, the cost of sorting a response-sized collection (typically not enormous) is negligible compared to the request's other costs (network, serialization itself).
- If insertion order specifically needs to be preserved (and the language's default set doesn't track it), switch to an ordered-set-like structure if the language provides one (some languages/standard libraries offer collections that combine set semantics with insertion-order iteration), avoiding a separate sort step while still gaining determinism.
- For very large collections where sorting cost genuinely matters, consider whether the ordering can be established earlier in the pipeline (e.g., if the data already comes from a sorted source like a database query with an explicit
ORDER BY, preserving that order through to the response rather than passing it through an unordered set at any intermediate step) rather than re-sorting a large collection at serialization time on every request.
Worked example
Confirming the mechanism: the endpoint's handler collects results into a set (used originally just to deduplicate, with no awareness that its iteration order would become externally visible), then serializes that set directly to the response. Diffing repeated calls to the same endpoint shows genuinely different orderings across calls, confirming set-iteration nondeterminism as the mechanism (as opposed to, for example, a database query lacking an explicit sort, which would tend to be consistent WITHIN one server/database session but could still differ across sessions or after a schema change, a related but mechanistically distinct possibility worth ruling out explicitly rather than assuming). Fix: after deduplicating via the set (keeping that step, since dedup itself is still correct and desired), explicitly convert to a list and sort it by a natural, stable key (the item's own ID) before serializing, at negligible added cost relative to the rest of the request, resolving the nondeterminism while preserving the original deduplication behavior.
Trade-offs and pitfalls
The temptation to "fix" this by simply switching the internal data structure to something that happens to iterate in insertion order today, without an EXPLICIT sort, risks re-introducing the same class of bug if the underlying collection or its implementation ever changes in a future language/runtime version; an explicit, intentional sort at the serialization boundary is more robust than relying on an implementation detail of whatever collection happens to be used internally, even if that detail is currently observed to be stable.
Your payment provider intermittently returns 502 errors, causing checkout failures for customers. Describe the debugging steps you would take to determine whether the root cause is your own integration, the provider itself, the network, or a configuration issue, and describe short-term mitigations to reduce customer impact while you investigate.
Sample Answer
Direct answer
First establish WHO owns the fault, before trying to fix anything: check whether the 502s originate from your own service (a bug in how you call the provider), the network path between you and them, your configuration (wrong endpoint, expired credentials, timeout misconfigured), or the provider itself. A 502 specifically means an upstream server returned an invalid response to a gateway, which already narrows the search: it usually means something between you and the provider's actual application server broke, not that your request was malformed (that would more often be a 4xx).
Structured elaboration
- Check your own logs for the exact request/response pair. Capture the full request you sent (headers, body, timing) and the full 502 response, including any body the provider returned. A 502 with a body from the provider's own error page suggests an issue near their edge, not deep in their systems.
- Check timing and correlation. Is the failure rate correlated with your request VOLUME (suggests rate limiting or a capacity issue on their side), with TIME OF DAY (suggests their scheduled maintenance or your own traffic pattern), or with a SPECIFIC request shape (suggests your own payload triggers an edge case in their processing)?
- Rule out your own network and configuration. Confirm DNS resolves to the expected endpoint, TLS handshakes succeed, and you're not accidentally hitting a sandbox/staging URL in production. Check whether a recent config or credential change on your side coincides with when the 502s started.
- Check the provider's status page and support channels. Many payment providers publish real-time incident status; if others report the same symptom at the same time, that's strong evidence the fault is upstream, not yours.
- Determine if it's truly intermittent or has a pattern. A steady low background rate of 502s can be normal for any third-party dependency at scale; a sudden step-change is what actually indicates an incident, on either side.
Worked example
A checkout service seeing 502s from a payment provider: logs show the failures are NOT correlated with request volume (rules out simple rate limiting) but ARE correlated with a specific payment method (Apple Pay tokens specifically), while card payments succeed at the normal rate. That pattern points at the provider's Apple-Pay-specific processing path, not a general outage or a problem in the checkout service's own code, since the same service, same network path, and same general request shape succeed for card payments. Confirmed by checking the provider's status page, which shows a partial incident affecting exactly that payment method.
Trade-offs and pitfalls
Short-term mitigation while you investigate: implement a retry with backoff for 502s specifically (a 502 is often transient), and if you can identify a stable pattern like the Apple-Pay-specific example, temporarily route that payment method to a fallback or clearly surface the failure to the user rather than silently retrying a request that will keep failing the same way. The trap is assuming "third-party" automatically means "not my problem to investigate further": even a genuine provider-side incident is worth root-causing on your end, both to build an accurate mitigation and because sometimes what looks like a provider outage is actually your own malformed request that only fails for a specific payload shape.
Given a function that sometimes throws an exception deep inside a library you cannot modify, explain how you would instrument the codebase to capture a stack trace and the relevant contextual variables without changing the library's own code. Describe approaches for at least two of Python, Java, or Node.js.
Sample Answer
Direct answer
You can capture a stack trace and the relevant local state without touching the library's source by instrumenting at the BOUNDARY where you call into it: wrap the call site in a try/catch (or its language equivalent) that re-raises after logging, install a global/uncaught-exception hook that fires regardless of where the exception originates, or attach a debugger/tracer that breaks on the exception type without modifying any code at all.
Structured elaboration
Three complementary approaches, from least to most invasive:
- Wrap the call site, not the library. Since you control the code that CALLS into the library, even if you can't modify the library itself, catching the exception at your own call site and logging the full stack trace, the exception type/message, and any locally-relevant variables (the arguments you passed in, the current request context) before re-raising gives you a permanent, low-overhead capture point. In Python:
try: library_call(...) except Exception: logger.exception("context: %s", relevant_vars); raise. In Java: a try/catch around the call that logs viae.printStackTrace()or a structured logger before rethrowing. In Node.js: wrapping in a try/catch for synchronous calls, or.catch()on the returned Promise for async ones. - Install a global uncaught-exception/unhandled-rejection hook. This catches cases where the exception surfaces somewhere you didn't anticipate wrapping: Python's
sys.excepthook, Java'sThread.setDefaultUncaughtExceptionHandler, Node'sprocess.on('uncaughtException', ...)andprocess.on('unhandledRejection', ...). These fire regardless of exactly where inside the library the exception originated, at the cost of being a last-resort catch-all rather than a precisely-scoped one. - Attach a debugger or tracer without changing code, when you need MORE than a stack trace, such as full local variable state at the moment of the throw. Most debuggers support "break on exception" (Python's
pdbwithpython -m pdb, orimport pdb; pdb.set_trace()combined with a signal, or a conditional breakpoint set on the exception type in a full IDE debugger; Java's debugger supports exception breakpoints directly in most IDEs) which pauses execution at the exact throw site, inside the library's own code, letting you inspect the full call stack and locals without ever editing the library.
Worked example
A third-party HTTP client library occasionally raises an unlabeled ConnectionError with no useful message, and you need to know what request triggered it. Wrapping the call site: try: response = third_party_client.get(url, timeout=5) except ConnectionError: logger.exception("third_party_client failed for url=%s, timeout=%s", url, 5); raise. This doesn't change a single line of the library, but every future occurrence now logs the exact URL and timeout that triggered it, alongside the library's own stack trace, immediately giving you the context needed to distinguish "this URL is consistently unreachable" from "this only fails under a specific timeout value."
Trade-offs and pitfalls
A global uncaught-exception hook is a safety net, not a substitute for a targeted wrap at the call site: it tells you SOMETHING failed somewhere, but without the call-site context (what arguments were in play, what business operation was in progress) it's often not enough to actually diagnose the issue, only to know it happened. The debugger approach is the most informative but requires either a reproducible local trigger or an interactive session, so it's best reserved for cases where the logged stack trace alone isn't enough to form a hypothesis.
Tell the story of a concrete bug or production failure you found. Explain how you detected it, how you reproduced it if that was possible, the debugging tools and techniques you used, the root cause, and the permanent fix you implemented.
Sample Answer
Direct answer
A concrete story: a service occasionally returned stale pricing data to a subset of users, detected via a customer complaint rather than any internal alert (since the values were plausible-looking, just wrong, not obviously broken); the root cause traced to a caching layer that keyed its cache entries incorrectly, causing two logically-distinct pricing contexts to collide and overwrite each other's cached value, and the permanent fix corrected the cache key's uniqueness rather than just adjusting the cache's expiry time.
Structured elaboration
How it was detected: a customer support ticket reported seeing a price that didn't match what should have applied to their account tier, with no corresponding error or alert on the engineering side, since the returned value was a real, validly-formatted price, just the WRONG one; this is a useful detail because it illustrates a class of bug (returning plausible-but-wrong data) that's structurally invisible to error-rate-based monitoring, and only surfaces via a downstream consumer noticing a substantive discrepancy.
How it was reproduced: confirming the report wasn't a one-off required identifying the PATTERN, not just the single instance; checking whether other users on the same account tier around the same time window also received an unexpected price showed a small but real cluster, ruling out "one weird one-off" and confirming a systemic, reproducible mechanism worth a full investigation.
Debugging tools and techniques used: traced the pricing-lookup code path for the affected requests, and found it flows through an in-memory cache keyed, it turned out, on account tier ALONE rather than on the combination of account tier AND region (pricing legitimately varies by both); when two users on the same tier but different regions made requests close together in time, the second request's result could overwrite the first's cache entry under the shared, insufficiently-specific key, and a THIRD user (same tier, either region) arriving shortly after could then receive whichever region's price happened to be cached most recently, regardless of their own actual region.
The root cause: a cache key that didn't include every dimension the underlying value actually varied by, a classic caching-correctness bug: the cache was implicitly promising "this value is valid for anyone with this tier," when the real invariant needed was "this value is valid for anyone with this tier AND this region."
The permanent fix implemented: updated the cache key to include region alongside tier, restoring the correct invariant; also added a specific integration test that exercises exactly this scenario (two regions, same tier, interleaved requests) to catch a regression of this specific mechanism in the future, since the original bug had shipped without any test covering this particular combination of dimensions.
What you learned that helps you avoid similar bugs: whenever introducing a cache, explicitly enumerate every dimension the cached value can legitimately vary by, and verify the cache key includes ALL of them, not just the ones that happen to be obvious or top-of-mind at implementation time; a caching bug of this shape is especially dangerous specifically because it fails SILENTLY (no error, no crash, just occasionally-wrong data) and is invisible to typical error-rate monitoring, which argues for treating "does this cache key capture every dimension of variation" as a deliberate design-review question on any future caching work, not something to verify only after a bug report arrives.
Trade-offs and pitfalls
The tempting quick fix, once the symptom (stale/wrong cached price) was understood, would have been to simply reduce the cache's TTL (expiry time), which would have reduced the WINDOW during which a collision could produce visibly wrong data without fixing the actual collision mechanism at all; the permanent fix specifically addressed the cache KEY's correctness, not the expiry duration, since a shorter TTL would have masked the bug's visible frequency without removing its root cause.
Provide a succinct example of a recent small bug you fixed. Include the minimal code before and after in a language of your choice, explain the root cause, why your fix works, how you tested it, and what you learned that will help you avoid similar bugs in the future.
Sample Answer
Direct answer
A recent example: a function computing a running total returned a subtly wrong result specifically when passed an empty input list, because it assumed at least one element would always be present; the root cause was an unchecked assumption about the input's shape, not a logic error in the calculation itself, and the fix was a one-line explicit guard plus a test covering the previously-unconsidered case.
Structured elaboration and worked example
Before:
def average_response_time(durations):
total = sum(durations)
return total / len(durations)
This looks correct and passes any test using a non-empty list. The bug: called with an empty list (a genuinely possible input, e.g., a time window with zero recorded requests), it raises ZeroDivisionError, which in this specific case was crashing a reporting job whenever a low-traffic service had a quiet hour with literally zero requests, a real, recurring production scenario the original code never considered.
After:
def average_response_time(durations):
if not durations:
return None # no data for this window; caller decides how to display that
total = sum(durations)
return total / len(durations)
Root cause: the function was written and tested against realistic-looking sample data, which always happened to be non-empty, so the implicit assumption "there's always at least one duration" was never challenged during development; the bug surfaced only once real production traffic included a genuinely empty window, an edge case that's easy to overlook precisely because it's uncommon rather than because it's hard to reason about once you think to check for it.
How it was tested: added an explicit unit test calling the function with an empty list and asserting it returns None rather than raising, alongside the existing non-empty-input tests, so the specific edge case that caused the production issue is now permanently covered and can't silently regress.
What was learned to avoid similar bugs: for any function that aggregates over a collection, explicitly consider and test the empty-collection case as a matter of habit, not just when a bug report forces the question; more broadly, the pattern generalizes to any input assumption implicit in code (non-null, non-negative, within some expected range) that reads as "obviously always true" during development but isn't guaranteed by the function's actual contract, and is worth deliberately challenging each such assumption with a test rather than trusting that realistic-looking sample data will happen to cover it.
Trade-offs and pitfalls
The fix here (returning None for an empty input) is a specific design choice, not the only valid one; an alternative would be raising a more informative, explicit exception (ValueError("cannot average an empty list")) if silently returning None risks a caller mishandling it further downstream without noticing. The right choice depends on what callers actually need to do with a "no data" result, and is worth deciding deliberately rather than defaulting to whichever felt fastest to write.
Unlock Full Question Bank
Get access to all 24 Debugging and Systematic Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.