Debugging and Systematic Troubleshooting Questions
Diagnosing defects methodically: reproducing failures, forming and testing hypotheses, reading stack traces and logs, bisecting changes, and reasoning about error handling and edge cases. Covers a disciplined root-cause approach that applies from local bugs to production issues, distinct from embedded hardware-level debugging. A universally probed engineering-craft skill.
Your payment provider intermittently returns 502 errors, causing checkout failures for customers. Describe the debugging steps you would take to determine whether the root cause is your own integration, the provider itself, the network, or a configuration issue, and describe short-term mitigations to reduce customer impact while you investigate.
Sample Answer
Direct answer
First establish WHO owns the fault, before trying to fix anything: check whether the 502s originate from your own service (a bug in how you call the provider), the network path between you and them, your configuration (wrong endpoint, expired credentials, timeout misconfigured), or the provider itself. A 502 specifically means an upstream server returned an invalid response to a gateway, which already narrows the search: it usually means something between you and the provider's actual application server broke, not that your request was malformed (that would more often be a 4xx).
Structured elaboration
- Check your own logs for the exact request/response pair. Capture the full request you sent (headers, body, timing) and the full 502 response, including any body the provider returned. A 502 with a body from the provider's own error page suggests an issue near their edge, not deep in their systems.
- Check timing and correlation. Is the failure rate correlated with your request VOLUME (suggests rate limiting or a capacity issue on their side), with TIME OF DAY (suggests their scheduled maintenance or your own traffic pattern), or with a SPECIFIC request shape (suggests your own payload triggers an edge case in their processing)?
- Rule out your own network and configuration. Confirm DNS resolves to the expected endpoint, TLS handshakes succeed, and you're not accidentally hitting a sandbox/staging URL in production. Check whether a recent config or credential change on your side coincides with when the 502s started.
- Check the provider's status page and support channels. Many payment providers publish real-time incident status; if others report the same symptom at the same time, that's strong evidence the fault is upstream, not yours.
- Determine if it's truly intermittent or has a pattern. A steady low background rate of 502s can be normal for any third-party dependency at scale; a sudden step-change is what actually indicates an incident, on either side.
Worked example
A checkout service seeing 502s from a payment provider: logs show the failures are NOT correlated with request volume (rules out simple rate limiting) but ARE correlated with a specific payment method (Apple Pay tokens specifically), while card payments succeed at the normal rate. That pattern points at the provider's Apple-Pay-specific processing path, not a general outage or a problem in the checkout service's own code, since the same service, same network path, and same general request shape succeed for card payments. Confirmed by checking the provider's status page, which shows a partial incident affecting exactly that payment method.
Trade-offs and pitfalls
Short-term mitigation while you investigate: implement a retry with backoff for 502s specifically (a 502 is often transient), and if you can identify a stable pattern like the Apple-Pay-specific example, temporarily route that payment method to a fallback or clearly surface the failure to the user rather than silently retrying a request that will keep failing the same way. The trap is assuming "third-party" automatically means "not my problem to investigate further": even a genuine provider-side incident is worth root-causing on your end, both to build an accurate mitigation and because sometimes what looks like a provider outage is actually your own malformed request that only fails for a specific payload shape.
Describe a systematic, repeatable approach you use to troubleshoot an unfamiliar technical problem end to end. Cover how you observe the symptom, form and prioritize hypotheses, gather and interpret evidence such as logs, metrics, and traces, isolate the root cause, implement and validate a fix, and decide when to escalate, roll back, or write up a postmortem.
Sample Answer
Direct answer
Troubleshooting an unfamiliar problem is a loop, not a single step: observe the symptom precisely, form a small set of testable hypotheses ranked by likelihood and cost to check, gather evidence that discriminates between them, isolate the true cause, implement and validate a fix, and decide whether the incident needs a rollback, an escalation, or a written postmortem. The loop repeats: each piece of evidence should narrow the hypothesis set, not just confirm what you already believed.
Structured elaboration
- Observe the symptom precisely. Write down exactly what is wrong, in falsifiable terms: not "the API is slow" but "p99 latency (the response time slower than 99% of requests, i.e. how bad the worst cases are, not just the average) on
POST /ordersrose from 80ms to 900ms starting at 14:32 UTC, affecting roughly 3% of requests." Vague symptoms produce vague hypotheses. - Form hypotheses before you start digging. List the plausible causes given what changed recently (deploys, config, traffic pattern, dependency versions) and what the symptom rules out. A hypothesis you cannot state is a hypothesis you cannot test.
- Prioritize by expected information gain divided by cost. A five-minute log grep that could confirm or kill three hypotheses at once beats a one-hour deep profiling session that only speaks to one.
- Gather evidence that discriminates. Logs tell you what happened at a point; metrics tell you the shape of the problem over time; traces tell you where time went inside one request. Pick the instrument that actually distinguishes your live hypotheses, not the one you're most comfortable with.
- Isolate the root cause, not just a correlated symptom. A dropped hypothesis should be dropped because evidence contradicts it, not because you got bored of it.
- Implement and validate the fix against the same evidence that revealed the problem. If you diagnosed via a specific metric, watch that metric recover before declaring victory.
- Decide what happens next. If customer impact is ongoing and the fix is unproven, roll back first and diagnose second. If a similar failure could recur, or the incident had real impact, write it up so the org doesn't relearn the same lesson.
Worked example
A "the checkout page is slow" report, applied through the loop: symptom precisely stated as "median load time is normal, but a subset of loads takes 8-12 seconds, starting after this morning's deploy." Hypotheses: (a) the new deploy added a blocking call, (b) a downstream dependency degraded independently, (c) the slow subset shares a common attribute (e.g., a specific region or a large cart). A single log query grouping slow requests by attribute would discriminate between (c) and the other two in minutes, before touching a profiler. Suppose it shows the slow requests all hit a newly added inventory-check call to a dependency with no timeout: that both confirms (a) and rules out (b)/(c) as primary causes. Fix: add a timeout and a fallback; validate by watching the p99 metric drop back to baseline over the next hour, not just the fix compiling.
Trade-offs and pitfalls
The biggest failure mode is skipping hypothesis formation and going straight to your favorite tool (attaching a profiler because you're comfortable with it, even when a five-minute log check would have ruled out two hypotheses first). The second is treating the first correlated signal as the cause without checking whether it's actually causal. Under time pressure it's tempting to fix the first plausible thing you see; that's fine as a mitigation, but the loop isn't complete until you've confirmed the metric recovered and understood why, or you will be back debugging the same symptom next week.
A nightly batch job that used to finish in one hour now takes three hours. Provide a prioritized performance-debugging checklist to identify the regression: quick sanity checks, metrics to review, sampling to reproduce the slowdown, configuration and resource changes to inspect, and deeper profiling approaches for CPU, memory, IO, and network bottlenecks.
Sample Answer
Direct answer
For a job that tripled in duration, work through a prioritized checklist that separates "did the WORK grow" from "did something get SLOWER at doing the same work": quick sanity checks on input volume first (the simplest explanation), then metrics review, then targeted reproduction, then configuration/resource changes, then deep profiling only if the cheaper checks don't explain it.
Structured elaboration
- Quick sanity checks (minutes, not hours):
- Compare input data VOLUME between a recent run and a historical baseline run; if row counts, file sizes, or record counts grew substantially, the job may simply be doing more work, not running less efficiently, and the "regression" is actually organic growth outpacing the original capacity plan.
- Check for any recent deploy or config change to the job around when the slowdown started; correlate timestamps before assuming a mysterious, uncaused regression.
- Metrics to review next:
- Per-stage duration breakdown (if the pipeline has multiple stages), to localize WHICH stage grew, rather than treating the job as one opaque block.
- Resource utilization during the run: CPU, memory, and I/O wait, to distinguish a CPU-bound slowdown from an I/O-bound one, which point toward very different causes.
- Shuffle/data-movement bytes, if running on a distributed engine (Spark, etc.), since a shuffle that grew disproportionately to input size often indicates a skew problem (one key or value shows up far more often than the rest, so the work piling up on it takes much longer than everyone else's share) or a join-explosion problem (a join unexpectedly multiplies rows, for example when the join key isn't as unique as assumed), not a proportional slowdown.
- Sampling to reproduce the slowdown at smaller scale: run the job against a smaller, representative slice of current data in a non-production environment, timed against the same slice size from a historical baseline, to confirm the slowdown reproduces outside of production noise and isn't itself a measurement artifact (contention with other jobs sharing the same cluster, for instance).
- Configuration and resource changes to inspect: confirm the job's allocated resources (executor count, memory, parallelism settings) haven't silently changed or been reduced by a shared-cluster resource-management change; confirm no configuration regression (a changed partition count, a disabled optimization flag) crept in via an unrelated deploy.
- Deeper profiling for CPU, memory, IO, and network bottlenecks, once the cheaper checks haven't explained it: CPU profiling to find hot code paths; memory profiling to check for growing GC pressure; I/O and network profiling to check for a downstream dependency that's become slower to read from or write to.
Worked example
Applying the checklist: input volume check shows data grew only about 15% since the baseline, ruling out "just more data" as the primary explanation for a 3x runtime increase. Per-stage breakdown shows one specific stage, a join against a reference table, went from a small fraction of total runtime to dominating it. Shuffle-bytes metrics for that stage show a huge spike disproportionate to the 15% data growth, pointing at a join-explosion or skew problem rather than a proportional slowdown. Investigating the join key's cardinality shows the reference table recently had a batch of records added with a NULL join key, and the engine's join semantics are producing a much larger-than-expected result set for null-key matches. That's the concrete root cause, found by following the metrics from coarse (which stage) to specific (which operation, why), rather than starting with a CPU profiler on the whole job.
Trade-offs and pitfalls
Starting with deep CPU/memory profiling before checking the cheap, coarse signals (input volume, per-stage timing, shuffle bytes) risks spending an hour profiling code that was never the bottleneck, when a five-minute metrics check would have pointed directly at the actual stage and mechanism. The prioritization in this checklist (cheap and broad first, expensive and narrow last) exists specifically to avoid that waste.
Describe the role of instrumentation (logs, metrics, traces) in effective debugging. Give a concise checklist of five things you would verify are in place before handing a service off to operations for production use.
Sample Answer
Direct answer
Before handing a service to operations, verify: (1) every meaningful failure path logs enough context to diagnose it without a redeploy, (2) the four golden signals (latency, traffic, errors, saturation) are exposed as metrics, not just logs, (3) a request can be traced end to end when it crosses more than one service, (4) alert thresholds exist and point to an actual runbook, not just a page with no next step, and (5) someone other than the author has actually looked at the dashboards and confirmed they answer "is this healthy right now" at a glance.
Structured elaboration
Each item on the checklist exists because of a specific failure mode it prevents:
- Actionable error logs. A log line that says "operation failed" with no request ID, no input summary, and no stack trace forces on-call to redeploy with more logging just to understand a 2am page. The bar: could someone who has never read this code diagnose the failure category from the log line alone?
- The four golden signals as metrics, not just logs. Logs answer "what happened in this one case"; metrics answer "is this normal right now." Without dashboards for latency, traffic, error rate, and saturation (CPU/memory/queue depth/connection pool usage), operations has no way to distinguish a healthy blip from a developing outage without grepping logs under pressure.
- Distributed tracing or correlation IDs. The moment a request crosses a service boundary, "check the logs" stops being a single grep and becomes "which of these five services' logs, and how do I know they're the same request?" A propagated request ID is the cheapest fix and the most commonly missing piece.
- Alert thresholds tied to a runbook. An alert that fires with no documented first step trains on-call to snooze it, which is worse than no alert at all: it becomes noise that hides the next real incident.
- A second set of eyes on the dashboards. The author of a service is the worst-positioned person to judge whether their own dashboard is readable to someone unfamiliar with the code, the same blind spot that makes self-review of documentation unreliable in general.
Worked example
A concrete pre-handoff review of a new payment-retry service: logs include payment_id, attempt_number, and the specific failure reason on every retry (satisfies #1); a Grafana panel shows retry rate, success rate, and queue depth (satisfies #2); the payment_id is propagated as a header to the downstream charge service so both services' logs can be joined (satisfies #3); the "retry queue depth > 500" alert links directly to a runbook section titled "Retry queue backing up" with three ranked likely causes (satisfies #4); a teammate who did not write the service opened the dashboard cold and correctly identified within 30 seconds whether the system was healthy (satisfies #5).
Trade-offs and pitfalls
The most common gap isn't missing instrumentation entirely, it's instrumentation that only makes sense to the person who wrote it: log lines with internal variable names instead of business-meaningful fields, dashboards with no annotations explaining what "normal" looks like, or alerts that reference a metric name with no context. The checklist is deliberately about READINESS for someone else to operate the system, not about whether instrumentation exists in principle.
A web service shows high CPU usage but low user-visible latency. Explain the possible causes for this discrepancy, how you would investigate whether the extra CPU is wasted work, background tasks, or a measurement artifact, and what remediation you would propose once you know which it is.
Sample Answer
Direct answer
High CPU with low user-visible latency means the CPU work isn't on the critical path the user is waiting on: it's either background work (batch jobs, garbage collection, async processing), wasted work (a busy-loop, redundant recomputation, an inefficient algorithm that happens to still finish fast enough), or a measurement artifact (CPU metrics counting time the process spends idle-but-scheduled, or double-counting across containers sharing a host). The investigation is about separating these three, not assuming the first plausible one.
Structured elaboration
- Check whether the CPU usage correlates with request volume or is constant/background. If CPU stays high even during low-traffic periods, it's very unlikely to be request-handling work; that points toward background jobs, scheduled tasks, or a stuck loop.
- Profile what's actually consuming CPU, using a sampling profiler (perf, py-spy, async-profiler depending on the runtime) rather than guessing from code review. Distinguish CPU time attributed to request handlers versus background threads, garbage collection, or monitoring/logging agents.
- Check for wasted work specifically: a cache that isn't actually being hit (so every request redoes expensive work that should have been cached), a retry loop that's spinning faster than intended, or an algorithm with much worse complexity than necessary that still completes within the user's timeout because the dataset happens to be small today.
- Rule out measurement artifacts. On containerized/shared hosts, CPU metrics can reflect cgroup accounting quirks (cgroups are the Linux kernel mechanism that caps and meters how much CPU/memory a container is allowed to use, and its usage counters can misreport in edge cases), CPU throttling counted as "usage," or a host-level metric that aggregates multiple co-located processes. Confirm the metric is scoped to the process you think it is.
- Distinguish "wasted but harmless today" from "a ticking time bomb." Work that's wasted but currently fits comfortably within capacity may not need urgent action; the same wasted work will become a real incident the moment traffic grows or the host loses spare capacity, so it's worth flagging even if latency looks fine right now.
Worked example
A service with 70% average CPU but P99 latency (the response time slower than 99% of requests, a common way to track worst-case rather than average experience) well within its SLA (service-level agreement, the target it's contractually or operationally expected to meet): profiling shows 40% of that CPU is spent in a background reconciliation job that runs every 30 seconds regardless of load, unrelated to request handling entirely. That explains the discrepancy directly: the CPU number reflects background work, not the request path the user experiences. The remediation is about the background job's efficiency and scheduling (should it run every 30s, or can it be event-driven instead), not about the request-handling code the on-call engineer might otherwise be tempted to optimize first.
Trade-offs and pitfalls
The main trap is treating "CPU is high" and "users are impacted" as the same finding when they can be completely decoupled, as in this example. Optimizing the wrong thing (request-handler code, when the real cost is a background job) wastes effort and leaves the actual capacity risk unaddressed. The remediation priority should follow from WHERE the profiler says the time goes, not from where it would be most convenient to look.
Unlock Full Question Bank
Get access to all 33 Debugging and Systematic Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.