Debugging and Systematic Troubleshooting Questions
Diagnosing defects methodically: reproducing failures, forming and testing hypotheses, reading stack traces and logs, bisecting changes, and reasoning about error handling and edge cases. Covers a disciplined root-cause approach that applies from local bugs to production issues, distinct from embedded hardware-level debugging. A universally probed engineering-craft skill.
You receive this Android crash stack trace from a user's device:
FATAL EXCEPTION: main
Process: com.example.app, PID: 12345
java.lang.NullPointerException: Attempt to invoke virtual method 'void android.widget.TextView.setText(java.lang.CharSequence)' on a null object reference
at com.example.app.ui.MainActivity.onCreate(MainActivity.java:45)
at android.app.Activity.performCreate(Activity.java:8000)
at android.app.ActivityThread.performLaunchActivity(ActivityThread.java:3000)
Identify the most likely root cause, what you would check in the source at that file and line, and two possible fixes.
Sample Answer
Direct answer
The trace shows a NullPointerException at MainActivity.onCreate(MainActivity.java:45) calling setText on a TextView that is null. The most likely root cause: findViewById for that TextView returned null, almost always because setContentView(...) was called with the wrong layout resource (or not called before the findViewById), or the ID referenced doesn't exist in the layout that actually got inflated (a common cause when the app has multiple layout variants, e.g., separate layout and layout-land XML files, and the ID was added to only one).
Structured elaboration
- Check line 45 of
MainActivity.javadirectly to see which view is being set and whichfindViewByIdcall produced it. - Check the order of
setContentViewandfindViewById. IffindViewByIdruns beforesetContentView, or targets a different content view than the one shown, it will returnnullfor every view ID, since there's no inflated view tree yet to search. - Check for layout-variant mismatches. If the app has alternate layout resources for different screen sizes, orientations, or API levels (
res/layout-land,res/layout-sw600dp, etc.), confirm the view ID referenced actually exists in ALL variants that could be inflated at runtime, not just the one the developer was looking at while writing the code. This is one of the most common real-world causes of this exact crash pattern, since it only manifests on devices/configurations that load the variant missing the ID. - Check for a recent layout XML change that renamed or removed the view ID without updating the corresponding
findViewByIdcall, which is a very common source of this bug after a UI refactor. - Confirm the reproduction path. Since this came from a specific user's device (a crash report, not a local repro), check whether the crash correlates with a specific device model, OS version, or locale, any of which could imply a specific layout variant or resource configuration being loaded.
Two possible fixes, addressing the two most likely root causes:
- If
setContentViewwas called with the wrong layout resource: fix the argument to point at the correct layout that actually contains the referenced view ID. - If the view ID genuinely doesn't exist in a particular layout variant (e.g., a landscape-only layout is missing a view present in portrait): either add the missing view to that variant, or guard the code with a null check before calling
setText, if the view is legitimately optional in some layouts.
Worked example
A crash report shows this NPE occurring only on tablet devices; checking the project's resource folders reveals a layout-sw600dp (tablet-specific) variant of the activity's layout that was created for a redesign but never had the TextView with that specific ID added to it, unlike the phone layout. The fix: add the missing view to the tablet layout (or, if it's intentionally absent there, guard the setText call with a null check specific to that code path).
Trade-offs and pitfalls
A defensive null-check around the setText call would stop the crash but silently hides a genuine layout inconsistency, which is the wrong fix if the view was SUPPOSED to exist in every layout variant: the app now fails silently (missing UI element, no visible error) instead of crashing loudly, which is often harder to notice and diagnose later. The null check is the right fix only when the view is deliberately optional in some configurations; otherwise, fix the actual layout/inflation mismatch.
Explain rubber-duck debugging and describe how you would use it collaboratively to help a teammate find a bug they are stuck on. Include when it is appropriate to escalate to active pair programming versus continuing to guide them through the technique.
Sample Answer
Direct answer
Rubber-duck debugging is the practice of explaining your code or problem, line by line, out loud, to an inanimate object (or any passive listener), on the theory that the act of ARTICULATING your assumptions forces you to notice the gap between what you believe the code does and what it actually does; used collaboratively, you become the "duck," a real listener whose job is to stay quiet and let the other person talk through it, only interjecting with clarifying questions, not answers.
Structured elaboration
Why it works: most debugging time is lost not to hard problems but to an unexamined assumption ("of course X is true here") that the explainer has never actually had to state out loud; the discipline of narrating each step forces exactly that statement, and the mismatch often becomes obvious the moment it's spoken, even before the listener says anything at all.
How to use it collaboratively to help a stuck teammate:
- Sit with them and ask them to explain the problem from the beginning, including what they expect to happen and what's actually happening, as if you know nothing about the code (even if you do); resist the urge to jump in with your own theory immediately, since the goal is for THEM to find the gap through their own narration, not for you to hand them the answer.
- Ask clarifying, not leading, questions when something seems glossed over: "what does this variable contain at this point?" rather than "isn't this variable actually null here?"; a clarifying question preserves the self-discovery effect, while a leading question short-circuits it, in effect just answering the question yourself, in a slightly slower way.
- Notice when they pause or hesitate on a specific line, since that hesitation is itself a signal, often the exact spot where their explanation doesn't fully hold together, even if THEY haven't consciously noticed it yet; gently returning to that point ("can you say more about what happens right there?") is more effective than moving on.
- Let silence do some of the work. Resisting the urge to fill every pause with your own guess gives them room to keep narrating and often arrive at the insight themselves, which is both faster (no back-and-forth debugging your OWN, potentially wrong, theory) and more valuable for their own learning than being handed the answer.
When to escalate to active pair programming instead of continuing to guide:
- When the explanation reveals a gap in UNDERSTANDING (not just an overlooked line) that narration alone won't resolve, e.g., a genuine misunderstanding of how a language feature or library behaves, where providing the missing knowledge is more useful than more questions.
- When you notice, through their narration, something concrete they clearly haven't seen (a specific line, a specific value) and pointing it out directly is faster and kinder than continuing an extended Socratic process once the value of self-discovery has been exhausted for this specific bug.
- When time pressure genuinely doesn't allow for the (often somewhat slower) self-discovery process, in which case switching to active pairing, working the problem together directly, is the pragmatic choice, with an explicit acknowledgment that you're switching modes for a good reason, not simply losing patience.
Worked example
A teammate stuck on why a function returns stale data: walking them through explaining the code line by line, they narrate "and then this checks the cache... and returns it if it's still valid... I wanted a sixty-second expiry, so I set the value to sixty thousand..." and pause, mid-sentence, on "sixty thousand," suddenly noticing they'd been assuming the codebase's convention was MILLISECONDS (so they multiplied their intended sixty seconds by 1000 before storing it), when the surrounding code that actually reads this field treats it as SECONDS directly, with no conversion; storing 60000 where the code expects 60 means the real expiry is 60000 seconds (about 16.7 hours), roughly 1000x longer than the sixty seconds they intended, which is exactly why the data was going stale. They found this entirely through their own narration; the only input needed was staying quiet and letting them keep talking through it.
Trade-offs and pitfalls
Rubber-duck debugging (with a real, silent listener) works best for a SPECIFIC kind of stuck-ness, an unexamined assumption, not for a genuine knowledge gap or a problem requiring information the person simply doesn't have; recognizing which kind of "stuck" you're looking at, and switching to direct guidance or active pairing when it's the latter, is what keeps the technique from wasting time on a problem it was never suited to solve.
Explain how you would diagnose an Android ANR (Application Not Responding). List the tools and commands you would use, what artifacts you would collect such as traces and stack dumps, and how to interpret a main-thread stack trace to find the blocked operation.
Sample Answer
Direct answer
An ANR (Application Not Responding) fires when the Android system detects the main/UI thread has been blocked too long, typically around 5 seconds for input dispatch (a foreground BroadcastReceiver.onReceive() gets roughly the same 5-second order of magnitude, not a shorter one; a background broadcast gets substantially longer). Diagnosing it means finding what was blocking the main thread at the moment the ANR was triggered, which requires capturing a main-thread stack trace from that exact moment, either from an ANR trace file the system writes, or from a live profiling session if you can reproduce it.
Structured elaboration
- Collect the ANR trace. Android writes a trace file (historically
/data/anr/traces.txt, or accessible viaadb shell dumpsys/ the device's bug report on modern versions) capturing the stack of every thread at the moment of the ANR, most importantly the main thread. Google Play Console also surfaces ANR reports with stack traces for production crashes you can't reproduce locally. - Read the main thread's stack trace specifically. The top frames show exactly what the main thread was doing when it got stuck: a synchronous network call, a large synchronous disk read, a lock wait, or a long-running computation left on the UI thread instead of a background thread.
- Use
adb shell dumpsys activityandAndroid Profiler(CPU profiler, specifically the "trace system calls" or method-tracing mode) to reproduce and observe main-thread activity live if you can trigger the ANR on demand, giving you a timeline rather than just a snapshot. - Check for common ANR causes directly, since certain patterns account for the large majority of real-world ANRs: a network or database call made synchronously on the main thread, a
BroadcastReceiver.onReceive()doing heavy work (its foreground timeout is roughly the same order of magnitude as input dispatch, about 5 seconds per Android's official ANR documentation, not shorter; background broadcasts get a substantially longer allowance, so this cause is just as urgent to rule out as a slow synchronous call), a deadlock between the main thread and a background thread both waiting on each other's locks, or aBindercall to another process that hangs. - Interpret a blocked-on-lock stack correctly. If the main thread's trace shows it waiting on a lock (
BLOCKEDstate in the trace, waiting to enter asynchronizedblock), find which OTHER thread currently holds that lock and what IT is doing; the actual root cause is often in that second thread, not visible directly on the main thread's own stack.
Worked example
An ANR trace shows the main thread's top frame inside SQLiteDatabase.query, called synchronously from onResume(). That's a direct, unambiguous cause: a database query is running on the UI thread, and if the query happens to take longer than usual (a missing index, database growth over time, a device under memory pressure causing disk I/O contention), it now exceeds the ANR threshold. The fix: move the query to a background thread or coroutine, and update the UI once results return, rather than blocking the thread responsible for keeping the app responsive to input.
Trade-offs and pitfalls
The main trap is treating "what was on the main thread's stack" as automatically the fix target, when the trace shows a lock wait: in that case the visible frame is a symptom, and the actual work causing the delay is happening on a different thread entirely, requiring you to read that thread's trace too. A second trap: ANRs that only happen on specific devices or under memory pressure can be very hard to reproduce locally on a fast development device, making production ANR reports (with their bundled stack traces) the primary and sometimes only usable evidence.
Your CI build fails only on the pipeline and never locally. Outline a methodical plan to find the difference: comparing environment variables, container images, OS versions, dependency versions, and filesystem semantics, and describe how you would produce a minimal reproducible example inside CI itself.
Sample Answer
Direct answer
When a CI build fails only in the pipeline, the fastest path is to systematically diff the CI environment against your local one along every axis that could plausibly differ, rather than guessing: environment variables, container/base image, OS and library versions, filesystem case-sensitivity and permissions, and the exact command CI runs versus the one you run locally. The end goal is a minimal reproduction that runs inside CI itself, because "works when I add print statements" isn't a fix, it's a workaround.
Structured elaboration
- Compare environment variables. Dump the full environment CI uses (most CI systems let you print it, or add a debug step that runs
env) and diff it against your local shell. Missing or unexpectedly-set variables (aNODE_ENV, a locale, a timezone, a feature flag) are a common silent cause. - Compare the base image or container. If CI runs in a container and you develop on bare metal (or a different container), pin down exact OS version, installed system libraries, and default locale/timezone, since these can silently change behavior (date parsing, sort order, floating-point rounding modes).
- Compare dependency versions exactly, not just "the same major version." A CI pipeline that does a clean install from a lockfile can resolve a transitive dependency (a dependency of one of your dependencies: something your package manager pulled in automatically, not something you installed directly) differently than a long-lived local
node_modulesorvenvthat never got a clean reinstall. - Compare filesystem semantics. CI runners are frequently Linux (case-sensitive filesystem) while local development happens on macOS (case-insensitive by default); an import or file reference with inconsistent casing works locally and fails only in CI.
- Reproduce inside CI, not around it. Add a debug step to the CI job itself, before the failing step, that dumps the environment, tool versions, and installed package list, so you're comparing CI's actual state rather than guessing from documentation. If possible, get an interactive shell into the exact CI container (many CI providers support this) and run the failing command by hand.
- Build a minimal reproducible example inside CI: strip the pipeline down to the smallest set of steps that still reproduces the failure, in a scratch branch if needed, so you have a fast iteration loop instead of waiting on the full pipeline each time.
Worked example
A test suite that passes locally on macOS but fails in CI (Linux) with a file-not-found error: dumping the CI environment shows the import path from Utils import helper, while the actual file on disk is utils.py. macOS's case-insensitive filesystem silently accepted the mismatch locally; Linux's case-sensitive filesystem in CI did not. The minimal repro: a two-line script with that exact import, run in a Linux container locally, reproduces it in seconds without needing the full CI pipeline.
Trade-offs and pitfalls
The trap is iterating by pushing small changes and waiting for the full CI pipeline to re-run, which can cost many minutes per guess. Getting a fast, local reproduction of the CI environment (even a bare Docker container matching CI's base image) pays for itself after two or three iterations. A second common trap: fixing the symptom by disabling the failing step (skip, retry, quarantine) without ever finding the actual environment difference, which resurfaces the same class of bug on the next similarly-shaped test.
Describe a systematic, repeatable approach you use to troubleshoot an unfamiliar technical problem end to end. Cover how you observe the symptom, form and prioritize hypotheses, gather and interpret evidence such as logs, metrics, and traces, isolate the root cause, implement and validate a fix, and decide when to escalate, roll back, or write up a postmortem.
Sample Answer
Direct answer
Troubleshooting an unfamiliar problem is a loop, not a single step: observe the symptom precisely, form a small set of testable hypotheses ranked by likelihood and cost to check, gather evidence that discriminates between them, isolate the true cause, implement and validate a fix, and decide whether the incident needs a rollback, an escalation, or a written postmortem. The loop repeats: each piece of evidence should narrow the hypothesis set, not just confirm what you already believed.
Structured elaboration
- Observe the symptom precisely. Write down exactly what is wrong, in falsifiable terms: not "the API is slow" but "p99 latency (the response time slower than 99% of requests, i.e. how bad the worst cases are, not just the average) on
POST /ordersrose from 80ms to 900ms starting at 14:32 UTC, affecting roughly 3% of requests." Vague symptoms produce vague hypotheses. - Form hypotheses before you start digging. List the plausible causes given what changed recently (deploys, config, traffic pattern, dependency versions) and what the symptom rules out. A hypothesis you cannot state is a hypothesis you cannot test.
- Prioritize by expected information gain divided by cost. A five-minute log grep that could confirm or kill three hypotheses at once beats a one-hour deep profiling session that only speaks to one.
- Gather evidence that discriminates. Logs tell you what happened at a point; metrics tell you the shape of the problem over time; traces tell you where time went inside one request. Pick the instrument that actually distinguishes your live hypotheses, not the one you're most comfortable with.
- Isolate the root cause, not just a correlated symptom. A dropped hypothesis should be dropped because evidence contradicts it, not because you got bored of it.
- Implement and validate the fix against the same evidence that revealed the problem. If you diagnosed via a specific metric, watch that metric recover before declaring victory.
- Decide what happens next. If customer impact is ongoing and the fix is unproven, roll back first and diagnose second. If a similar failure could recur, or the incident had real impact, write it up so the org doesn't relearn the same lesson.
Worked example
A "the checkout page is slow" report, applied through the loop: symptom precisely stated as "median load time is normal, but a subset of loads takes 8-12 seconds, starting after this morning's deploy." Hypotheses: (a) the new deploy added a blocking call, (b) a downstream dependency degraded independently, (c) the slow subset shares a common attribute (e.g., a specific region or a large cart). A single log query grouping slow requests by attribute would discriminate between (c) and the other two in minutes, before touching a profiler. Suppose it shows the slow requests all hit a newly added inventory-check call to a dependency with no timeout: that both confirms (a) and rules out (b)/(c) as primary causes. Fix: add a timeout and a fallback; validate by watching the p99 metric drop back to baseline over the next hour, not just the fix compiling.
Trade-offs and pitfalls
The biggest failure mode is skipping hypothesis formation and going straight to your favorite tool (attaching a profiler because you're comfortable with it, even when a five-minute log check would have ruled out two hypotheses first). The second is treating the first correlated signal as the cause without checking whether it's actually causal. Under time pressure it's tempting to fix the first plausible thing you see; that's fine as a mitigation, but the loop isn't complete until you've confirmed the metric recovered and understood why, or you will be back debugging the same symptom next week.
Unlock Full Question Bank
Get access to all 12 Debugging and Systematic Troubleshooting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.