User Research Planning and Fieldwork Questions
Designing and running a rigorous user or product research study end to end: formulating research objectives and hypotheses, choosing qualitative vs quantitative and mixed-method approaches for evaluating user needs and product experiences, and designing research instruments (screeners, discussion guides, surveys, usability tasks). Covers defining a sampling and recruitment strategy, screening and scheduling participants, and running fieldwork such as moderated or unmoderated usability sessions, diary studies, and interviews. Includes sample-size reasoning, avoiding method and recruitment bias (including inclusive and accessible recruiting), and trading research speed against rigor under real product timelines. The planning-and-execution discipline that determines whether user research findings are trustworthy enough to act on.
You need to run a contextual inquiry to see how people really use a product in their own environment, but this round has to happen over video calls instead of in person. What do you lose by not being physically there, and how would you adapt your recruitment, task design, and observation approach to still capture genuine context of use?
Sample Answer
Direct answer
The biggest loss going remote is everything outside the camera's frame: you can't wander the space yourself, notice the sticky note on the monitor nobody mentioned, or pick up on ambient cues (what's playing on the TV, who else is in the room) the way you would standing there in person. The fix is to treat the participant's own camera as a deliberately steered window rather than a passive replacement for your eyes, and design recruitment, tasks, and observation around that constraint instead of pretending it isn't there.
Recruitment adaptations
Screen for participants who can actually show you their environment: a phone with a working camera and enough mobile data or wifi to hold a video call while walking around, not just a desktop webcam pointed at a face. Ask upfront whether they're comfortable panning their camera around their space, since this is a bigger ask than a normal video call and some people will decline or need reassurance about what's actually being recorded.
Task design adaptations
Break the session into short segments rather than one long continuous task, since a remote participant physically has to move and re-aim a device between steps, which a moderator in the room would never have to ask for. Open with a two to three minute guided "space tour" using their phone camera before any task starts, so you see the room, their setup, and any artifacts (the notebook, the second device, the sticky notes) before you're mid-task and it's awkward to stop and ask them to show you something.
Observation adaptations
Ask the participant to narrate what's just outside the frame ("what's on the desk next to your laptop right now?") rather than assuming you'll catch it visually, since you won't. Request photos of their setup ahead of the call as a backup, in case the live camera work is clumsy or the connection drops. Treat pauses, background noise changes, and mutes as signals worth asking about directly, since they might mean the participant moved, got interrupted, or is hiding something they're embarrassed about, information an in-person observer would have caught without asking.
Worked example
For a mobile-banking contextual inquiry, I would ask the participant to have their wallet, a recent paper statement, and the phone they normally use to check their balance within reach before the call starts. I'd open with a two-minute phone-camera tour of wherever they normally do this task (kitchen table, couch, desk), then have them walk through checking their balance while thinking aloud, periodically asking them to angle the camera to show their hands and immediate surroundings, not just the app screen, since the screen alone tells you what they tapped but not what else was competing for their attention in that moment.
Trade-offs & pitfalls
Remote contextual inquiry skews toward participants who are comfortable being on camera and technically capable of managing a video call while also completing a task, which is a form of selection bias you should name rather than ignore. It also raises social desirability (people behave more self-consciously when they are the one holding the camera pointed at themselves than when an observer is quietly present in the corner of a room). Mitigate both by offering a camera-off option for parts of the session where narration alone is enough, collecting pre-session photos so less depends on live camera work, and where the stakes justify it, running a small number of in-person sessions specifically to sanity-check that the remote sessions aren't missing something systematic.
An online satisfaction study has already reported its results, and you notice afterwards that the people who responded were overwhelmingly the product's heaviest users. How far off are those numbers likely to be, what can still be salvaged from them, and how do you keep the same thing from happening on the next one?
Sample Answer
Direct answer
When respondents are overwhelmingly the product's heaviest users, the results are very likely biased toward more positive satisfaction than the true user base would report on average, since heavy usage and satisfaction tend to correlate: people who dislike a product tend to use it less and are also less likely to bother filling out a survey about it. I'd treat the headline satisfaction number as an upper bound on the true average, not a representative estimate. The exact size of the gap isn't knowable without more data, but real value can still be salvaged by comparing respondents against the full population on usage data already on hand, and by scoping what heavy users say honestly rather than generalizing it.
Why the direction of the bias is predictable even when the size isn't
Non-response to a voluntary survey is rarely random: people with a strong, often positive, relationship to a product are more willing to spend a few minutes rating it, while people who are lukewarm, confused, or actively dissatisfied are both less likely to still be using the product heavily and less likely to bother responding at all. That makes the reported score most likely inflated relative to the true population average, though without more data the exact number of points can't be pinned down.
What can still be salvaged
Compare the respondent sample against the full user base on whatever usage data is already available, login frequency, feature adoption, tenure, which tells concretely how skewed the sample is. The open-ended comments from heavy users are still real and useful, just scoped honestly as "what our most engaged users think," which is legitimate input for retention-focused decisions, but shouldn't be generalized to "what our users think" broadly, especially not for decisions aimed at casual or at-risk users. If there are any responses at all from lighter users, even a small number, I'd read those specifically and separately rather than averaging them into the overall score, since they're the closest thing available to the missing signal.
A statistical option if a single corrected number is needed
If the true population's split across usage bands is known (say, what share of all users are heavy versus light), the survey responses can be reweighted so each band contributes in proportion to its real share of the population rather than its share of respondents (reweighting means mathematically adjusting how much each response counts toward the average). This pulls the corrected average back toward what a representative sample would likely show, but only works if there's at least some response in each band to reweight, an average with zero light-user responses can't be rescued this way since there's nothing there to upweight.
Keeping it from happening again
Set a target quota by usage band before sending the survey, and recruit or oversample specifically to fill the under-represented bands, rather than sending one blast email and hoping responses land proportionally. Track response rate by usage band as the survey runs so a skew is visible in real time rather than discovered afterward. Consider incentivizing or specifically targeting lighter users, who have the least built-in motivation to respond to a survey about a product they use less.
Worked example
Suppose the survey drew 200 responses, and internal usage data shows those 200 people log in on average 5 times a week, while the full active user base averages 1.5 logins a week. That's direct, checkable evidence the sample skews heavy, and it tells concretely which direction, and roughly how far, the reported satisfaction score likely overstates the true average, even without knowing the exact correction.
Trade-offs and pitfalls
A common wrong turn is quietly treating the number as fine because "it's still directionally useful," without ever stating the skew out loud to whoever is using the result, which lets an inflated number quietly drive a decision meant for the average user. Reweighting is also not a magic fix: it can only correct for the usage-band imbalance actually measured, and it does nothing for other unmeasured differences between respondents and non-respondents, for example if dissatisfied heavy users also disproportionately skip the survey, reweighting by usage band alone won't catch that.
Walk me through the research methods you reach for during discovery. For each one, what question does it answer well, roughly how many participants does it need, and where does it mislead you?
Sample Answer
Direct answer
During discovery I reach for methods that answer "what is actually going on and why," not "did this specific design work." My default toolkit is user interviews, contextual inquiry (watching people work in their real environment), diary studies for behavior that unfolds over days, and a review of any existing analytics. I hold off on usability testing and controlled experiments (A/B tests) until there is a concrete design to react to, because those are validation-phase tools, not discovery tools.
Structured elaboration
For each method: what it answers well, roughly how many participants it needs, the artifact it produces, and where it sits in the product lifecycle.
| Method | Answers well | Participants (rough) | Artifact produced | Lifecycle |
|---|---|---|---|---|
| User interviews | Motivations, pain points, mental models (how someone expects a thing to work) | 5 to 8 per user segment; teams stop when new interviews stop surfacing new themes, called saturation | Synthesis memo: recurring themes with illustrative quotes | Exploratory (discovery); a few extra sessions are common during validation to explain a confusing result |
| Contextual inquiry | Real workflow and environment constraints someone won't think to mention unprompted | 4 to 6 site visits, since each is longer and costlier than an interview | Annotated field notes plus a workflow diagram | Exploratory only |
| Diary study | Behavior that unfolds over days, not in one sitting (how someone actually triages messages over a week) | 8 to 12 recruited, budgeting for a few dropping out mid-study | Day-by-day log per participant | Exploratory, occasionally repeated post-launch to check adoption over time |
| Exploratory survey | How common a pattern is, once interviews have named it | 20 to 30 for reading open text for new themes; 100+ once estimating a percentage with any confidence | Top-line stats plus open-text excerpts | Straddles both: early to size a hypothesis, later to measure whether an intervention moved a stated attitude |
| Analytics review | What people actually did, at scale, with no self-report distortion | Uses existing usage data, not a recruited sample | Funnel or dashboard view | Continuous, and central to validation once something has shipped |
Usability testing and A/B tests belong later. Usability testing needs a prototype to react to, and A/B tests need a live variant to route traffic to; both tell you whether a specific solution works, not what the underlying problem is.
Worked example
Say support tickets and a hunch suggest new users don't understand a permissions screen. I'd run 6 interviews first, a plain-language walkthrough of their mental model, which surfaces a specific confusion about what "admin" means to them. I'd confirm the pattern isn't a fluke with a quick 25-person exploratory survey asking users to define "admin" in their own words, then size how common the confusion is before recommending a redesign.
Trade-offs and pitfalls
Interviews and contextual inquiry both suffer from self-report bias (what people say does not always match what they do), and contextual inquiry adds an observer effect, meaning people behave differently while watched. Diary studies trade a small sample for depth over time, so don't expect statistical confidence from them. Analytics tells you what happened but never why, so treat a surprising number as a hypothesis to interview around, not a finished conclusion. The most common beginner mistake is reaching for a usability test or A/B test during discovery, before there is anything concrete to test.
Your team proposes using customer support logs and product telemetry as the primary source of generative insight. Critically evaluate that proposal: what does it genuinely get you, what would worry you about it, and how would you cover the questions those sources cannot answer?
Sample Answer
Direct answer
Support logs and product telemetry, the passive, automatically-recorded data on what users click, complete, or abandon, are genuinely useful as a cheap, high-volume starting point for generating hypotheses. What worries me is treating them as sufficient on their own: they systematically miss certain users, the record itself can quietly corrupt over time, and behavior without a stated reason can't support a claim about motivation or cause. I'd treat them as a first pass and layer a small, targeted primary study on top before making a real decision.
Structured elaboration
Who is systematically missing: users who churned quietly before generating much telemetry, survivorship bias, meaning the data only reflects people who stuck around long enough to be recorded; users on older devices or app versions where tracking was never updated; anyone outside the primary supported locale if instrumentation was built English-first; offline or low-connectivity users whose events never sync; and, in support logs specifically, the large group who got frustrated and left without ever filing a ticket, versus the smaller, more vocal group who did.
How the record itself corrupts: instrumentation gaps happen when a team ships a feature and forgets to add or fix the tracking for it, so real usage exists with no record of it at all. Event-definition drift happens when engineers quietly redefine what an event means, say "signup complete" starts firing at a different step after a refactor, without updating a shared definition, so a chart that looks continuous is actually comparing two different things across the point where the definition changed.
Why behavior alone can't carry a motivational or causal claim: a click tells you what happened, not why, and it can't distinguish "this confused them" from "they intended exactly this and it worked." Attributing a cause to a raw behavioral pattern, "users abandon here because the button is unclear," is a guess dressed up as a finding until someone actually asks the people who left.
Worked example: layering a primary study on top
Telemetry shows a spike in a specific error-adjacent event for a segment of users. Rather than concluding what caused it, I'd trigger a short in-app micro-survey right after that event fires, catching intent at the moment it happens instead of days later from memory, then follow up with 5 to 6 contextual interviews or a reproduction of the flow in a moderated usability session with people who match that segment, specifically to test the leading hypothesis the telemetry suggested.
Trade-offs and pitfalls
The augmentation adds time and cost the raw data doesn't, which is exactly why teams are tempted to skip it under deadline pressure, and exactly when the biggest wrong-conclusion risk shows up. The instrumentation-drift risk is easy to miss because the chart still looks clean; it's worth explicitly asking an engineer whether an event's definition has changed before trusting a trend line that spans a big code change.
Tell me about a time you convinced a stakeholder to accept a slower or more rigorous piece of research than they wanted. How did you make the case, and how did it turn out?
Sample Answer
Direct answer
In a past role I pushed for two extra weeks before a rushed launch, and that extra time is what let us catch a real usability problem before it reached everyone, rather than shipping a metric win alongside a support cost we hadn't budgeted for.
Situation
Three weeks before a marketing campaign slot, leadership wanted to ship a redesigned onboarding flow fast enough to catch the promotional window. My qualitative research (a round of customer interviews) and the instrumentation needed for a proper controlled test were both still incomplete.
Task
I needed to decide whether to let the rushed timeline stand or make the case for more time, knowing the pushback would come from marketing and from the executive who owned the launch date.
Action
Before I went to the executive sponsor, I aligned the two partners whose evidence would carry more weight than mine alone: the engineering lead, who confirmed an unresolved instrumentation gap meant we'd be flying blind on a key drop-off point, and the support lead, who flagged which ticket categories were most likely to spike if we shipped the current confusing step unchanged. With that evidence assembled, including partial interview notes pointing at a specific unresolved point of confusion and the instrumentation gap itself, I put together a short risk-versus-cost comparison: the two-week delay costs a marketing slot; shipping now on an untested assumption risks a common workflow breaking and driving up support volume. I proposed a middle path, two more weeks to finish the outstanding interviews and instrumentation, then a release behind a feature flag with a small initial rollout so we could catch problems before everyone saw them.
Result
The executive accepted the compromise once engineering and support were already visibly on board, not because I was still the lone voice asking for delay. The extra research surfaced one real point of confusion that we fixed before the wider release. Activation ended up landing solidly ahead of where the rushed version was projected, and we didn't see the support-ticket spike the rushed version had risked.
Looking back
I'd build a research-versus-launch tradeoff conversation into the roadmap earlier, so we're not negotiating time at the last minute, and I'd get engineering's and support's read on risk before the crunch rather than during it, since that's what actually moved the decision in the end, not my own argument alone.
Unlock Full Question Bank
Get access to all User Research Planning and Fieldwork interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.