User Research Planning and Fieldwork Questions
Designing and running a rigorous user or product research study end to end: formulating research objectives and hypotheses, choosing qualitative vs quantitative and mixed-method approaches for evaluating user needs and product experiences, and designing research instruments (screeners, discussion guides, surveys, usability tasks). Covers defining a sampling and recruitment strategy, screening and scheduling participants, and running fieldwork such as moderated or unmoderated usability sessions, diary studies, and interviews. Includes sample-size reasoning, avoiding method and recruitment bias (including inclusive and accessible recruiting), and trading research speed against rigor under real product timelines. The planning-and-execution discipline that determines whether user research findings are trustworthy enough to act on.
You need to learn about a user group that is hard to reach physically and has patchy connectivity, for example workers on remote offshore sites. How would you recruit them, and how would you adapt the study so the constraints do not quietly decide your findings?
Sample Answer
Direct answer
Reach a hard-to-reach, low-connectivity group through the organizations and local intermediaries who already have physical or trusted access to them, and adapt the study to run on the connection people actually have rather than the one you wish they had, so the study's results reflect their real needs, not just whichever subset happened to have a strong enough signal to take part.
Structured elaboration
Recruiting through intermediaries: for offshore workers, that means site operators, HSE (health, safety, and environment) staff, or unions who can grant access and identify willing participants across shifts and roles. For a similar low-bandwidth, emerging-market population, that means local community organizations, NGOs, or local partners who already have trust and can help with recruiting, translation, and framing the study appropriately. In both cases, go through the intermediary for access, but interview the actual worker or user directly, an intermediary is a door, not a proxy for the person you actually need to hear from.
Adapting the study to survive a bad connection: use short, offline-capable formats, store-and-forward tools that record locally and sync later, brief structured surveys of 5 to 8 questions rather than long ones, short voice notes or SMS-based prompts where even that is more reliable than an app, and keep any on-site sessions short and focused, since connectivity or shift schedules may cut them off with little warning.
Local norms and language: translate materials properly, not through an on-the-fly interpretation, and design task scenarios that make sense in the local context, a scenario written for one market's norms may simply not map onto another's.
Safeguards for a group with limited power to decline: when access runs through an employer, a site manager, or a local partner organization, there's a real risk that participation feels mandatory even when technically framed as voluntary, because saying no to a request that came through your boss carries more weight than saying no to a stranger. Address this directly: make clear, repeatedly, that participation and specific answers will not be shared with the employer or partner organization in an identifiable way, keep the employer or on-site manager out of the room during the actual interview, and give people a genuinely private way to decline or stop that doesn't require telling the person who granted access.
Not letting the constraints quietly decide the findings: log connectivity and device conditions for every submission, so "this person didn't answer that question" can be told apart from "this person's connection dropped there." Remember that anyone whose connection is bad enough to prevent participation at all is invisible to a remote-only study, which can make a low-connectivity group look more capable or satisfied than it really is; combine remote data collection with at least some in-person or asynchronous, low-bandwidth channel to catch that group too.
Worked example
For offshore rig workers, partner with the site's HSE manager to identify willing participants across day and night shifts, run short 30 to 45 minute on-site contextual interviews during downtime plus an offline diary tool that caches responses locally and syncs when connectivity returns. Tell each participant explicitly that individual answers will not be shared with their supervisor, and log device model and connection status alongside every diary entry so a missing entry can be told apart from a "nothing to report" entry.
Trade-offs and pitfalls
Relying entirely on remote or asynchronous tools systematically drops the very worst-connectivity people from the data, which understates the problem being measured. Skipping the "employer isn't in the room" safeguard is the fastest way to get answers that are polite rather than honest.
Your team is about to research a mental-health feature that may surface sensitive personal disclosures from participants. What would you put in place before sessions start to protect participants and prepare your moderators, and how would you handle a participant who becomes distressed mid-session?
Sample Answer
Direct answer
Before sessions start I'd put three things in place: real clinical input into the plan and the moderator training (not researchers improvising alone), a written, rehearsed escalation protocol so nobody has to invent a safety decision live, and consent language that's upfront about the topic and about the limits of confidentiality. Mid-session, if a participant becomes distressed, the moderator pauses the interview protocol, responds to the person first, follows the pre-agreed script, and never simply presses on with the guide.
Structured elaboration
Groundwork before recruiting: a licensed mental-health clinician helps design the risk-assessment approach and escalation criteria; an ethics or institutional review board (a panel that reviews research involving human participants for safety and consent adequacy) signs off; legal reviews the consent language and any mandatory-reporting obligations (situations where policy or law requires notifying someone regardless of the participant's wishes, for example imminent risk of harm, or certain disclosures involving minors).
Consent and framing: tell participants upfront, not buried in fine print, that questions may touch on mental health and that they can skip anything or stop entirely, and state plainly what happens if they disclose something indicating risk of harm to themselves or others, so there's no surprise later about the limits of confidentiality.
Phrasing sensitive questions to reduce harm and social-desirability bias: ask about specific recent behavior rather than abstract self-judgment, "in the last week, how many times did you open the app when you were feeling anxious" rather than "are you an anxious person," and normalize the topic before asking ("a lot of people using this feature tell us X, does any of that match your experience") so participants don't feel singled out. The same care applies to other sensitive domains: for financial hardship, ask about specific recent behavior ("have you had to delay a bill payment in the last month") rather than "are you bad with money."
Escalation path and outside resources: a written, rehearsed protocol naming exactly who on the team can authorize pausing or ending a session, a short list of region-appropriate crisis resources ready to hand over verbally, and a "warm handoff" option (staying connected while the participant reaches out to a resource, if they want that), rather than reading a phone number and disconnecting.
Storage, access, and anonymization for sensitive transcripts: paraphrase rather than transcribe a disclosure verbatim where a paraphrase preserves the insight, restrict access to the smallest possible group, encrypt at rest, keep contact and safety-follow-up information separate from research content, and set a shorter retention window for sensitive raw material than for routine usability recordings.
Minors, when relevant: a mental-health feature study may include participants under 18, which needs both guardian consent and the minor's own assent (agreeing in their own words, in age-appropriate language, separate from the guardian's signature). If approvals run behind a tight timeline, the honest options are running the adult-eligible sessions first and treating minors as a later wave, or shrinking scope, not skipping proper assent to hit a date.
Moderator preparation: train moderators specifically on the escalation script (rehearsed, not read cold under stress), trauma-informed interviewing basics (following the participant's pace, not pressing for detail), and a mandatory debrief after any session that escalated, so the moderator isn't left carrying it alone.
Worked example
If a participant becomes distressed or discloses something concerning mid-session, the moderator's first move is to stop the interview protocol and check in directly: "I want to pause, are you okay to continue, and is there anything you need right now." If the participant wants to stop, the session ends there and they're paid in full regardless. If what's disclosed meets the pre-agreed risk threshold, the moderator follows the written script: acknowledge what was shared, offer the region-specific resource verbally, and offer to stay connected while the participant reaches out to it. The moderator documents only what's needed for the safety record, not a full transcript of the disclosure, and immediately loops in the study's clinical lead per the escalation plan rather than deciding alone how serious it is.
Trade-offs & pitfalls
An escalation plan nobody has actually rehearsed tends to get improvised anyway under real stress. Consent language written purely to minimize legal liability can read as alarming and discourage honest participation, so it needs to stay clear and reassuring, not just airtight. And guardian consent and child assent are safeguards, not paperwork overhead, so they shouldn't flex under deadline pressure the way something like sample size might.
Your team wants to know whether a new feature in a complex product is easy to discover, and nobody agrees on what that means. How would you turn 'ease of discovery' into something you can actually measure, and what would you accept as good enough?
Sample Answer
Direct answer
Break "ease of discovery" into a small set of behavioral measures pulled straight from instrumentation, the logging that records what users actually do, paired with one small attitudinal or task-based measure that explains why those numbers look the way they do, and agree the success thresholds for both before the data lands, so nobody re-litigates what "good enough" means after seeing the result.
Structured elaboration
Behavioral measures from instrumentation, the "what happened": discovery rate, the share of eligible users who find or use the feature within a set number of sessions; time-to-first-use, how long it takes from being exposed to the feature to first successful use; and click-path length, how many steps or screens someone needs before reaching it. These are cheap to collect at scale, since they piggyback on events you're likely already logging or can add with a small engineering lift, but they can't explain why a number is low.
A small attitudinal or task-based measure, the "why": a first-click test or a brief moderated task with a handful of participants, asking them to find the feature without help, to see whether people even notice the entry point and what they expected to find there, plus maybe one short self-report question like "how easy was that to find" on a simple scale. This costs more per participant, but it explains the behavioral numbers, for example telling you whether a low discovery rate is caused by confusing labeling versus users genuinely not needing the feature.
Success thresholds agreed before the data lands: pick specific numbers with the team up front, for example "we'll call this a success if at least six of ten unmoderated first-click testers find the entry point without help, and the instrumented discovery rate clears an agreed floor within the first month." Stating these as numbers everyone commits to beforehand stops the bar from shifting depending on whether the actual result looks good or bad.
The trade-off between signal quality and instrumentation cost to engineering: more granular events, every hover, every scroll depth, every dead click, give a richer diagnostic picture, but cost real engineering time to build and maintain, and can add privacy or data-volume overhead. A cheaper, coarser plan, a handful of key events like exposed, clicked, and completed, gets directional signal fast and is usually enough to answer "is this bad enough to be worth investigating further," saving the more expensive fine-grained instrumentation for the cases that genuinely need root-causing.
Worked example
Take a bulk-export feature buried in a settings menu. Behavioral measure: of 10,000 users who opened settings in the last month, 300 also clicked bulk export at least once, a 3 percent discovery rate. Attitudinal or task-based check: 8 first-click tests asking people to find where they'd export all their data, and only 2 of the 8 find it without help, both after searching through multiple menus, a 25 percent unaided success rate. The team's pre-agreed threshold was to redesign the entry point if fewer than half of first-click testers succeed unaided, or if instrumented discovery is under a 10 percent floor after a month. Both numbers, 25 percent unaided success against a 50 percent bar, and 3 percent discovery against a 10 percent floor, miss the agreed bar clearly, so the team proceeds to redesign rather than debating whether 3 percent "feels" acceptable.
Trade-offs and pitfalls
Using only instrumentation, fast and cheap at scale, tells you that discovery is low but not why, leaving you unable to fix it with any confidence. Using only the qualitative check, rich but tiny, risks generalizing a handful of first-click testers to a user base they may not represent. And picking the "good enough" threshold after seeing the numbers quietly turns a measurement exercise into a negotiation, which is exactly what agreeing the bar in advance is meant to prevent.
What does triangulation mean in user research, and why would you spend part of a tight budget on a second method rather than more sessions of the first? Give an example where the second source changed the conclusion.
Sample Answer
Direct answer
Triangulation means deliberately checking a finding with a second, different method or independent data source, instead of trusting one method's result on its own. Spending part of a tight budget on a second method, rather than more sessions of the first, works because more sessions reduce sampling noise within a method, they don't remove that method's blind spot. Eight interviews and eighteen interviews both still carry self-report bias, meaning people describe their own behavior less accurately than they think; only a different kind of evidence, like usage data or a task-based test, can catch what interviews structurally can't see.
Structured elaboration: three ways to combine methods
| Design | Order | Answers | Product example |
|---|---|---|---|
| Sequential-explanatory | Quantitative first, then qualitative | "We saw the pattern, now what's causing it?" | Analytics show a spike in drop-off at account setup; follow-up interviews reveal the "company name" field rejects a character many small-business names contain |
| Exploratory-sequential | Qualitative first, then quantitative | "We think we found a need, how common is it?" | Interviews turn up freelancers manually re-entering the same client details across projects; a survey then measures what share of users do this weekly |
| Concurrent (both at once, compared afterward) | Run together, cross-check | "Do independent signals agree?" | A usability test's task success rate is run alongside a post-task rating; if most people finish the task but rate it frustrating, that gap is itself the finding |
Two concrete analytics-plus-qualitative pairings worth having ready:
- Metric: step-level drop-off rate in a funnel. Follow-up question, asked of someone who just left at that step: "Walk me through what you were trying to do right before you left this screen."
- Metric: unusually high variance in time-on-task, some users take three times longer than others. Follow-up question, in a moderated session: "Show me how you'd do this again, and think out loud while you do it," to locate exactly where the slow group gets stuck.
Worked example: where the second source changed the conclusion
Say a post-task satisfaction survey shows most users rate a checkout confirmation step "easy," and the team is about to deprioritize redesigning it. Session-level analytics for that step show a meaningful share of sessions include a back-and-forth pattern, users returning to the previous step before completing. A handful of follow-up interviews with people who showed that pattern reveal they weren't sure the payment had gone through and were re-checking, not struggling with the interface. The self-reported "easy" rating was true (the interface was fine) but incomplete; the second source, behavioral data plus targeted interviews, surfaced a trust problem the survey alone made invisible, and the conclusion changed from "leave it alone" to "add a clearer confirmation state."
Trade-offs and pitfalls
Triangulation costs more setup time and forces you to reconcile disagreement, which can slow a decision down when speed matters most. It's most worth the budget when the decision is expensive to get wrong, or when you already suspect the first method has a specific blind spot, self-report on sensitive topics being a common one. Don't triangulate reflexively on every low-stakes question; that spends budget without buying much extra confidence.
You are the first researcher at a company that has never run a moderated usability session. Walk me through how you would set up and run a remote study end to end, from the week before the first session to the day after the last one.
Sample Answer
Running a company's first moderated usability session end to end is really a sequence of small, boring reliability decisions made in advance, so that nothing on session day or in the write-up depends on memory. Here's the shape, week before to day after.
Two weeks before: define and set up
Write down the research questions, the tasks, and the success metrics (task completion, time on task, and optionally a short standard survey like the System Usability Scale, a 10-item questionnaire that produces a 0-100 usability score). Choose video conferencing software with screen-share and cloud recording, and confirm it supports getting explicit consent to record, which you'll ask for verbally at the start of every session, not just imply from a calendar invite. Assign roles: a moderator who runs the session and asks neutral prompts, a note-taker who timestamps what happens, and one or two silent observers, typically stakeholders, who stay muted, stay off camera if possible, and do not interject; brief them on this rule explicitly beforehand, since the instinct to jump in and explain a design decision is strong and it biases the participant.
One week before: recruit with buffer and finalize the script
Recruit more participants than you need, roughly 20 to 30 percent over your target, since no-shows and disqualifications are close to guaranteed on a first attempt. Build 10 to 15 minutes of buffer between sessions for notes and a quick tech reset. Write the moderator script in order: intro and purpose, verbal consent to record, warm-up questions to build rapport, an explicit think-aloud reminder ("as you go, please say out loud what you're looking at and thinking, even if it seems obvious"), the task prompts themselves, in-task probes, and a wrap-up debrief. Run one internal pilot session against the script to catch timing problems and awkward wording before a real participant sees it.
During each session: the moment nobody plans for
Decide in advance what the moderator says when a participant gets stuck or asks for help, since this is where new moderators improvise badly. The rule: don't rescue them, and don't let them sit in silence indefinitely either. A concrete line: "Take your best guess, there's no wrong answer here." If they're still stuck after roughly 60 to 90 seconds, a second line: "Let's move on to the next task, and I'll note this one as a blocker," logging the stuck point as a finding rather than treating it as a failed session. For a technical failure mid-session (dropped screen-share, frozen video), the fallback is dropping to audio-only and asking the participant to narrate what they're doing, or rescheduling the remaining tasks if the connection doesn't recover within a couple of minutes.
After each session: close the loop before memory does
Confirm the recording actually saved before the participant disconnects. File the note-taker's timestamped notes the same day. Have the moderator and note-taker do a five-minute debrief immediately after while the session is fresh, since the exact wording of a confusing moment fades within hours, not days.
The day after the last session
Make sure what exists is complete before anyone starts drawing conclusions from it: the recording, the timestamped notes, and any survey responses for every session, filed in one place. That completeness check, not the analysis itself, is the last step of fieldwork; what you do with the material afterward is a separate job.
Trade-offs
Over-recruiting too aggressively wastes budget on no-shows that didn't materialize; under-recruiting risks ending up with too few usable sessions to say anything with confidence. An observer who breaks the "muted, no interjecting" rule even once can visibly change how candid the participant is for the rest of that session.
Unlock Full Question Bank
Get access to all User Research Planning and Fieldwork interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.