User Research Planning and Fieldwork Questions
Designing and running a rigorous user or product research study end to end: formulating research objectives and hypotheses, choosing qualitative vs quantitative and mixed-method approaches for evaluating user needs and product experiences, and designing research instruments (screeners, discussion guides, surveys, usability tasks). Covers defining a sampling and recruitment strategy, screening and scheduling participants, and running fieldwork such as moderated or unmoderated usability sessions, diary studies, and interviews. Includes sample-size reasoning, avoiding method and recruitment bias (including inclusive and accessible recruiting), and trading research speed against rigor under real product timelines. The planning-and-execution discipline that determines whether user research findings are trustworthy enough to act on.
Tell me about a time you convinced a stakeholder to accept a slower or more rigorous piece of research than they wanted. How did you make the case, and how did it turn out?
Sample Answer
Direct answer
In a past role I pushed for two extra weeks before a rushed launch, and that extra time is what let us catch a real usability problem before it reached everyone, rather than shipping a metric win alongside a support cost we hadn't budgeted for.
Situation
Three weeks before a marketing campaign slot, leadership wanted to ship a redesigned onboarding flow fast enough to catch the promotional window. My qualitative research (a round of customer interviews) and the instrumentation needed for a proper controlled test were both still incomplete.
Task
I needed to decide whether to let the rushed timeline stand or make the case for more time, knowing the pushback would come from marketing and from the executive who owned the launch date.
Action
Before I went to the executive sponsor, I aligned the two partners whose evidence would carry more weight than mine alone: the engineering lead, who confirmed an unresolved instrumentation gap meant we'd be flying blind on a key drop-off point, and the support lead, who flagged which ticket categories were most likely to spike if we shipped the current confusing step unchanged. With that evidence assembled, including partial interview notes pointing at a specific unresolved point of confusion and the instrumentation gap itself, I put together a short risk-versus-cost comparison: the two-week delay costs a marketing slot; shipping now on an untested assumption risks a common workflow breaking and driving up support volume. I proposed a middle path, two more weeks to finish the outstanding interviews and instrumentation, then a release behind a feature flag with a small initial rollout so we could catch problems before everyone saw them.
Result
The executive accepted the compromise once engineering and support were already visibly on board, not because I was still the lone voice asking for delay. The extra research surfaced one real point of confusion that we fixed before the wider release. Activation ended up landing solidly ahead of where the rushed version was projected, and we didn't see the support-ticket spike the rushed version had risked.
Looking back
I'd build a research-versus-launch tradeoff conversation into the roadmap earlier, so we're not negotiating time at the last minute, and I'd get engineering's and support's read on risk before the crunch rather than during it, since that's what actually moved the decision in the end, not my own argument alone.
What goes into the research plan you put in front of stakeholders before discovery starts, and how do you keep it short enough that they read it and specific enough that they can approve it?
Sample Answer
Direct answer
A pre-discovery research plan needs six things on it: the decision it unlocks, the research questions, the method and sample size, the timeline, the success criteria, and who has to sign off. Keep it to about one page by cutting anything a stakeholder doesn't need in order to say yes, and keep it specific by writing every section as something they can approve or reject, not a vague intention.
Structured elaboration
What goes in:
- The decision this research unlocks (what will we do differently depending on the answer)
- Research questions, split into primary (must be answered) and secondary (nice to have)
- Method and sample size (how you'll answer the questions, and how many people)
- Timeline with real dates, not "a few weeks"
- Success criteria: what result would count as a clear answer
- Who needs to approve, and what they're approving (scope, screener, budget)
How to keep it short enough that people actually read it:
- One page of decision-relevant content up front; push detail a stakeholder doesn't need to approve (the full discussion guide, screener wording, consent language) into an appendix or a linked doc.
- Bullets over paragraphs, and cut any section that doesn't change what a stakeholder will say yes or no to.
- Replace research jargon with a plain restatement, so a stakeholder outside the research team can read it in one pass.
How to keep it specific enough that people can actually approve it:
- Every research question should be paired with the decision it feeds, so a stakeholder can judge "is this worth answering" rather than "does this sound reasonable."
- Success criteria stated as a number or an explicit signal ("at least 4 of 6 primary users complete the task unaided") rather than a vague goal like "get feedback."
- Timeline as calendar dates, and sample size stated as a number, so a stakeholder can judge feasibility rather than trust it blindly.
Worked example
A one-pager for a discovery study on automated invoicing for small-business users might read:
- Decision this unlocks: whether to build automated invoicing this quarter or continue with manual workflows.
- Primary research question: will target users adopt automated invoicing, and at what price point.
- Method: 12 moderated interviews plus a screener survey of 150 respondents to size interest.
- Timeline: recruiting weeks 1 to 2, fieldwork weeks 3 to 4, synthesis and readout week 5.
- Success criteria: a clear yes or no on the adoption hypothesis, and at least three validated pain points ranked by frequency.
- Sign-off needed from: the product lead (scope) and the design lead (screener and prototype).
That's roughly 150 words a stakeholder can read start to finish and either approve or push back on with a specific objection, rather than a 10-page plan they skim and rubber-stamp.
Trade-offs and pitfalls
Cutting too much detail can make the plan so vague that a stakeholder technically approves it but disputes the scope later, once real findings show up; keep the decision and success criteria concrete even while trimming everything else. Naming a method without a sample size stops a stakeholder from judging how much confidence to expect. And skipping "who signs off" invites the plan to get re-litigated later by people who were never asked to approve it in the first place.
A product manager proposes a small text change on a call-to-action button. How would you choose the single metric that decides whether the change worked, and what else would you watch while the test runs?
Sample Answer
Direct answer
Pick the single decision metric on two tests at once: business alignment (does moving this number actually matter to the goal the CTA serves) and sensitivity (is the metric noisy or "far" enough from the change that a small copy tweak realistically cannot move it in a readable window). For a CTA copy change, that usually means click-through rate (the share of viewers who click the button) rather than a downstream metric like trial-to-paid conversion, unless you have the traffic and time to wait for a signal that far downstream.
Why the "obvious" business metric can be the wrong pick
Imagine the real business goal is trial-to-paid conversion, and someone proposes measuring that directly. If baseline trial-to-paid conversion is 5%, and a CTA wording change could plausibly move click-through by a few points but has no direct mechanism to change what happens weeks later during a billing decision, then trial-to-paid conversion is both too far from the lever you pulled and too diluted by everything that happens in between (product experience, price, competing offers) to move detectably from copy alone. Click-through rate sits right next to the change, so it is sensitive enough to actually register the effect if one exists, while trial-to-paid conversion becomes a guardrail: something you watch to make sure a click-through win is not being bought at the expense of the thing you actually care about.
Guardrails and unintended consequences
Watch for a CTA that gets more clicks but pulls in the wrong people: increased bounce rate on the page it leads to, no change (or a drop) in the metric further down the funnel, or a rise in immediate back-button behavior. Also watch performance split by device and traffic source, since a wording change can read very differently on mobile than desktop.
Setting the threshold and decomposing for diagnostics
Fix the minimum detectable effect, and what result counts as "the change worked," before the test starts, not after you see the data: for example, agree in advance that only a relative lift of 10% or more in click-through rate justifies shipping, because anything smaller would not be worth the engineering and design cost of the rollout. If click-through does move, decompose it into the funnel it sits inside (impression to click, click to next-step) and by segment (device, new vs returning visitor, traffic source) to confirm the lift is broad-based rather than one anomalous segment carrying the whole result, which is the difference between a real effect and a fluke.
What is the difference between generative and evaluative research, and how do you decide which one a project needs right now? Name one method you would run for each, and say what would make you switch.
Sample Answer
Generative research figures out what to build and why, by exploring needs and mental models before a solution exists; evaluative research checks whether a specific design already works. Which one a project needs right now depends on whether a design exists to test yet, not on preference.
The core difference
- Generative, also called exploratory or formative: uncovers unmet needs, motivations, and mental models. One method: unstructured or semi-structured interviews with a target user group.
- Evaluative, also called confirmatory or summative: tests whether a design meets a bar for usability, comprehension, or impact. One method: moderated usability testing with a working prototype or live product.
One day versus four weeks
With one day, generative means 3 to 4 rapid remote interviews rather than a proper diary study, and evaluative means an informal think-aloud test with whoever you can grab, sometimes called a hallway or guerrilla test, rather than a full moderated study with a screener. With four weeks, generative can be a real contextual inquiry or diary study observing people in their own environment over days, and evaluative can be a properly recruited, benchmarked usability study or a well-powered A/B test.
How budget and regulatory risk shift the choice
A tight budget pushes toward generative methods, they need fewer participants and simpler setup, a handful of interviews rather than test infrastructure. High regulatory risk, a healthcare or financial product for example, pushes the other way, toward more rigorous evaluative work even under time pressure, because the cost of shipping something that fails silently is much higher than the cost of moving slower.
When one project needs both, and the order
A new feature almost always needs generative first, understand the need before you've committed to a specific design, and evaluative second, check the specific thing you built against that need. Example: a mobile banking budgeting feature. Sequence: generative (diary studies and contextual interviews on how people currently budget) leads to a rough design, then evaluative (moderated usability testing on early wireframes, then a benchmarked or A/B-tested version once the design stabilizes).
Worked example
Applied concretely: in weeks 1 to 2, six diary-study participants log how they track spending today, generative. By week 2 a clear pattern emerges, people want a warning before overspending, not a report after the fact, so week 3 moves into evaluative mode, testing a low-fidelity prototype of an in-the-moment spending warning against that specific need.
What would make you switch mid-project
If generative interviews keep circling the same unresolved question after several rounds with no new insight, that's a signal to stop exploring and evaluate a concrete design instead, exploration has diminishing returns once you're hearing the same things repeatedly. Conversely, if evaluative testing on an assumed-solved problem keeps surfacing a need you didn't expect, that's a signal to step back into generative mode rather than iterating the same design.
Trade-offs and pitfalls
The common mistake is running evaluative research on a design nobody validated the need for, which produces a usable interface for the wrong problem, or staying in generative mode indefinitely because it always feels like there's more to learn, without ever committing to something testable.
How do you make sure the sample you recruit is actually diverse and does not just skew toward power users and early adopters? Say how you would know afterwards whether it worked.
Sample Answer
Direct answer
Default recruiting channels, your own customer list, in-app intercepts, referrals, or a community forum you already run, are exactly the channels a power user or early adopter is most likely to be reachable through, so avoiding that skew means deliberately recruiting through channels outside the product itself, and you find out afterward whether it worked by comparing your recruited sample's makeup against your actual user base, not by feel.
Structured elaboration
Why default channels skew: someone who responds to an in-app prompt, follows the product on social media, or is active in a community forum has, by definition, engaged with the product more than average. A casual or lapsed user who opened the app twice and stopped almost never sees or responds to those channels. This is survivorship bias built into the recruiting method itself, before any screening even happens.
Structural fixes: recruit through channels that reach people regardless of current engagement, such as a random sample pulled directly from the user database rather than a self-selecting sign-up form, a general panel or paid ad rather than an in-app prompt, and explicit quotas by usage tier or tenure (targeting a mix of "used once and never returned," "casual," and "power user," rather than accepting whoever volunteers first). For reaching a specific underrepresented group, go through trusted community channels and partner organizations that already have credibility with that group, rather than a generic panel, and offer incentives that make sense to that community, not the standard product credit or early access that only appeals to people who already want more of the product, such as cash, a donation to a community organization, or practical support like transit or childcare costs.
Screening without profiling: ask about behavior and frequency of use, not identity or demographic proxies for it. A screener question like "how often do you use feature X" identifies a power user directly; a question that indirectly filters by income, neighborhood, or another demographic proxy can quietly exclude the exact underrepresented group you're trying to reach, even when usage tier was the only thing you meant to screen on.
Measuring whether it worked, after the fact: pull the recruited sample's actual usage-tier or tenure distribution and compare it side by side against the real user base's distribution from analytics. A mismatch is the checkable evidence that the channel skewed the sample, and it tells you to run a supplemental recruiting round targeting lapsed or casual users specifically, rather than trusting the first round's findings as representative.
Worked example
Suppose the true active-user base is roughly 15% power users, 55% casual or regular users, and 30% people who used the product once or twice and stopped. A first recruiting round sourced entirely from an in-app prompt produces a 20-person sample of 11 power users (55%), 7 casual users (35%), and 2 one-time users (10%), close to an inversion of the real base. That comparison is the concrete evidence to go back and specifically recruit from the lapsed or one-time-user list, pulled from the database rather than the app, before treating the findings as representative of typical users.
Trade-offs and pitfalls
Fixing this costs more effort than accepting whoever volunteers, since lapsed and casual users are inherently harder to re-engage. A screener question written to be too clever about avoiding an "obvious" bias can accidentally introduce a new proxy filter, so review every screening question by asking what else it might correlate with beyond what you intend to measure.
Unlock Full Question Bank
Get access to all User Research Planning and Fieldwork interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.