User Research and Discovery Questions
Planning and running user research and discovery to inform product direction. Covers research methodology, prioritizing a research roadmap, integrating findings into strategy, and partnering research with design and product. Assesses rigor in generating and applying user insight.
Your team shipped a redesigned checkout flow based on a research insight, and now needs to know whether it actually worked. Walk through how you'd turn that insight into a measurable KPI, design the experiment or metric-tracking plan to test it, and explain how you'd set a baseline, target, and timescale, plus what qualitative follow-up you'd run alongside the quantitative read to catch anything the metric alone would miss.
Sample Answer
Situation & chosen insight
I prioritized the insight: "High first-time onboarding drop-off on step 2 (profile details) is blocking activation." I translate that into a measurable KPI and an experiment.
KPI & baseline
- Primary KPI: Onboarding completion rate (users who finish all onboarding steps within 7 days of signup).
- Baseline: 42% completion (last 30 days).
Target & timescale
- Target: Increase to 55% within 8 weeks post-launch (roughly +13 percentage points, "pp" for short).
- Secondary metrics: Time to complete onboarding, conversion to activated user (first paid action), and drop-off rate at step 2.
Experiment design
- Variant A = current flow (control). Variant B = redesigned flow (progressive disclosure + inline validation + save-for-later).
- Randomized A/B test with a 50/50 UA split (User Allocation split: the percentage of users randomly routed into each variant).
- Minimum sample size: sized to reliably detect a conservative minimum lift of 4 percentage points at 80% statistical power. Power is the probability the test actually detects a real effect if one truly exists; 80% is the standard bar, meaning we accept a 20% chance of missing a real lift in exchange for a smaller, cheaper sample than 90% or 95% power would need.
- The arithmetic: using the standard two-proportion sample-size formula, n = 2 x (z_alpha/2 + z_beta)^2 x p_bar(1-p_bar) / (p1-p2)^2, with baseline p1 = 0.42, target p2 = 0.46 (a 4pp lift), p_bar = 0.44, z_alpha/2 = 1.96 (95% two-sided confidence) and z_beta = 0.84 (80% power): n = 2 x (2.80)^2 x (0.44 x 0.56) / (0.04)^2 = 2 x 7.84 x 0.2464 / 0.0016, about 2,415, which rounds to the N ~ 2,400 per arm target (adjust with real traffic and recheck as data comes in).
- Run for at least 2 full weeks and until sample size/statistical significance reached; monitor for seasonality.
Measurement & tracking plan
- Instrument events: signup_started, onboarding_step_completed (with step id), onboarding_completed, time stamps, user cohort id.
- Build dashboard showing funnel, completion rate, step-level drop-offs, time-to-complete, and segmented by device, channel, and new vs returning.
- Statistical test: two-proportion z-test with confidence intervals; also track lift and p-value (the probability of seeing a difference this large by chance alone if the redesign actually had no effect; conventionally, below 0.05 is treated as a significant result).
Qualitative follow-ups
- Recruit users from both variants who dropped off and who completed for short usability interviews (5–8 users per cohort).
- Review session recordings and heatmaps focused on step 2.
- Use feedback to iterate on microcopy, input affordances, and error messaging.
Outcome interpretation
- If KPI target met with significance and secondary metrics improve, roll out change. If lift is small or mixed, use qualitative insights to refine and run a follow-up test.
You're kicking off research for a new feature with a roadmap review in one week. Walk through how you'd decide whether to run qualitative interviews, send a quantitative survey, or do both, and explain what would change your answer if you only had two days instead of a week.
Sample Answer
Direct answer
With a week, the deciding factor is what kind of unknown you have: if you don't know why users behave a certain way or what shape the solution should take, run qualitative interviews first, since a genuinely unknown problem needs exploration before it can be measured. If you already understand the problem and need to know how many people are affected or which option performs better, a quantitative survey is the stronger first move. With a full week, a short round of each, a handful of interviews to sharpen the survey questions, then the survey to size what you found, usually beats either method alone. With only two days, drop the survey: getting a valid, representative response volume takes longer than two days to field, so a rushed survey produces a number that looks precise but isn't trustworthy. Run 4 to 5 rapid interviews instead, and pull whatever analytics or support-ticket data already exists.
Structured elaboration
Three questions to ask before choosing a method:
- What is actually unknown? An exploratory unknown ("why" or "what") points to qualitative work; a confirmatory unknown ("how many" or "which option wins") points to quantitative work.
- How much time do you actually have to reach a valid sample? Interviews can be scheduled and run within days; a survey needs days just to field and collect a stable response rate before you can trust the resulting percentage.
- Does the answer already exist? Check product analytics, support tickets, and past research before commissioning anything new. Often the quantitative half of the answer is already sitting in existing data, and the real gap is a handful of interviews to explain it.
When each method fits:
- Qualitative fits questions like "why are new users confused during onboarding," a brand-new feature concept with no existing usage data, or understanding a workflow nobody has documented yet.
- Quantitative fits questions like "what percentage of users would use this" or "which of two designs converts better," and sizing a problem you already understand well enough to measure before committing engineering time.
- Both together fit high-stakes decisions where you need to explain why something is happening and defend a number to leadership.
Worked example
One-week plan: day 1, recruit and draft a 5-question interview guide; days 2 to 3, run 6 contextual interviews (about 45 minutes each) about the workflow the feature targets; day 3 evening, draft a 6-question survey informed by the interview themes; days 4 to 5, field the survey to existing users through an in-app prompt, targeting at least 150 responses for a stable percentage; day 6, analyze; day 7, present both the synthesized themes and the survey top-line numbers to the roadmap review.
Two-day plan: day 1 morning, pull existing analytics and support-ticket data on the problem (a free, instant signal); day 1 afternoon and evening, run 4 to 5 rapid interviews with existing customers reached through customer success or a community channel, with no formal recruiting process; day 2, synthesize and present the roadmap review with the directional qualitative signal plus the existing analytics, explicitly labeled as directional rather than statistically powered. Skip a freshly launched survey entirely: one fielded for a single day only reaches the fastest-responding, most-online slice of users, which biases the result in a way that's easy to mistake for a real number.
Trade-offs and pitfalls
The most common mistake is treating a rushed two-day survey as if it carried the same evidentiary weight as a properly fielded one. Qualitative-only research skips confirming how common a pain point actually is, risking a build for a vocal minority; quantitative-only research skips the "why," risking solving the wrong problem precisely and confidently. And the biggest pitfall of all is picking the method you personally prefer running rather than the one the type of unknown actually calls for.
Explain when to use low-fidelity (paper), mid-fidelity (clickable) and high-fidelity prototypes in discovery. For each fidelity level, describe the minimum effort, typical validation questions it's good for, risks, and one example scenario where it is the right choice.
Sample Answer
Low-fidelity (paper/sketch)
- Minimum effort: 10–60 minutes; hand-drawn screens, flows, paper cutouts. Cheap to iterate during interviews or workshops.
- Good validation questions: Do users understand the concept/flow? Is the problem/opportunity framed correctly? Which features matter most?
- Risks: Users may dismiss it as “not real”; can’t test interaction timing or microcopy; ambiguous visuals can lead to misleading feedback.
- Example: Early discovery for a new onboarding flow. Test whether users grasp steps and value proposition before building clickables.
Mid-fidelity (clickable prototypes)
- Minimum effort: Hours–days; digital mockups with basic styling and clickable hotspots (Figma/Proto.io).
- Good validation questions: Is navigation intuitive? Does the task flow accomplish goals? Where do users hesitate/confuse?
- Risks: Looks more finished, which can bias feedback toward polish; limited backend realism (data/state edge cases).
- Example: Validating an account settings workflow with target users to refine labels, error states, and branching paths before engineering effort.
High-fidelity (pixel-perfect, interactive with realistic data)
- Minimum effort: Days–weeks; detailed UI, polished copy, realistic interactions, maybe instrumented for analytics.
- Good validation questions: Will users convert/complete tasks at acceptable rates? Do microinteractions, copy, and performance affect outcomes? Usability at scale.
- Risks: High cost and sunk effort if concept changes; stakeholders may conflate prototype with final product; longer iteration cycles.
- Example: Pre-launch A/B test of a redesigned checkout with realistic payment flows to estimate conversion lift and technical edge cases.
Use fidelity progressively: start low to validate problem and concepts, move to mid to refine flows, and use high to de-risk launch and measure real outcomes.
Explain the Jobs-to-Be-Done (JTBD) framework and describe when JTBD is more useful than personas for framing user problems. Provide a concrete JTBD-style statement and an example scenario (e.g., email management, grocery shopping) where JTBD leads to clearer product opportunities.
Sample Answer
Jobs-to-Be-Done (JTBD) is a user-centric framework that views customers as “hiring” a product or service to make progress on a specific task or goal in a context. Instead of profiling who the user is, JTBD focuses on the circumstance, desired outcome, and constraints: what job needs completing, when, and why.
Why use JTBD vs personas:
- JTBD is outcome-oriented and stable over time; personas describe people and can conflate behaviors with goals.
- JTBD surfaces unmet functional, emotional, and social jobs that reveal product opportunities across user segments.
- Personas are useful for messaging/UX empathy; JTBD is better for prioritizing features and business value.
Concrete JTBD-style statement:
“When I need to clear my inbox before starting work, I want to triage and defer low-priority emails quickly so I can focus on high-impact tasks without missing anything important.”
Example scenario, email management:
- Persona thinking: “Busy professional, 35–44, needs better email.” That suggests targeting a demographic with general UX tweaks.
- JTBD thinking: “Job = Rapid morning inbox triage under 10 minutes.” This leads to specific product opportunities: smart batching (auto-group newsletters), one-tap snooze-to-task, priority summary cards, and an accuracy-optimized classifier for what truly requires action. Those features map directly to measurable outcomes (time spent, emails deferred, task completion), making trade-offs and experiments clearer for a product roadmap.
Limitations and when personas still matter:
- A JTBD statement can be too abstract to act on by itself: knowing the job is "rapid morning inbox triage" doesn't tell you who is most underserved by today's product or which segment to design for first, so a lightweight persona or segment definition often still sits alongside JTBD rather than replacing it.
- JTBD foregrounds the functional job, but two users can share the identical functional job ("triage my inbox fast") for very different emotional or social reasons, such as avoiding anxiety versus looking responsive to a manager, and a JTBD statement alone can flatten that difference unless you deliberately capture the emotional and social dimensions too.
- Use personas when the decision is about who to target or how to talk to them (messaging, onboarding tone, segmentation for marketing); use JTBD when the decision is about what to build or which feature wins a trade-off, since it ties the choice to progress toward an outcome rather than to a demographic profile.
As a senior researcher, you're given competitor analysis, recurring support-ticket themes, a declining NPS trend, and a recent round of user interviews. Synthesize these into three strategic product initiatives, rank them by impact and feasibility, and give one success metric you'd track for each.
Sample Answer
In short
The job is to find where the four signal sources agree, not to average them: a theme that shows up in support tickets, the NPS (Net Promoter Score, a survey-based loyalty metric) decline, competitor gaps, and interview quotes at the same time is far more trustworthy than any one source alone, and that convergence should drive the ranking, not which signal is loudest.
Framework
Cross-reference the four sources for overlap before naming initiatives: cluster support-ticket themes, map interview quotes to the same clusters, check whether the NPS decline concentrates in a particular segment or moment (new users versus power users, a specific workflow), and note where a competitor already solves the same pain. A theme appearing in 3 or 4 of the sources is a strong candidate; a theme from only one source needs more evidence before it becomes an initiative.
Three ranked initiatives, ranked by impact times feasibility:
- Fix the highest-friction part of onboarding (impact 9, feasibility 8, score 9 x 8 = 72): if ticket volume and interview quotes both concentrate on the first-week experience and NPS is weakest among new users, this is usually the highest-confidence, fastest-to-ship initiative. Success metric: 7-day activation rate, the percent of new users completing the core first action.
- Resolve the top reliability or performance complaint (impact 8, feasibility 6, score 8 x 6 = 48): if tickets and interviews both point to a specific failure mode, a sync error, a slow step, that frustrates existing, higher-value users, this addresses churn risk directly. Success metric: related support tickets per 1,000 monthly active users, or MAU (the count of distinct users active in a typical month).
- Close the clearest competitive gap for power users (impact 6, feasibility 6, score 6 x 6 = 36): if interviews and competitor analysis both name a specific missing capability that power users want, this is a retention and win-back play rather than a growth play. Success metric: adoption of the new capability among the top decile of most active users within 30 days.
Worked example
Suppose support tickets show 22% of volume relates to account-setup confusion, interviews (8 of 12 participants) independently raise the same confusion unprompted, and NPS is 15 points lower among users under 30 days old than among tenured users. All three signals converge on onboarding, so it ranks first with high confidence. By contrast, a "missing export feature" appears in only 2 of the 12 interviews and isn't visible in ticket data at all; it goes on a watch list rather than becoming one of the three ranked initiatives, since a single source isn't enough to commit resources against.
Trade-offs and pitfalls
The main trap is anchoring on whichever source is most vivid, usually a strongly worded interview quote, and retrofitting the other data to support it; check convergence across sources deliberately rather than trusting your first impression. The second trap is presenting these as three independent tactical fixes instead of naming what they imply about the competitive landscape and where the product should place longer-term bets, a more strategic framing worth using when the audience is executives rather than the immediate product team.
Unlock Full Question Bank
Get access to all 7 User Research and Discovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.