Product Sense and Design Questions
Reasoning about what to build and why, for a given user and business context: reading a vague brief, generating and evaluating candidate feature or product concepts, and defending the chosen concept with structured logic (user needs, business impact, feasibility). Typical prompts are open-ended 'design a product for X' or 'improve product Y' briefs, scoping a Minimum Viable Product (what ships first and why), Jobs-to-be-Done style problem framing, picking a small set of metrics to judge a proposed concept's health, and making a single interaction-shape judgment call within the chosen concept (for example, opt-in vs. default-on, or exposing a power feature vs. keeping it hidden). Also covers anticipating what could go wrong with a proposed concept, such as adoption failure or harm to an underserved subgroup, before it ships. Assesses taste, creativity, and the ability to turn an ambiguous brief into a coherent, defensible product proposal. Out of scope: the multi-phase design process itself (ideation through validation, the double diamond), platform or technical-roadmap prioritization mechanics (RICE, ICE, MoSCoW, Cost of Delay scoring drills), system or infrastructure architecture, and narrating the candidate's own past work (STAR-style storytelling).
Identify one thing about Spotify's user experience you would improve. Describe the problem, your proposed solution at a high level, and the measurable outcomes you'd use to evaluate success.
Sample Answer
Direct answer
I would improve Spotify's autoplay continuation, the feature that keeps playing similar music after
a queue or playlist ends. The problem: when someone is using music to hold a specific mood or pace
(a workout, a focus session), autoplay often continues on genre or artist similarity and lands on a
track with a mismatched tempo or energy, breaking the exact mood the listener built the playlist to
sustain. The fix: weight continuation choices by acoustic-feature closeness (tempo, energy, and
valence, the audio characteristics that describe a track's pace and mood) to the listener's
recently played tracks, not just artist or genre similarity, and give users a simple toggle between
"keep the vibe going" and "surprise me."
Structured elaboration
The underlying job to be done is "keep a mood going without having to manually queue more music,"
not "discover something new right now." Autoplay today mostly solves for the second job (finding
something plausibly related) even in moments when the user actually wants the first. The fix
targets that mismatch directly: instead of ranking continuation candidates mainly by collaborative
signal ("people who liked this track also liked that one") or shared genre and artist, weight
candidates by proximity to the rolling average tempo, energy, and valence of the last several tracks
played. A simple toggle exists because forcing one universal default trades one failure mode for
another: some listeners genuinely want autoplay to introduce them to something different, and a
system that only ever protects the current vibe would quietly undercut Spotify's other core value,
discovery.
Worked example
Run it as an A/B test: control keeps today's autoplay behavior, treatment uses acoustic-feature
weighting for continuation candidates. Illustrative results, meant to show what success would look
like rather than a claim about real data: skip rate on the first two autoplay tracks drops from 34%
in control to 24% in treatment, average listening session length after the original playlist ends
increases from 9 minutes to 13 minutes, and thumbs-down rate on autoplay-selected tracks drops from
6% to 4%. Crucially, the autoplay opt-out rate stays flat at roughly 5% in both arms, which is the
guardrail check that the fix actually worked rather than just making people turn the feature off and
manually curate instead.
Trade-offs and pitfalls
Over-optimizing for mood consistency risks making autoplay boringly repetitive, undercutting music
discovery, which is exactly why the "surprise me" toggle needs to exist rather than a single
tightened default. Acoustic-feature similarity is also only a proxy for mood: two tracks can share
tempo and energy and still feel emotionally mismatched, so this is a real, measurable improvement,
not a fully solved problem. Finally, the change needs to be checked against overall listening time
as the primary business metric; fixing the narrower mood-continuity complaint would not be worth
shipping if it quietly reduced how much people listen overall.
You and a designer disagree about where to spend a fixed slice of engineering time: a prominent UI callout that would increase a feature's discoverability, or backend work that would make the feature itself noticeably better. How would you decide between the two, and what would you look at quickly to avoid just guessing?
Sample Answer
Bottom line
Don't debate this on taste. Split it into two separate questions, get a cheap, fast read on each from data you likely already have, and only then compare the two using a rough return-on-investment (ROI, value created per unit of effort spent) estimate.
How to decide
-
Separate the two failure modes the callout and the backend work each fix:
- The callout fixes a discoverability problem: people never find or try the feature.
- The backend work fixes a quality problem: people find it, but it disappoints them once they use it.
A feature can suffer from either, both, or neither, and the fix only works if it targets the actual bottleneck.
-
Pull cheap signals before guessing, all things you can usually get from existing analytics and support data in a day or two, with no new code:
- Funnel data: of all active users, what percent ever open the feature (tells you if discoverability is the bottleneck), and of those who open it, what percent complete it or return to it (tells you if quality is the bottleneck).
- A quick read of 30 to 50 recent support tickets or session recordings tagged to the feature, to see whether drop-off looks like people getting lost in the navigation versus people hitting an actual defect or a "this isn't good enough" moment.
-
Turn those signals into a rough ROI comparison:
ROI=engineering costexpected incremental value
Expected incremental value is roughly (number of additional or improved conversions the change would produce) times (value per conversion). Engineering cost is the effort, in developer-weeks, for each option. You don't need precision here, you need the two ratios to be different enough to point clearly one way.
- If the two ROI estimates are close, don't force a binary call: a small, fast experiment for each (a feature-flagged callout to a slice of traffic, and an equally-scoped fix for the top quality complaint) settles it with real data faster than another round of debate.
Worked example
Say a trip-planning app has 100,000 monthly active users, and a cost-splitting feature is buried three taps deep. Funnel data shows only 4% of users ever open it (4,000 people), and of those, 55% complete it (2,200). That split alone is informative: 96% of users never even see the feature, which points at discoverability as the bigger-volume bottleneck.
Suppose a similar UI callout shipped in another part of the product previously lifted a comparable feature's open rate from 4% to 9%, a real precedent to anchor the estimate on. Applied here, that's roughly 5,000 additional openers a month, and if completion rate holds at 55%, about 2,750 additional completions a month.
Now the backend option: support tickets show a specific failure (splits don't handle uneven groups well) affecting a chunk of users who try to complete the flow. Suppose fixing it would lift completion rate among the existing 4,000 openers from 55% to 70%, that is 4,000 x 0.15 = 600 additional completions a month, from the same audience size that already exists today.
At comparable engineering cost (call it one sprint for either option), the callout produces roughly 4 to 5 times more additional completions by volume. On pure ROI, the callout wins here, but that number alone isn't the whole decision.
Trade-offs and pitfalls
- A callout is more visible and more demoable, which biases people toward preferring it in a review even when the numbers say otherwise. Don't let "which one is easier to show off" substitute for "which one the data supports."
- If the feature genuinely has a quality problem, driving more traffic into it with a louder callout can backfire: more people try it, more people bounce or leave frustrated, which can hurt overall app ratings even as the "open rate" metric goes up. When completion quality is clearly broken, it's often safer to fix or at least de-risk that first before spending a scarce, high-visibility UI slot to drive volume into it.
- A prominent callout also has an opportunity cost: that UI real estate could be promoting something else, so the comparison isn't really "callout vs. backend," it's "this feature's callout vs. everything else that wants that same spot."
- This doesn't have to be all-or-nothing. If both problems are real and the effort is divisible, a smaller version of each (a modest UI nudge plus a scoped fix for the worst quality complaint) can capture most of the value without betting the whole sprint on one lever.
Using the Jobs-to-be-Done framework, propose JTBD statements for a ride-sharing app's 'quiet ride' feature. For each JTBD, list acceptance criteria, minimal technical requirements, priority relative to other features, and an experiment to validate demand.
Sample Answer
Direct answer
A strong candidate names the underlying job in plain, functional-plus-emotional terms before proposing anything else. Jobs-to-be-Done, or JTBD, is a framework that describes what "job" a customer is hiring a product to do, centered on the underlying need rather than a feature request or a demographic label. A single "quiet ride" request usually hides two or three genuinely different jobs, and each deserves its own acceptance criteria, technical scope, and validation plan rather than one compromise design.
JTBD 1: the everyday decompression ride
"When I'm commuting after a long day, I want to arrive without having made small talk I didn't want, so I can use the ride to unwind instead of performing politeness."
- Acceptance criteria, written as numbers so a pilot result can actually fail them: riders can select a quiet-ride preference at booking; at least 80% of riders who selected it rate the ride's quietness 4 or better out of 5 in the post-ride survey; drivers acknowledge the preference at pickup on at least 90% of matched trips. The 90% acknowledgment bar is the load-bearing one, because below roughly nine trips in ten the rider learns the setting is unreliable and stops trusting it, which is a worse outcome than never having shipped it. The 80% satisfaction bar is deliberately looser, since a rider can rate a ride badly for reasons that have nothing to do with noise. Both are pre-pilot hypotheses and both get re-cut after the first market, but they get written as numbers first, because "most riders" and "the large majority" are criteria no result could fail.
- Minimal technical requirements: a preference toggle in the booking flow, a matching flag and acknowledgment prompt in the driver app, plus a survey question and analytics event to measure adherence.
- Priority: high. It's the broadest job (most riders have had an unwanted-conversation ride at some point) and the cheapest of the two to build.
- Experiment to validate demand: expose the option to a slice of riders, measure selection rate and satisfaction versus a holdout, paired with a short survey on why riders did or didn't select it.
JTBD 2: the work-mode ride
"When I need to take a call or focus during a ride, I want assurance there won't be background conversation or music, so I don't have to apologize to whoever's on the line."
- Acceptance criteria: riders can opt into a stricter no-conversation, no-music mode; at least 60% of riders who select it during weekday business hours report in the follow-up survey that they used the ride for work. 60% is the point at which the job named in the statement is genuinely the majority use rather than a story told about a feature people are selecting for some other reason, which would mean this second job is really the first one wearing a different label and does not need its own tier.
- Minimal technical requirements: an additional preference tier beyond basic quiet mode, with the preference persisted across a rider's saved trip settings.
- Priority: medium-high. A narrower audience (business commuters) but often willing to book more predictably or pay a small premium during weekday hours.
- Experiment to validate demand: a targeted beta invite to riders who frequently book weekday business-hours trips, tracking opt-in and repeat use.
Worked example
Suppose the JTBD 1 experiment runs for two weeks in one metro market, with the option exposed to a randomly chosen 5% of riders. Riders, not trips: the preference is a rider-level setting, and a rider who saw the toggle on Monday but not on Thursday would be sitting in both arms at once. The reading then has to be exposed arm against holdout arm, everyone in each arm counted, not selectors against non-selectors, and that distinction is most of the experiment. A meaningful minority actively select quiet ride when offered, for example about 18% of exposed riders choose it once shown the option, and it is tempting to compare those 18% against the riders who declined and report the gap as the feature's value: selectors rate post-ride quietness at roughly 4.6 out of 5 against about 3.9 for riders who declined, so the feature appears to be worth 0.7 of a point. That comparison measures a preference, not a treatment: riders who choose quiet ride are the ones who wanted quiet, so of course they rate quietness higher, and the riders who declined are not an unexposed group, they are exposed riders who said no. Read the right way, average post-ride quietness across the whole exposed arm comes in at about 4.0 out of 5 against 3.9 for the holdout arm (0.18 x 4.6 + 0.82 x 3.9 = 4.03, since the 82% who declined get the same ride they always got). That is a lift of roughly a tenth of a point rather than the 4.6-versus-3.9 the selector-only comparison would have shown, and it is the only version attributable to shipping the feature; the 18% selection rate is reported alongside it as its own result rather than folded into the satisfaction comparison.
The one thing this design cannot settle is the matching-side cost, and it is worth saying so rather than quietly claiming it. Driver acceptance and wait time are properties of a market's driver pool, not of an individual rider, and both arms here are served by the same pool in the same city, so any constraint the exposed arm creates is absorbed by cars that also serve the holdout. At 5% exposure with 18% opt-in, under 1% of the market's trips carry the constraint at all, far too little to move a fleet-level number in either arm. "Drivers accepted matched trips at about the same 97% rate as before" is therefore evidence that the pilot was too small to hurt anything, not evidence that the feature is free at full rollout. Answering that question needs a market-level design, matched cities or a switchback where the feature flips on and off market-wide on alternating blocks, run once rider demand is established. So the honest read of this pilot is an 18% opt-in rate and a real but modest arm-level satisfaction lift, enough to justify moving to the narrower work-mode test next, with the marketplace-liquidity question explicitly still open.
Trade-offs and pitfalls
Treating "quiet ride" as one feature instead of two distinct jobs produces a single compromise design that under-serves everyone: too strict for casual commuters, not reliable enough for someone on a work call. Enforcement is the hardest part, since a feature depending entirely on driver compliance needs a plan for drivers who don't follow it, or rider trust collapses fast. And prioritizing the narrower, more monetizable job before validating the broad one risks investing in a segment before you even know the base feature works; the emergency and safety button must also stay fully reachable and unaffected by any quiet-mode setting. The easiest way to get a false green light on an opt-in feature like this is to compare the riders who opted in against the riders who did not: that gap is mostly self-selection, and it will look impressive whether the feature works well or barely works at all.
Using the Jobs-to-Be-Done framework, describe the primary job a personalized news feed is hired to do for users. Name three user pains, three measurable outcomes you'd track, and one concrete hypothesis for a change that addresses a key pain. How would you validate that hypothesis before and after building it?
Sample Answer
Direct answer
A strong candidate names the underlying job before jumping to metrics or a fix. Jobs-to-be-Done, or JTBD, describes what job a product is hired to do for the user, independent of any specific feature. A personalized feed's real job usually isn't "show me content," it's closer to "let me feel caught up and informed without spending time I don't have filtering through noise," and every pain, metric, and hypothesis that follows should trace back to that job.
Primary job, pains, and outcomes
Primary job: help me quickly find news that's relevant and trustworthy so I can feel informed and act on it (share, discuss, decide), without wasting time filtering out noise myself.
Three user pains: overload, where the feed surfaces too many similar or irrelevant stories and the user has to filter manually; staleness or low relevance, where stories feel generic or out of step with what the user actually cares about right now; and trust and diversity, where users worry the feed shows a narrow slice of viewpoints or content of uncertain credibility, undermining the "informed" part of the job.
Three measurable outcomes:
| Outcome | What it measures | Which pain it tracks |
|---|---|---|
| Time-to-first-meaningful-engagement | Time from opening the feed to a genuine read, save, or share, not just a scroll | Overload: faster time to value means less time spent filtering |
| Relevant-engagement rate | Share of shown stories the user reads, saves, or shares, weighted by time spent | Staleness and low relevance: directly reflects whether the feed surfaces the right stories |
| Return frequency | How often a user comes back to the feed | The job overall: a proxy for whether the feed delivers enough value to be worth the habit |
Those three are the outcomes you are trying to move. The third pain, trust and diversity, needs a fourth measure of a different kind: a guardrail, a number you are trying to hold flat rather than push up, because relevant-engagement rate can rise precisely by narrowing the feed. Without one, the metric set reports a win in exactly the case where the product got worse on a pain you named yourself. Two concrete, cheap options:
- Source-set breadth per user: the count of distinct sources a user engaged with over a rolling 28 days, reported as the median across users. Ship-blocking condition: it must not fall relative to the control group.
- Periodic trust survey: a single question ("does this feed show you a fair range of viewpoints?") sampled to a small slice of users each week, tracked as a trend rather than a one-off reading.
Pick one and pre-commit its threshold before the experiment starts, or it becomes something to argue about after the results are in.
Hypothesis and validation
One concrete hypothesis: adding a short, visible "why you're seeing this" line next to unfamiliar or non-followed sources will reduce the perceived-noise pain and raise relevant-engagement rate, because users are more willing to engage with an unfamiliar story once they understand why it appeared.
Before building: run a quick concept test with a clickable mockup or a manually curated stand-in for the real feature, with a handful of representative users, watching whether the explanation changes their willingness to engage with unfamiliar stories, and asking whether it reads as helpful or as an excuse.
After building: run a controlled experiment comparing the feed with and without the explanation, measuring relevant-engagement rate, return frequency, and the diversity guardrail over a few weeks, paired with a short survey on perceived trust and relevance, so you can tell whether any lift comes from genuinely better understanding rather than novelty.
Size that experiment before you run it rather than after. Detecting a 3 percentage point lift on a 22% baseline relevant-engagement rate, at the usual 80% power and 5% significance level, needs roughly 3,100 users per arm. A three-arm test (control plus two phrasings) therefore needs about 9,400 users, and a smaller test that comes back "no measurable lift" has told you nothing, since it could not have detected the effect you were looking for either way.
Worked example
Suppose the pre-build concept test with 10 to 12 users shows most, 8 of the 11 users tested, say the explanation makes them more willing to try an unfamiliar story, while a couple, 2 of the 11, say it reads as an excuse for showing them something they didn't want. That mixed signal is still useful: it tells you the wording matters as much as the feature itself, so the post-build experiment should test at least two phrasings, not just presence versus absence. Suppose that post-build experiment then runs for three weeks at about 3,100 users per arm: the better-performing phrasing lifts relevant-engagement rate from a baseline of 22% to 25% and return frequency from 2.1 to 2.3 sessions per week, the weaker phrasing shows no measurable lift on either metric, and median source-set breadth holds flat against control in both arms. That last part is what lets you ship: the engagement lift did not come from quietly narrowing the feed.
Trade-offs and pitfalls
Chasing relevant-engagement rate alone risks over-fitting the feed to what a user already likes, quietly recreating the trust and diversity pain even while the engagement number rises, which is exactly why the diversity guardrail has to be pre-committed with a threshold rather than added as a post-hoc explanation. A concept test with a handful of users can catch an obviously bad explanation but can't tell you the real effect size, so treat it as a filter before the real experiment, not a substitute for it. And return frequency is a slow-moving metric; don't expect it to move before the shorter-term engagement metric does, and don't panic if it lags behind.
How would you define 'product sense', and what does having it actually look like day to day, translating observed user needs and business goals into a concrete plan rather than just an opinion? Give one example where you (or someone you admire) showed it well.
Sample Answer
Direct answer
Product sense is the calibrated judgment to turn incomplete, ambiguous information about users and
the business into a specific, defensible decision about what to build next, fast enough that waiting
for a full study would have cost more than it was worth. It is not an unexplainable gut feeling.
Day to day, it looks like constantly building a real model of the user from every available signal,
then translating a vague observation into a concrete, falsifiable plan rather than stopping at an
opinion.
Structured elaboration
The distinction the question is really asking about is where "opinion" stops and "product sense"
continues. An opinion stops at "I think users want X." Product sense continues: "...because of
evidence Y, so we should build Z, and we'll know we're right if metric M moves within N weeks." That
extra clause, the evidence and the falsifiable check, is the actual skill, not the initial hunch.
Day to day this shows up as a few habits:
- Grounding: treating support tickets, sales calls, usage data, and direct product use as a
continuously updated mental model of the user, so that when a decision has to be made in a meeting
with no time to commission new research, there is real evidence to reason from rather than a
guess invented on the spot. - Specificity: turning a vague signal ("users are churning") into a precise, testable claim
("users who don't complete onboarding step three within their first session churn at three times
the rate of those who do, so the plan is X"). Specificity is what separates product sense from an
opinion that merely sounds confident. - Holding the tension explicitly: weighing what a user wants against what the business can
afford to build, and stating the trade-off out loud (a feature might delight a segment that is
2% of revenue but cost two engineer-months) rather than pretending the tension does not exist. - Fast, falsifiable framing: proposing a plan that can be shown wrong quickly, a small test and
a specific metric, rather than an unverifiable big bet. Good product sense includes accepting that
you might be wrong and building in a fast way to find out.
Worked example
Here is an illustration of the pattern, not a specific attributed story: a product leader at a
food-delivery company notices that "where is my order" support contacts spike specifically during
peak dinner hours, not in overall order volume. The opinion-shaped response would be "we need a
general order-tracking redesign." The product-sense response is narrower and falsifiable: add a
live map and dynamic ETA specifically to the tracking screen during peak hours, on the hypothesis
that the underlying need is reassurance under uncertainty, not a full UI overhaul, and set a bar
that is actually specific: "where is my order" contacts per 1,000 peak-hour orders down 25% within
one quarter. Where that number comes from is the part that matters. In this illustration peak-hour
contacts run at roughly twice the off-peak rate, and that gap is the anomaly the whole hypothesis
rests on, so closing half of it is the smallest move that would be visible above week-to-week noise
and would still pay for the work. Writing "a meaningful drop" instead would have been exactly the
failure the rest of this answer is arguing against: a bar with no number in it cannot be missed, so
it can never tell the team the hypothesis was wrong, and a check that cannot fail is an opinion
wearing the costume of a plan. The pattern is: observed signal, translated into a
specific, falsifiable, resource-proportionate plan, not a sweeping redesign justified by the same
observation.
Trade-offs and pitfalls
The most common failure is mistaking confidence for product sense: a loudly stated opinion is not
the same as a well-reasoned, falsifiable one, and interviewers (and teams) can be fooled by
conviction alone. A second failure is letting a past win calcify into an unexamined prior, applying
a pattern that worked for a previous user base to a current one without checking whether the
underlying need actually transferred. A third is using "product sense" as a shield against ever
validating anything, when a cheap test was available and skipped purely to avoid the discomfort of
being provably wrong. Good product sense includes knowing when a quick, cheap test is worth running
anyway, not treating judgment and validation as opposites. The everyday form of that shield is the
unquantified success bar: naming the metric and the horizon but never the size of the move you would
accept, which reads as rigor in the room and commits to nothing a quarter later.
Unlock Full Question Bank
Get access to all 9 Product Sense and Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.