Product Sense and Design Questions
Reasoning about what to build and why, for a given user and business context: reading a vague brief, generating and evaluating candidate feature or product concepts, and defending the chosen concept with structured logic (user needs, business impact, feasibility). Typical prompts are open-ended 'design a product for X' or 'improve product Y' briefs, scoping a Minimum Viable Product (what ships first and why), Jobs-to-be-Done style problem framing, picking a small set of metrics to judge a proposed concept's health, and making a single interaction-shape judgment call within the chosen concept (for example, opt-in vs. default-on, or exposing a power feature vs. keeping it hidden). Also covers anticipating what could go wrong with a proposed concept, such as adoption failure or harm to an underserved subgroup, before it ships. Assesses taste, creativity, and the ability to turn an ambiguous brief into a coherent, defensible product proposal. Out of scope: the multi-phase design process itself (ideation through validation, the double diamond), platform or technical-roadmap prioritization mechanics (RICE, ICE, MoSCoW, Cost of Delay scoring drills), system or infrastructure architecture, and narrating the candidate's own past work (STAR-style storytelling).
You are evaluating adding social features to increase engagement. Compare three product models: follow model, shared spaces/groups, and in-app messaging. For each model, list expected user behaviors, required technical investments, potential risks, and the top three metrics to validate impact.
Sample Answer
Direct answer
All three models raise engagement through a different mechanism: the follow model drives content
discovery and broadcast, shared spaces build belonging and community, and messaging drives direct
utility and coordination. Which one to build first depends on which underlying problem the app
actually has right now, not a general belief that "social features help." Most social apps that
succeed end up layering all three eventually, usually starting with whichever one resolves their
current cold-start problem (not having enough existing users, content, or connections yet for the
core experience to feel valuable to a brand-new person) fastest.
Structured elaboration
| Model | Expected user behavior | Technical investment | Key risks | Top 3 metrics |
|---|---|---|---|---|
| Follow model | Users curate a one-directional feed of people or creators they find interesting | Follower-graph storage, feed-ranking logic, notification fanout at scale | Feels noisy or spammy; follower counts get gamed by bots; painful cold start for new users with zero follows | Follow-through rate (search to follow), % of feed sessions with meaningful engagement, follower-graph growth |
| Shared spaces / groups | Users join topical or interest-based groups and participate in threaded discussion | Group creation and discovery tooling, in-group content ranking, spam and abuse detection at group scale | Degrades into low-quality or toxic content without active moderation; fragments into many small dead groups | % of groups with active weekly posting, member-to-active-poster ratio, 30-day group retention |
| In-app messaging | Users have 1:1 or small-group direct conversations, high-frequency utility use | Real-time infrastructure (sockets or push), spam and harassment detection for content that is no longer public | Can cannibalize the core app if it becomes the primary reason to open it; private abuse is much harder to moderate than public content | Messages sent per daily active user, response rate and time, % of daily actives with an active conversation in the last 7 days |
Worked example
Suppose usage data shows most new-user drop-off happens immediately after signup, specifically
because the feed is empty (nobody to follow yet). That is direct evidence pointing at shared
spaces first, not the follow model: dropping a new user into an existing topical community gives
them "people like me are here" belonging on day one, whereas a follow model needs a large base of
people worth following to pay off, which a new user has no way to discover on their own.
The step that argument skips, and the one an interviewer will push on, is that groups have a
cold-start problem of their own. On the day the feature ships there is no existing community to drop
anyone into, and an open "create your own group" surface launched into a vacuum produces exactly the
risk named in the table above: fragmenting into many small dead groups, which is a worse first
impression than the empty feed you were trying to fix. The reason groups still beat the follow model
here is not that they escape cold start, it is that their cold start is solvable by hand and the
follow model's is not: a team can seed ten to twenty topical spaces itself and pay staff or invited
power users to post in them for a few weeks, and no team can hand-manufacture thousands of people
worth following. So the sequencing decision comes with a launch design attached: ship a small,
closed set of seeded groups that are already being posted in rather than open group creation, hold a
density bar (every group a new user can see has had a post in the last 48 hours, and no group is
visible below some floor of active members), and only open self-serve group creation once organic
posting sustains that bar without staff effort. If the bar cannot be held with seeding, that is the
signal that groups are not the answer either and the real problem is upstream of any social feature. Messaging
would be premature here too: it needs the other two to exist first so users have any shared context
to message each other about. If, say, 70% of day-one churners bounced right after seeing an empty
feed with zero follows, that single data point is enough to sequence groups ahead of the follow
model, even though the follow model is the more familiar pattern to reach for by default.
Trade-offs and pitfalls
Building messaging before there is a reason to message someone (no shared context, no existing
relationship) tends to sit unused. Groups require real, ongoing moderation investment; skipping
that investment to hit a launch date is how groups rot into low-quality spaces within weeks. The
follow model is the slowest of the three to pay off since it needs supply (people worth following)
to exist before demand shows up, which makes it a poor first bet for cold-start problems even though
it is often the first thing teams reach for. The single most common wrong turn is trying to ship all
three at once: it dilutes engineering focus into three shallow, half-built experiences instead of
one that actually works.
You just shipped a change meant to improve developer onboarding on your API platform. What metrics would tell you whether it's actually working, and how would you tell a leading signal from something that's just a lagging vanity number?
Sample Answer
Direct answer
A strong candidate first defines what "working" means for this specific change, usually that developers reach a successful first integration faster and stick around afterward, and then separates candidate metrics into leading indicators that move quickly and predict the outcome versus lagging indicators that only confirm it much later, stress-testing any leading metric by asking whether it could rise for a reason that has nothing to do with real success.
Leading versus lagging, and the vanity test
A leading indicator changes soon after the change ships and sits causally close to the behavior you actually care about. A lagging indicator confirms the outcome weeks or months later, by which point you've already waited a long time to learn anything. For a developer-onboarding change, useful candidates and how to classify them:
- Time-to-first-successful-API-call (leading): tightly connected in time to the change itself.
- Onboarding-flow completion rate (leading to mid): did people finish the quickstart, not just start it.
- Docs page views or quickstart visits (vanity, not a real leading signal on its own): this can rise because the flow confuses people into re-reading the same page, or because of an unrelated traffic spike; it doesn't tell you the flow got easier.
- 7 or 30-day developer retention (lagging): confirms the change mattered, but takes weeks and is affected by many other things that happen in between.
- Onboarding-related support tickets (leading, inverse signal): a drop suggests real friction actually went away.
The test for whether a candidate leading signal is real or vanity: does it move within days of the change and sit close, causally, to the behavior you want (reaching a working integration), and can you think of an obvious way it could move without the underlying experience actually improving. If a metric fails either check, keep measuring it, but don't treat it as evidence on its own.
Worked example
Suppose before the onboarding redesign, the median time from signup to first successful call was 40 minutes, and afterward, among the first 200 developers who sign up post-launch, the median drops to 15 minutes, while quickstart page views also jump. The page-view jump alone would be a false positive if treated as success. To be confident the drop in time-to-first-call is real, you'd want two more things: developers aren't re-visiting the page more times per session than before (meaning they need it less, not more), and a comparison against a similar earlier cohort (a group of developers who signed up in a comparable prior period, used here as a stand-in comparison group), or a small holdout that didn't see the change, to rule out a coincidence like a marketing campaign landing in the same window.
Trade-offs and pitfalls
Waiting only for lagging metrics like 30-day retention to declare success means learning slowly and losing weeks of course-correction time. Treating any metric that goes up as proof is the classic vanity-metric trap; always ask what else could explain the movement. And without a comparison group, whether a real holdout or a matched earlier cohort, even a convincing-looking before-and-after change can be confounded by something unrelated that shipped in the same window.
Describe the hierarchy of metrics you would set up to monitor product health (e.g., north star, leading indicators, lagging metrics). Then, for a social consumer app, propose a three-level metric tree including a north-star and two leading indicators with definitions and why they matter for problem solving.
Sample Answer
Direct answer
The hierarchy has three layers: a single north star metric at the top representing overall value
delivered, leading indicators beneath it that the team can move now and that predict where the
north star is heading, and lagging or outcome metrics that confirm value was actually captured over
a longer horizon (revenue, long-term retention). For a social consumer app, I would set the north
star as weekly connected users, users who both post or react and receive a reaction within the
week, supported by two leading indicators: new-user activation rate and content-response rate.
Structured elaboration
The point of the hierarchy is diagnostic, not decorative: when the north star moves, leading
indicators tell you where in the system the movement is coming from, before the (slower, noisier)
lagging metrics would ever reveal it.
- North star: weekly connected users, users who receive at least one reaction to something
they posted or shared during the week. This is deliberately a two-sided metric: pure posting
volume can be inflated by a few very active users, while requiring a received reaction ties the
metric to real reciprocated value on the network, not just activity. - Leading indicator 1, new-user activation rate: the percent of signups who complete a first
meaningful action (post plus receive one reaction) within 24 hours. It matters for problem
solving because it isolates whether the top of the funnel is healthy, independent of what is
happening to the existing user base. - Leading indicator 2, content-response rate: the percent of posts that get at least one
reaction or comment within 24 hours. It matters because it is a direct read on network density:
a network that has stopped responding to new content will drag the north star down within weeks,
and this metric will show the problem well before that happens.
Worked example
Say weekly connected users sit at 40,000 out of 100,000 weekly active users (40%), and it drops to
34,000 the following month. Checking the leading indicators: activation rate held steady at 55%
(new users are onboarding fine), but content-response rate on posts from days 3 through 7 of the
week fell from 62% to 48% (existing-user engagement is fading). That pattern points the
investigation at the feed ranking algorithm or a notification regression affecting existing users,
not at onboarding, and it does so weeks before a lagging metric like monthly retention would have
surfaced the same story.
Trade-offs and pitfalls
A north star built from too many combined conditions becomes unintelligible and hard for any single
team to move; keep the definition to the smallest set of conditions that still captures real,
reciprocated value. Lagging metrics like revenue or long-term retention should not drive weekly
team decisions: they move too slowly and arrive too late to be actionable at that cadence, which is
exactly why the leading-indicator layer exists. Finally, any leading indicator that a team is
directly rewarded for moving needs a guardrail alongside it (for example, reports or blocks per
session next to content-response rate), or the team will find the fastest way to inflate the number
rather than the healthiest one.
For an early-stage marketplace connecting freelance tutors with students, what would you pick as the north star metric, and why does it beat the other obvious candidates you're rejecting? Name three supporting metrics you'd track alongside it.
Sample Answer
Direct answer
I would pick weekly completed tutoring sessions as the north star metric: the count of sessions
that actually happen, not sessions booked, tutors signed up, or dollars changing hands. It beats
gross marketplace volume (GMV), total registered users, and average rating because it is the one
number that can only go up when real value gets exchanged on both sides of the marketplace at once.
Structured elaboration
A good north star metric for a two-sided marketplace has to satisfy three things: it reflects value
delivered to both sides simultaneously, it moves in the direction the business actually wants
(growth, not just activity), and it is hard to inflate by moving only one side of the market.
Here is why the obvious alternatives lose:
- GMV (gross merchandise value, total dollars transacted): GMV can rise purely from a price
increase even while fewer sessions happen and the marketplace is getting thinner. It measures
revenue, not liquidity (whether supply and demand are actually matching), and liquidity is the
thing an early marketplace lives or dies on. - Total registered tutors or students: a supply-side or demand-side vanity count. Someone can
sign up and never book or teach a single session; the number can look healthy while the
marketplace is actually dead. - Average session rating: a real quality signal, but a lagging one that says nothing about
growth, and it is trivially misleading at low volume (a 5.0 average built on three total sessions
tells you nothing). It belongs as a guardrail, not a north star.
Three supporting metrics I would track alongside weekly completed sessions:
- Fill rate: the percent of booking requests that convert into a completed session. This is
the earliest signal of marketplace liquidity and tells you whether growth is being held back by
a matching problem before the north star even moves. - Tutor utilization: average booked hours per active tutor per week. Low utilization predicts
tutor churn (a tutor earning too little leaves the platform) well before it shows up in the
completed-session count. - Repeat booking rate: the percent of students who book a second session within 30 days of
their first. This is the leading indicator that today's completed sessions will still be
happening next month, since a marketplace that only ever acquires new demand and never retains it
is not sustainable.
Worked example
Say the marketplace has 500 active tutors and 2,000 active students, and this week produces 1,200
completed sessions at an average price of $30, for $36,000 in GMV. Now suppose the team raises the
average price to $36 per session and completed sessions drop 10% to 1,080 because price-sensitive
students book less. New GMV is 1,080 x $36 = $38,880, higher than before, even though the
marketplace just got measurably less healthy: fewer students got tutored, and fill rate and repeat
booking almost certainly moved down alongside it. A team optimizing for GMV would read this as
success. A team watching completed sessions, fill rate, and repeat booking rate would catch the
regression immediately. That is the concrete case for rejecting GMV as the north star, not just an
abstract preference.
Trade-offs and pitfalls
Picking a metric that lives on only one side of the marketplace (say, tutor hours worked) biases
the team to over-invest in supply even if demand is the real constraint, or vice versa; completed
sessions forces both sides to be healthy simultaneously. The metric is also gameable if you are not
careful: a bad actor could mark sessions "completed" without real teaching happening, so it needs a
verification signal (a post-session confirmation or rating flow) and a no-show/cancellation rate
tracked as its own guardrail. Finally, this exact north star will need to evolve as the marketplace
matures. Early on, completed sessions is really a liquidity proxy; at scale, a business might add a
second layer (session quality, or GMV per active tutor) without demoting completed sessions from the
top of the tree.
You and a designer disagree about where to spend a fixed slice of engineering time: a prominent UI callout that would increase a feature's discoverability, or backend work that would make the feature itself noticeably better. How would you decide between the two, and what would you look at quickly to avoid just guessing?
Sample Answer
Bottom line
Don't debate this on taste. Split it into two separate questions, get a cheap, fast read on each from data you likely already have, and only then compare the two using a rough return-on-investment (ROI, value created per unit of effort spent) estimate.
How to decide
-
Separate the two failure modes the callout and the backend work each fix:
- The callout fixes a discoverability problem: people never find or try the feature.
- The backend work fixes a quality problem: people find it, but it disappoints them once they use it.
A feature can suffer from either, both, or neither, and the fix only works if it targets the actual bottleneck.
-
Pull cheap signals before guessing, all things you can usually get from existing analytics and support data in a day or two, with no new code:
- Funnel data: of all active users, what percent ever open the feature (tells you if discoverability is the bottleneck), and of those who open it, what percent complete it or return to it (tells you if quality is the bottleneck).
- A quick read of 30 to 50 recent support tickets or session recordings tagged to the feature, to see whether drop-off looks like people getting lost in the navigation versus people hitting an actual defect or a "this isn't good enough" moment.
-
Turn those signals into a rough ROI comparison:
ROI=engineering costexpected incremental value
Expected incremental value is roughly (number of additional or improved conversions the change would produce) times (value per conversion). Engineering cost is the effort, in developer-weeks, for each option. You don't need precision here, you need the two ratios to be different enough to point clearly one way.
- If the two ROI estimates are close, don't force a binary call: a small, fast experiment for each (a feature-flagged callout to a slice of traffic, and an equally-scoped fix for the top quality complaint) settles it with real data faster than another round of debate.
Worked example
Say a trip-planning app has 100,000 monthly active users, and a cost-splitting feature is buried three taps deep. Funnel data shows only 4% of users ever open it (4,000 people), and of those, 55% complete it (2,200). That split alone is informative: 96% of users never even see the feature, which points at discoverability as the bigger-volume bottleneck.
Suppose a similar UI callout shipped in another part of the product previously lifted a comparable feature's open rate from 4% to 9%, a real precedent to anchor the estimate on. Applied here, that's roughly 5,000 additional openers a month, and if completion rate holds at 55%, about 2,750 additional completions a month.
Now the backend option: support tickets show a specific failure (splits don't handle uneven groups well) affecting a chunk of users who try to complete the flow. Suppose fixing it would lift completion rate among the existing 4,000 openers from 55% to 70%, that is 4,000 x 0.15 = 600 additional completions a month, from the same audience size that already exists today.
At comparable engineering cost (call it one sprint for either option), the callout produces roughly 4 to 5 times more additional completions by volume. On pure ROI, the callout wins here, but that number alone isn't the whole decision.
Trade-offs and pitfalls
- A callout is more visible and more demoable, which biases people toward preferring it in a review even when the numbers say otherwise. Don't let "which one is easier to show off" substitute for "which one the data supports."
- If the feature genuinely has a quality problem, driving more traffic into it with a louder callout can backfire: more people try it, more people bounce or leave frustrated, which can hurt overall app ratings even as the "open rate" metric goes up. When completion quality is clearly broken, it's often safer to fix or at least de-risk that first before spending a scarce, high-visibility UI slot to drive volume into it.
- A prominent callout also has an opportunity cost: that UI real estate could be promoting something else, so the comparison isn't really "callout vs. backend," it's "this feature's callout vs. everything else that wants that same spot."
- This doesn't have to be all-or-nothing. If both problems are real and the effort is divisible, a smaller version of each (a modest UI nudge plus a scoped fix for the worst quality complaint) can capture most of the value without betting the whole sprint on one lever.
Unlock Full Question Bank
Get access to all 21 Product Sense and Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.