Customer and User Obsession Questions
Reasoning from the customer or end user inward when making decisions in any role: building empathy for users, including buyers who are not the people who use the product; identifying and prioritizing customer pain points and workarounds; collecting and acting on customer feedback, including how representative it is, how it compares with usage data or satisfaction scores, and how it reaches the people who can act; bringing the voice of the customer into roadmap and technical trade-offs; advocating for users against internal or deadline pressure; and balancing customer needs against business goals and engineering constraints, such as a large account's bespoke request, power users versus everyday users, or retiring a feature. Includes stories of changing course because of customers or recovering after letting one down. Assesses whether a candidate starts from the customer's problem rather than from features or technology. Study design and fieldwork, research synthesis, personas and journey maps, analytics and experimentation mechanics, reliability engineering, customer-success operations, market and competitive research methods, and live handling of angry customers are covered elsewhere.
You want in-product feedback from users without annoying them or hearing only from the angriest ones. How would you design when and how you ask, and what would you do with what comes back?
Sample Answer
Direct answer
I would ask rarely, at the right moment, of a sample, and give people an always-available way to volunteer feedback, then compare the two so loud, angry voices do not stand for everyone. What comes back is tagged, linked to usage context, counted by theme, and answered, so people see that it matters.
Design: when to ask
- At a natural completion point, such as right after a user finishes a task, not mid-task. Never interrupt checkout, error recovery or a first session.
- Eligibility rules (the conditions a user must meet before being shown the prompt): only users who have done the relevant thing a few times, and suppress (hide the prompt from) anyone who dismissed a prompt recently.
- Frequency cap (a limit on how often the same user is asked): one prompt per user per fixed window (for example 60 days, a design choice to tune).
- Sample, do not blast: show the prompt to a random slice of eligible users.
Illustrative sizing: with 20,000 eligible users a week and a 10% sample, 2,000 see the prompt. At an assumed 10% response rate (responses divided by users shown the prompt), that is about 200 answers a week.
Design: how to ask
- One quick question (a rating or thumbs up or down) plus an optional free-text box.
- Dismissible in one tap, and the dismissal is remembered.
- Capture context automatically (page, app version, plan, session ID) with clear consent, so users do not have to describe where they were.
Example (illustrative): right after a user exports a report successfully, a small card appears at the bottom of the screen, not a blocking pop-up. It reads "How was exporting that report?" with thumbs up and thumbs down. A tap opens an optional box, "Anything we should know?", and a "Not now" link closes the card for 60 days. The question is about the task just finished, not about the product in general.
Avoiding the angriest-only problem
A permanent "send feedback" button draws people with strong feelings, usually bad ones. A sampled prompt at a moment of success reaches satisfied and neutral users too. Keep both channels, label them in reports, and compare responders with the user base (plan, tenure, usage) to see who is missing.
Variants
| Context | Flow |
|---|---|
| Machine-learning (ML) product with a wrong output | "Report a problem" next to the result, with the input and output attached (privacy-redacted, meaning names, emails and other personal data are removed before storing), a short reason list and optional note |
| AI assistant | Flag (thumbs down with reason), correct (user supplies the right answer), escalate (hand to a human with conversation context) |
| Mobile app | Same rules, but app-store ratings use the platform's own review prompt, which limits itself. Google Play states it enforces a time-bound quota (calling the review flow more than once in under about a month may show nothing) and advises against triggering it from a button, so route users who volunteer feedback to the store page instead. Apple also rate-limits its system prompt; check the current documentation for both before building |
For AI assistants, route by severity: harmful or high-stakes errors escalate immediately, ordinary quality complaints go into the weekly theme review.
What to do with the responses
- Tag by theme, attach usage context, and count distinct users per theme.
- Weekly review: urgent issues to engineering, repeated themes to the backlog, how-to questions to docs.
- Reply to people who left contact details, and tell them when it ships.
- Track the health of the channel itself: response rate by segment and share of respondents who are negative versus the base's behaviour.
Pitfalls
Over-asking trains users to dismiss the prompt. Treating the average rating as a verdict hides who answered. Collecting feedback and never replying teaches users to stop giving it.
In many B2B products the person who pays is not the person who uses it. Tell me how you would handle a case where the buyer's requirements pull against what the end users need.
Sample Answer
Direct answer
Treat the conflict as a hidden disagreement about outcomes, not features. In business-to-business (B2B) products the economic buyer (the person who controls the budget) usually asks for a requirement that protects some goal such as control, compliance or visibility, while end users need to get daily work done quickly. A product users avoid does not get renewed, and a product the buyer cannot justify does not get bought. So I find the buyer's real goal, measure the actual cost to users, and look for a design that delivers the goal with the least user burden. Only if none exists do I escalate it as an explicit trade-off.
Steps
- Name the conflict in one sentence with both sides. "The buyer requires X; users report Y."
- Ask the buyer why. "What would you be able to show or prevent if this were in place?" Requirements are often proxies: a stand-in request for the real goal behind it. "Mandatory fields" may really mean "an audit trail".
- Quantify user cost. Extra time per task, share of users affected, whether they already avoid the step, adoption trend. For example: 8 extra seconds per task, 15 tasks per shift is 2 minutes per user per shift; 120 of 150 users affected; 30% of entries already skipped or filled with the default.
- Lay out options and who pays for each:
| Option | Buyer gets | Users pay | Fits when |
|---|---|---|---|
| Deliver the goal another way (defaults, passive capture meaning data recorded automatically as a side effect of normal work, report from existing data) | The outcome | Little | The requirement is a proxy |
| Role-based rules (strict only for roles that need it; for example an approval step only for purchases over $5,000) | Control where it matters | Only some users | The need applies to a subset |
| Admin setting, tunable or off by default | Control | Depends on setting | Buyer segment values it |
| Phase it (light now, stricter after adoption) | Partial, on a date | Gradual | Deal timeline versus workflow |
| Decline or descope (cut the request down or drop it) | Nothing if dropped; a reduced version if descoped | Nothing if dropped; less burden if descoped | It is a preference that harms a core flow |
- Decide with both in the room. Show the buyer the usage evidence; show users the reason. A hard requirement (legal, regulatory, written into the contract) outranks preference; a preference loses to demonstrated workflow harm.
Worked example (illustrative)
An operations director buying a warehouse system requires pickers to enter a reason code every time an item is short, to reduce shrinkage reports (shrinkage is stock lost or unaccounted for, such as theft or miscounts). A reason code is a short label from a fixed list, such as "damaged" or "miscounted". Pickers hit shorts many times a shift, the extra typing slows them, and they start choosing the first code in the list just to move on, so the data becomes unreliable. Resolution: preselect the most likely reason from scan context (the scan trail: the record of what was scanned, where and when), ask only when the situation is unusual, and give the director a weekly report built from the scan trail. The buyer gets trustworthy data, pickers keep pace, and the requirement's real goal is better served than by the original rule.
Trade-offs and pitfalls
- Siding with whoever signs the contract buys the deal and loses the renewal. Siding with users while ignoring the buyer ignores the pressure to justify spend.
- Saying yes to a custom version for one buyer creates a fork (a separate customer-specific version of the product) you maintain for years. Prefer a setting that other buyers can use. A segment is a group of similar customers (for example by size or industry), and a setting aimed at a segment spreads the cost of building it.
- What would flip my call: a regulatory or contractual requirement keeps the field, and the work shifts to minimising its burden. A buyer who will not move on a deal outside the segment you serve well may be one to decline.
- Solutions Architects can often bridge in this deal with configuration and rollout design; Product Managers own whether the pattern becomes a roadmap item.
You are about to invest in solving a problem you believe customers have, but you have only anecdotes so far. How would you check whether it is a real, widespread pain worth acting on, what would count as validation, and where could each check mislead you?
Sample Answer
Direct answer
Anecdotes tell you where to look, not how big it is. I would set the validation bar before collecting evidence, then use a mix of qualitative and quantitative checks (qualitative means what people say and do in depth, quantitative means counts across many), and treat each result as stronger or weaker according to what the person did, not what they said. A pain is worth acting on when it is frequent enough, severe enough, hits enough of the target group and is poorly served by what they do today.
Terms used below
- Problem interview: a conversation about the customer's past experience of a problem, not a pitch of your solution.
- Representative sample: respondents who resemble the whole target group, not just the friendly or vocal ones.
- Silent churners: customers who leave without ever complaining, so they never appear in tickets.
- Prevalence: how common the problem is across the whole group.
- Smoke test: a cheap fake offer (such as a sign-up page for a feature that does not exist yet) to see who responds.
- Commitment test: asking for something real in return, such as time, money or an introduction.
- Leading question: a question that suggests the answer you want.
Set the bar first (examples to be tuned)
For instance: at least 6 of 8 interviewees from the target segment describe a specific recent incident without prompting, tickets mentioning the problem come from at least 3% of accounts (illustrative), and at least a quarter of surveyed users report it at least monthly, and some people have already paid in time, money or a workaround. Writing this down first stops us from reading support for our idea into whatever we find.
Six checks, what counts as validation, and where each misleads
| Check | Validation signal | How it can mislead |
|---|---|---|
| 1. Trace the anecdotes | They come from several unconnected sources | One loud source repeated, or sales hearing only unhappy people |
| 2. Problem interviews with a varied sample | Specific recent stories, asked as "last time this happened, what did you do?" | Leading questions, politeness, and a sample drawn from friendly customers |
| 3. Existing behavior data (support tickets, search terms, drop-off points) | The pain appears across many accounts, not a few | Only shows pain people already report; misses silent churners |
| 4. Short survey on frequency and severity | At least the bar you set (for example 25% of 200 or more randomly chosen users) reports both | Low response rates can skew toward people with strong feelings (self-selection), although the response rate alone does not prove bias, so compare respondents with the known make-up of your user base and follow up with non-responders where you can. Sample question: "In the last 30 days, how many times did you have to rebuild or re-check an invoice by hand?" with options 0, 1 to 2, 3 or more |
| 5. Cost and workaround evidence | People use spreadsheets, scripts or manual steps and can say what it costs | A cheap annoyance can still produce elaborate workarounds |
| 6. Commitment test (sign-up page, pilot request, price conversation) | People give something up: time, money, an introduction | Smoke tests measure interest in a pitch, not the underlying pain |
Leading versus good interview questions
- Leading: "Wouldn't it be great if you could find invoices instantly?" (it hands over the answer, and people politely agree).
- Good: "Tell me about the last time you needed an old invoice. What did you do first? What happened next? What did it cost you?" Follow with "Why?" or "Can you show me?" to turn a vague answer into a story.
Signal strength
Opinions are weakest, stated intent is better, observed behavior is stronger, and commitment is the strongest. For B2B (selling to businesses), test willingness to switch vendors: ask what they would stop using, who must approve, when the current contract ends and what migration costs them. "We would love that" is cheap; "send me a pilot agreement" is not.
Worked example (illustrative)
Three account managers say customers cannot find invoices. The team tags tickets and finds 120 of 4,000 last quarter, or 3.0% of tickets, mention invoices. Counting distinct accounts, as the bar requires, those tickets come from 96 of 2,000 accounts, or 4.8% (illustrative), so the account bar of 3% is cleared. The team also interviews 8 customers across three sizes. Six describe losing time at month end. One says finance already built a spreadsheet to track invoices. Result: real for a defined segment, not yet proven broad. The next step is a commitment check with the segment that feels it most.
Pitfall
Declaring validation after several friendly interviews. Small samples find themes, not prevalence.
Think of a moment when you learned something about a user's problem firsthand (an interview, a support call, shadowing) that you would not have seen in the data. What did you do with it?
Sample Answer
Direct answer
I would tell one specific story: where I watched a user, what I saw, why the data could not have shown it, and what I changed because of it. The best stories show the firsthand observation changed a decision, not just my opinion.
How to structure the story
- Setting: which channel (an interview, a support call, or shadowing, which means sitting in on a customer doing real work without steering them), and why you were there.
- The observation: what the user did or said, in concrete detail.
- Why the data missed it: the gap, usually that the behavior happened outside the product or the metric measured the wrong thing.
- What I did: verification with more users, a clear problem statement, and a change.
- Result and caution: what happened, and how I kept one visit from becoming a false certainty.
Worked example (an illustrative story)
Product analytics said the export step had a high success rate: people clicked export and got their file. In a call a customer-success colleague invited me to join, I watched a user export a report, open a spreadsheet, spend a long stretch reformatting columns, and paste a summary into a slide for their manager.
The data could not show this because it all happened outside the product, and "export succeeded" hid the real job: producing something a manager can read. I asked three more customers for a 20-minute screen share of their last export-to-report workflow. Two showed the same pattern, which I treated as a direction, not proof. I wrote the problem statement as "users need a manager-ready summary, not raw data" and brought a short clip to planning.
We shipped a small scheduled summary with the four columns people kept. The real test was behavior afterwards: whether those customers still reformatted by hand, which I checked in follow-up conversations, and whether they kept sharing the summary. Illustrative result: of the 12 customers who received the summary in the first month, 9 stopped reformatting by hand when I asked them about it three weeks later, and 7 were still forwarding the summary to their managers, so I kept it and added the fifth column the other three asked for. The caution: this is a dozen customers who were willing to talk to me, so I read it as encouraging, not as proof for the whole base.
Pitfalls
- A story where the data and the visit agreed has little to teach. Pick one where the visit added something.
- "I always talk to users" without a specific observation and a specific change sounds rehearsed.
- Over-reading one visit. Say how you checked it.
Usability testing shows a core flow confuses users, but leadership wants to ship for a revenue opportunity. How do you make the case to delay or change it, and what do you do if you lose?
Sample Answer
Direct answer
I would not open with "we should delay". I would make the case in revenue and risk terms, show options each priced in days, recommend the smallest change that protects the revenue date, and ask for a staged rollout with a pre-agreed trigger. If I lose, I commit to the decision, instrument it, and reopen it only on agreed data.
Step 1: make sure the finding is solid
Usability testing means watching real users attempt tasks. Check how many participants failed, on which task, and whether they were blocked or only slowed. Illustrative finding used below: 5 of 8 participants could not find where to enter the promotion code at checkout. Five failures out of eight is a strong signal but a small sample, so confirm with analytics (the funnel drop-off, meaning the share of users who quit at each step of the flow, at that step) or a quick unmoderated test (participants do the task alone, with their screen recorded and no facilitator, so you can add many more of them cheaply).
Step 2: translate it into leadership's currency
Leadership is optimizing for revenue, so express confusion as lost revenue. Illustrative numbers: 10,000 users start the flow each month, baseline completion (today's share of starters who finish) is 70%, and beta data suggests completion could fall to 62% (an estimate to be confirmed, not a measurement).
| Quantity | Calculation | Result |
|---|---|---|
| Completions lost per month | 10,000 x (0.70 - 0.62) | 800 |
| Value lost per month at $20 per completion | 800 x $20 | $16,000 |
| Cost of a one-week delay, if the promo is worth $48,000 a month | $48,000 x 7 / 30 | $11,200 |
Why the delay costs money rather than just moving it: the promo runs in a fixed window, so each day the launch slips is a day of promo revenue that does not happen later (illustrative: $48,000 a month, so 7/30 of it for a week). The two costs have different shapes. The delay is paid once ($11,200). The confusion is paid every month until it is fixed ($16,000 a month), so a one-week fix costs less than a single month of confusion and the gap widens each month the problem stays.
Step 3: the one-slide executive framing
One slide with the decision wanted as the title, a 20-second clip next to the number as evidence, three options with days and revenue effect, and a recommendation with a tripwire (a pre-agreed trigger that reverses or pauses the plan).
| Option | Effect |
|---|---|
| A. Ship as is | Date held, completion likely near 62% |
| B. Delay one week to fix | About $11,200 of delayed promo revenue |
| C. Ship on the date with a small copy and layout fix, to 10% of users first (a staged rollout) | Date held, and we measure |
What the fix in option C is: relabel the small "Have a code?" link as a visible "Promo code" field and place it above the pay button. It changes wording and position only, not payment logic, so it is cheap (illustrative: 2 to 3 days of work), which is why it fits inside the launch date. The 10% of users who get it are a cohort (a group of users followed together), and the other 90% stay on the old layout until the numbers say it is safe to widen.
Recommendation: C. Tripwire: pause the promo placement if completion stays below 65% across the first 300 starters in the 10% cohort, which separates the estimated 62% from the 70% baseline. It flips to B if the 10% cohort breaches that line. Why 300 starters and not three days: 10% of 10,000 monthly starters is about 1,000 a month, roughly 33 a day, so three days is only about 100 starters, where completion is measured to within roughly 4.8 points (standard error), wider than the 3-point gap between 62% and 65%, so a three-day read would fire or stay silent largely by chance. 300 starters (about nine days) cuts that to roughly 2.8 points, which makes it a screen, not a verdict. A firm read of 62% against the 65% line needs about 700 starters (about three weeks; this is a one-sided 95% read, where the 3-point gap equals 1.64 standard errors, and a conventional two-sided 95% read would need about 1,000 starters, roughly a month at 10%), so if the date allows, widen the cohort to 30% to get there in about a week. Assign the cohort at random so the comparison with the other 90% is fair.
If I lose
Commit and ship without sandbagging (quietly half-hearted execution that lets the decision you opposed fail). Put in the instrumentation (the tracking that records which steps users complete, so the result is visible), log the fix in the backlog with an owner and a date, share the data at the agreed checkpoint, and escalate only if the tripwire is hit on data everyone accepted. Do not say "I told you so".
The same method in other versions of this fight
The same move (price the harm in leadership's units, offer a smaller option) carries over, with illustrative numbers:
| Version | What I would bring or ask for |
|---|---|
| Promo feature vs stability | Price an incident against the promo gain. If there is an estimated 20% chance of a 6-hour outage at $5,000 of sales an hour, the expected cost is 0.20 x 6 x $5,000 = $6,000; ask for the top stability fix to be inside the promo scope |
| Retention bug vs cosmetic launch | If the bug causes 100 extra cancellations a month at $240 a year each, that is $24,000 of annual revenue lost every month it lives, while a cosmetic launch can slip a sprint (a fixed work period, often two weeks); ask to swap the order |
| Revenue widget hurting discoverability | Measure what it displaces (clicks to the key feature before and after); propose a placement test where half of users see the widget at the top and half lower, and compare clicks |
| Beta feedback negative while stakeholders want to expand | Tie expansion to exit criteria (conditions agreed before the beta, such as 7-day retention of at least 40% and under 5% of beta users reporting a blocker); share verbatim quotes with counts; expand one cohort at a time |
| Reliability-risk launch delay | State probability and impact of failure, name what evidence you need (a load test, which simulates many users at once to see whether the system holds, and a rollback plan, the written steps to undo the release quickly), and offer a feature flag (a switch that turns the new code on for chosen users only) |
Pitfalls
- An ultimatum ("it cannot ship") ends the conversation; options keep it going.
- Do not quote an estimated 62% as fact. Label it, then use the staged rollout to measure it.
- Delay is not the only lever. Changing scope, placement or rollout size often protects both goals.
Unlock Full Question Bank
Get access to all 22 Customer and User Obsession interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.