Direct answer
Rank by estimated customer pain per unit of effort, and build pain from independent evidence: how many people hit the problem (reach), how badly (severity), and how much work the fix takes (effort). Loudness is controlled by counting distinct affected users from behaviour data instead of counting messages, and the ranking is checked against qualitative evidence from interviews, support transcripts and Net Promoter Score (NPS) comments. NPS is a survey asking customers how likely they are to recommend you on a 0 to 10 scale, with a free-text reason.
Step 1: deduplicate and tag
Cluster the 150 reports into distinct underlying problems, then tag each with the area, the symptom and a severity class: 3 = loses data or blocks a core task, 2 = slows a task but a workaround exists, 1 = annoyance or cosmetic.
Step 2: measure reach from behaviour, not volume
Use telemetry (usage data the product records automatically): error logs, funnel drop-offs (the share of users who quit at each step of a multi-step flow such as checkout), usage events. Tickets overcount persistent complainers and undercount quiet sufferers. One account sending nine emails is one account. Upvotes (votes from users on a public request) are volume too, so use them only as a hint, not as reach.
Step 3: score
text
priority = reach (% of active users) x severity (1 to 3) / effort (engineer-days)
Enter reach as a plain number of percentage points (12% goes in as 12, 0.4% as 0.4, never 0.12 or 0.004). Plain English: more people, hurt more, for less work, ranks higher. Treat scores in the same ballpark as ties and break them with judgement.
Step 4: let qualitative evidence change the ranking (illustrative numbers)
Tag each interview, NPS comment and transcript to the same issue clusters, with explicit rules: a transcript showing lost work or abandonment raises severity to 3; one showing a trivial workaround lowers it. Record the supporting quote so the change can be audited.
| Issue | Reach | Severity (tickets only) | Effort | Score | Severity after transcripts | Score after |
|---|
| A: dashboard colour bug (one account, nine emails) | 0.4% | 1 | 5 | 0.08 | 1 | 0.08 |
| B: export fails on large files | 12% | 2 | 8 | 3.0 | 3 | 4.5 |
| C: confusing sidebar label (30 upvotes) | 4% | 1 | 1 | 4.0 | 1 | 4.0 |
Before the transcripts, C outranks B. Reading five transcripts shows users redoing hours of work after failed exports, so B moves to severity 3 and overtakes C. The loudest report, A, is cosmetic (severity 1 on the Step 1 scale) and stays last.
Conflicting requests from strategic customers
Strategic customers are the accounts the company has named as most important to win or keep. Do not override the score silently. Reserve a small explicit quota, agreed in advance, for documented churn risk (a customer likely to cancel), and judge those items on evidence (renewal date, stated blocker). Everything else competes on the score.
Checking the ranking reflects real pain
- Spot-check with a few users from the top and the bottom of the list.
- After fixing the top items, confirm tickets and funnel drop-offs for those issues fell.
- Compare who is complaining with who is affected in the telemetry.
Variation: complaints about model quality (ML-powered features)
The same scoring still applies. For a feature powered by a machine learning model, add these checks in the first 24 hours: (1) did complaints start after a model, data or prompt release (a change to the instructions given to a language model)? (2) replay saved examples to tell a regression (something that used to work and now does not) from an old weakness; (3) compare today's inputs with the training period (data drift: real inputs shift away from what the model learned); (4) rule out the product around the model (truncated text, meaning input cut off before it reaches the model; timeouts; a stale cache, meaning an old saved answer served instead of a fresh one) before blaming the model; (5) check whether complaints cluster in a segment.
Example (illustrative): 40 tickets say "the assistant summaries are wrong since last week". Check 1 shows all 40 began within two days of a prompt release. Check 2 replays 50 saved inputs and finds the new prompt drops the last paragraph of long documents, so it is a regression. Check 4 shows the same truncation in a log. The fix is to roll back the prompt, and check 2 is what caught it.
Variation: player-value scoring (games)
Weight reach by player value, for example new players (whose retention, meaning whether they keep playing, decides the game's future) versus long-tenured or paying players, as a capped multiplier so it does not recreate a revenue-ranked skew.
Example (illustrative): a bug hits 3% of players, severity 2, effort 4, so the base score is 3 x 2 / 4 = 1.5. If most affected players are paying, apply a multiplier of 1.5 (the cap): 1.5 x 1.5 = 2.25. Uncapped, a spending-weighted multiplier of 5 would give 7.5 and let paying players outrank every other issue, which is the skew the cap prevents.