Research Collaboration and Stakeholder Management Questions
Working across teams and the wider community: partnering with engineers, product, and academic collaborators, engaging stakeholders, prioritizing competing requests and hypotheses, and securing resources and support. Covers negotiating scope and timelines, keeping expectations calibrated through negative or inconclusive results, agreeing who owns what at handoffs, setting terms for joint work with outside labs, and re-planning or deciding when good enough is enough when competing work lands. Interviewers look for researchers who create leverage through collaboration rather than working in isolation. Excludes generic prioritization frameworks, stakeholder mapping on delivery projects, earning credibility with skeptical audiences, setting research-organization strategy, the mechanics of shipping a research prototype, and user-research study design.
Product wants to optimize average session length. Your research suggests day-30 retention tracks long-term value better. How do you resolve the disagreement with the team, and what evidence would you bring?
Sample Answer
Direct answer
I would reframe it from "my metric versus yours" to "which number best predicts the outcome we all want", then settle it with evidence both sides agreed to in advance. Day-30 retention (the share of new users who are still active 30 days after they start) is closer to long-term value, while average session length is easy to move in ways that may hurt users. I would not argue from authority. I would propose a test where both are measured, with a decision rule we write down first.
Steps
-
Agree on the goal. Ask product what session length is a proxy for (engagement? revenue? satisfaction?). Usually the real goal is long-term value, which is what retention approximates.
-
Show how the two relate in existing data. Compare cohorts (groups of users who started in the same period): do users with longer early sessions retain more? Careful: heavier users naturally have both, so this is correlation. It shows the metrics are linked, not that raising one raises the other.
Also show the link the whole argument rests on: in past cohorts, did day-30 retention predict later value (for example revenue or active days at month 6 or 12) better than average session length did? Without that, "retention tracks long-term value" is an opinion, not evidence. Compare, say, the correlation of each metric with month-12 value across past launches.
-
Explain the failure mode. Session length can rise because the product got slower or more confusing (people take longer to find things), or because of notification nudges that bring people back once and burn them out. This is the Goodhart problem: a measure that becomes a target stops measuring what it was chosen to measure.
-
Propose an A/B test (randomised experiment) with session length as a diagnostic and day-30 retention as the deciding metric, plus guardrails (metrics that must not get worse, with a limit agreed in advance, e.g. support complaints must not rise more than 10% and uninstall rate must not rise at all beyond noise).
-
Offer a faster proxy if 30 days is too slow: check whether day-7 retention predicted day-30 in past launches, and use it as an early read.
Worked example (illustrative numbers)
A feature variant is tested on 10,000 users per arm. Average session goes from 12 to 14 minutes (about +17%). Day-30 retention goes from 40% to 38%.
The retention gap is 2 percentage points. The standard error (how much a measured gap would typically wander by chance if you reran the test) comes from adding the sampling variance of each arm, p x (1 - p) / n, then taking the square root:
SE = sqrt(0.40 x 0.60 / 10,000 + 0.38 x 0.62 / 10,000) = sqrt(0.0000476) = 0.0069, about 0.69 points
A 95% interval is the gap plus or minus 1.96 standard errors, the range that would contain the true gap in about 95 of 100 repeats of the test. Here 1.96 x 0.69 = 1.35, so the interval is 2 - 1.35 to 2 + 1.35, roughly 0.6 to 3.4 points (percentage points, meaning the 40% and 38% are subtracted directly). Zero is outside the interval, so retention probably fell, even though sessions got longer. Product's metric says ship; the retention metric says the extra minutes cost users.
The decision rule written before the test makes this concrete (illustrative): ship only if the whole 95% interval for the retention change sits above -0.5 points, meaning we can rule out a loss bigger than half a point, and every guardrail holds. The interval here is about -3.4 to -0.6 points, so the rule says do not ship. The test size was adequate for this question: detecting a 2-point difference from a 40% baseline with 80% power needs about 9,300 users per arm, and this test had 10,000. Presenting both side by side, with the interval, makes the decision a shared reading of the data, not a debate.
Trade-offs and pitfalls
- Do not just call the metric "bad". Session length is a legitimate diagnostic; the dispute is about what we optimise.
- Retention is slow and noisy, and small tests may not detect small differences; say so rather than overclaim.
- Retention can also be gamed (reminder spam), so pair it with a quality guardrail.
- What would change my mind: if the data show retention does not depend on early session length and our product is read-heavy where time-on-task is the real value, I would accept session length as a secondary goal.
What I would say to product
"Session length went up 17%, and I am glad it moved. But the people in that arm were 2 points less likely to be here at day 30, and that gap is bigger than chance. Can we agree that retention decides this, and keep session length as a signal that helps us see why?"
What do you do early in research planning to make sure the product and engineering partners will actually adopt the result? Give an example where this changed what you built or how you defined done.
Sample Answer
Direct answer
Early in planning I define "adopted" with the partners who will have to use the result, before I choose methods. I ask them what they would need to see to put it into production, what constraints it has to fit, and who will own it afterward. Then I design to those answers. That conversation changes what gets built and gives the project a definition of done that includes the partners' needs, not only a model score.
Structured elaboration
Questions I ask in the first weeks:
- Decision and owner: what decision will this inform, and who owns the system afterward?
- Constraints: latency, cost per request, interpretability, data availability at serving time, review requirements.
- Evidence bar: what result would convince you to ship, and what would convince you not to?
- Integration: what format and interfaces do you need, and who will review them?
- Checkpoints: when do we look at an early version together?
I turn the answers into a short written definition of done shared by everyone.
Worked example (a story to adapt)
I was asked to improve a classifier that flags risky transactions. My plan was to maximize offline accuracy. In the first meeting the engineering lead told me the serving system (the live software that runs the model on each request) had a strict response-time budget (the maximum latency, meaning wait time per request, it allows; here 50 ms) and could not call the heavy features I intended to use, and the operations team said reviewers would only trust alerts that came with a reason (interpretability: a human-readable explanation of why each alert fired).
- I redefined done as: beats the current rules on the agreed review-precision measure (of the alerts the system raises, the share later confirmed as truly risky by investigators; the rules scored 55%), fits the 50 ms response-time budget, and gives a short reason code (a label such as "amount far above this account's usual") per alert.
- I dropped the most expensive features and built a smaller model, plus a simple explanation field.
- I shared an early version with the operations team after four weeks and changed the reason wording based on their feedback.
The offline score (review precision replayed on saved historical transactions with their confirmed outcomes, the agreed measure; it cannot show whether reviewers would trust the reason codes, so it is re-checked live in the reviewers' queue after launch) was somewhat lower than the bigger model's would have been (illustrative numbers: 62% review precision for the small model versus 66% for the big one, and about 35 ms versus 120 ms per request, so only the small model fit the 50 ms budget), and the engineers shipped it within one planning cycle because it already met their constraints. Had I built the larger model first, the likely result was a strong score that could not be deployed.
What changed in what I built: feature set, model size and an output the reviewers could use. What changed in done: from a score to a score, a latency fit and a reason code.
Trade-offs and pitfalls
- Asking partners for requirements is not enough; ask what would make them say no.
- Constraints discovered late are the usual reason strong research never ships.
- Keep room for the research question: partners can constrain the solution without defining the science.
Product hands you a goal: lift 30-day retention by 5%. How do you turn that into a research scope you can commit to, and who do you need in the room to do it?
Sample Answer
Direct answer
I would not accept "lift 30-day retention by 5%" as a research scope until it is pinned down. I would sit with product, data and engineering to define the metric, the baseline, the population (which users the goal covers), the lever, and how we will measure success, then translate the goal into a scoped set of hypotheses and a commitment on what research can deliver by when. I would commit to the experiments and learning, and negotiate the outcome target separately.
Structured elaboration
- Clarify the goal. "5%" can mean different things. 30-day retention here means the share of new users who are still active 30 days after signing up. Ask: is the lift relative or absolute, which user group, which time window, and what is the current baseline?
- Check detectability. Detectability means whether a test can reliably tell a real lift from random noise. Can we even see an effect of that size with the traffic we have?
- Find the levers. A lever is something we can change that plausibly moves the metric. Break retention into where users drop off and which of those a model can influence (recommendations, notifications, onboarding), then pick hypotheses.
- Define what research commits to: a measured test of the top hypotheses by a date, not a guaranteed lift.
- Agree what "done" means and who acts on a success.
Worked example (illustrative numbers)
Say baseline 30-day retention is 40%. A 5% relative lift means 40% x 1.05 = 42%, a 2-point change. A 5-point absolute lift means 45%. The second is far harder. Which one product means changes the scope entirely.
The number also drives feasibility. Two terms first. Power of 80% means that if the lift is real, the test will detect it 80% of the time. 5% two-sided significance means we accept a 5% chance of calling a lift real when it is only noise. With those settings, the users needed per group (control and treatment) are:
n per group = (1.96 + 0.84)^2 x [p1(1-p1) + p2(1-p2)] / (p2 - p1)^2
Here 1.96 and 0.84 are the standard constants for 5% two-sided significance and 80% power, and p1, p2 are the two retention rates. The intuition: the smaller the gap, the more users you need, and the need grows with the square of how small the gap is.
- 40% versus 42%: (2.80)^2 = 7.84 (7.85 if the constants are left unrounded, the value used below). The spread term is 0.40 x 0.60 + 0.42 x 0.58 = 0.2400 + 0.2436 = 0.4836. The gap squared is 0.02^2 = 0.0004. So n = 7.85 x 0.4836 / 0.0004, about 9,500 per group, roughly 19,000 in total.
- 40% versus 45%: the spread term is 0.2400 + 0.2475 = 0.4875 and the gap squared is 0.05^2 = 0.0025. So n = 7.85 x 0.4875 / 0.0025, about 1,500 per group.
If the product only gets a few thousand new users in a test window, the 2-point lift cannot be detected in one test, and I say so before committing. Retention at day 30 also needs 30 days of follow-up per user, so the test cannot read out until 30 days after the last user enrolls; that waiting time belongs in the commitment too.
Who needs to be in the room
- Product manager: owns the goal, the definition of retention and what the lever could be.
- Data analyst or data scientist: owns the baseline and the metric pipeline.
- Engineers: own what can ship and the cost of the lever.
- Experimentation owner: owns test design and traffic allocation.
- Design or growth if onboarding or messaging is a lever.
Trade-offs and pitfalls
- Committing to the outcome lets luck decide whether the quarter was a success; commit to learning and a decision.
- A vague goal lets every stakeholder hear their own meaning of success.
- Check that the metric cannot be moved by something that harms the product, such as aggressive notifications. Add a guardrail.
A product manager asks for an ML feature in two weeks to hit a launch date. Your honest estimate for rigorous evaluation is twelve weeks. How do you respond, and what would you offer instead?
Sample Answer
Direct answer
I would not say yes or no to "two weeks". I would say: "I can give you something real in two weeks, but not a validated result. Here are three options and what each one risks." Then I let the product manager (PM) choose with the evidence in front of them. Rigorous evaluation means testing the model on data it never saw (held-out data), against a simple baseline (the current rules or a trivial model it must beat), across the user groups that matter (subgroups), before anyone relies on it. Twelve weeks is my estimate for doing that properly, and I will say what is in it.
Structured elaboration
- Find the real constraint. Ask what the launch date is tied to (a marketing event, a contract, a quarter-end) and what happens if the feature is late versus wrong. A fixed date with a soft scope is very different from a hard promise to a customer.
- Unpack the twelve weeks. Show what the time buys: data collection and labeling, a baseline, offline evaluation on held-out data, checks on subgroups and failure cases, then an online test. Ask the PM which pieces they believe are optional. That turns "researcher is slow" into a shared view of risk.
- Offer options, not a refusal. Each option states what ships, what is still unproven, and how we limit damage if the model is bad.
- Recommend one. I recommend the staged option below, because it hits the date and keeps the evidence bar.
- Say what would change my call. If a quick look shows a simple rules baseline already meets the need, or the feature is easy to switch off and low-stakes, I would accept a shorter evaluation.
Worked example (illustrative numbers)
| Option | What ships at week 2 | What is still unproven | Total time to full confidence |
|---|---|---|---|
| A. Ship as asked | Model to all users | Everything beyond a quick offline check (scoring saved data, with no real users involved) | 12 weeks, but after exposure |
| B. Staged (recommended) | Feature behind a flag (an on/off control that shows the feature to only chosen users), internal users plus a small slice of traffic, with a simple fallback (the old behavior users get if the model is off) and a clear kill switch (one action that turns it off for everyone) | Subgroup behavior, long-run effect | 2 + 4 + 6 = 12 weeks |
| C. Move the date | Nothing at week 2 | Nothing | 12 weeks |
In option B, weeks 1 to 2 build a baseline and a first model plus the flag and fallback. Weeks 3 to 6 run the held-out and subgroup evaluation while the small slice gives early signal. Weeks 7 to 12 run the full online test (a live experiment comparing users who get the model with users who do not) and the go or no-go review. The PM gets a launch-day story ("early access") without the whole user base absorbing a possibly bad model. A guardrail metric (a number that must not get worse, such as error complaints or latency) defines when the flag is turned off. For example, I would write down in advance: "if error complaints rise by more than 20% over the control group, or median latency worsens by more than 100 ms, we switch the flag off." These thresholds are illustrative; the point is that they are agreed before launch.
Trade-offs and pitfalls
- Do not pad the estimate to look careful; tie every week to a risk it removes.
- Do not agree to two weeks and quietly cut evaluation. The cost shows up later as a rollback and lost trust.
- Do not argue in statistics jargon. Speak in consequences: "if this is wrong, who is affected and how fast can we undo it?"
- Close the loop in writing: what was agreed, what is unproven, and the date of the next checkpoint.
What I would say to the PM
"I can put something real in front of internal users and a small slice of customers in two weeks, and it will be labeled early access. I cannot tell you it is validated, because I have not yet checked it on data it has never seen or on every user group. If it misbehaves we turn it off with one switch and users see the old experience. Which of these three options fits what the date is protecting?"
How do you decide when a good-enough research result is the right call instead of chasing the optimal one? Walk me through a situation in an industry lab where you would stop early.
Sample Answer
Direct answer
A good-enough result is the right call when the decision it feeds can already be made, and the expected extra value of more work is smaller than its cost, including the other work you give up. In an industry lab, research exists to inform a product or strategy decision, so the stopping rule is tied to that decision, not to the best number you could possibly reach.
How I decide
- Name the decision and its bar. What does the product team need to see to act? For example, "the new ranking model must beat the current one by enough to justify the engineering cost."
- Check whether you are past the bar with enough confidence that more tuning would not change the call.
- Estimate the marginal gain. Look at the learning curve (a plot of the metric against effort or data, which flattens when more work stops helping): if the last two weeks of effort improved the metric by an amount smaller than run-to-run noise, more effort is mostly chasing noise.
- Count the opportunity cost. The same people-weeks could answer another question.
Worked example (illustrative)
A team compares a new ranking model to production. Product's bar: a win of at least 1.0 point on the offline quality metric (the score measured on saved data, not real users) before running a live test. After 4 weeks the baseline scores 41.2 and the new model 43.1, a lead of 1.9 points. Repeat runs with different random starting points (random seeds) of the same model vary by a standard deviation of about 0.4 points, and each model was run with 3 seeds, so the 41.2 and 43.1 are 3-seed averages. Two more weeks of tuning would, by past experience, add about 0.2 points, which is less than that seed-to-seed noise (all numbers illustrative).
Is 1.9 clearly past the bar of 1.0? The margin over the bar is 0.9 points. One run of each model would differ by a standard deviation of sqrt(0.4^2 + 0.4^2) = 0.57 points, so 0.9 is only about 1.6 of those, which would not be convincing. Averaging 3 seeds per model shrinks it to 0.57 / sqrt(3) = 0.33 points, so 0.9 is about 2.8 of those, enough to justify a live test (rough check, assuming runs vary independently).
Decision: stop. Write up the result and its limits, hand it to engineering for a live test (an experiment on real traffic, which gives better evidence than more offline tuning), and move the 2 freed weeks to the next question on the list. The live test, not extra tuning, is the more valuable use of time.
When I would not stop early
- The result is near the bar but the confidence interval (the range of values the true result could plausibly take) crosses it. Here a lead of 1.9 against a bar of 1.0, with 3-seed noise of about 0.33, is clear enough to proceed; a lead of 1.1 would sit only 0.1 above the bar, about 0.3 noise units, and would not be.
- The error is concentrated in a user group that matters (for example a safety-relevant slice), so a good average hides a bad slice (a subset of users or inputs, such as one language or device).
- A wrong "good enough" would be expensive to reverse, such as a decision that locks in architecture.
Trade-offs and pitfalls
Stopping tuning is not stopping checking: the live test is the safeguard, and it is cheap relative to two more weeks of offline tuning. Stopping early can leave value on the table; polishing can burn a quarter for a gain nobody can measure. State the assumption you are stopping on and what you would check later, so the call can be reversed cheaply.
Unlock Full Question Bank
Get access to all 11 Research Collaboration and Stakeholder Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.