Research Collaboration and Stakeholder Management Questions
Working across teams and the wider community: partnering with engineers, product, and academic collaborators, engaging stakeholders, prioritizing competing requests and hypotheses, and securing resources and support. Covers negotiating scope and timelines, keeping expectations calibrated through negative or inconclusive results, agreeing who owns what at handoffs, setting terms for joint work with outside labs, and re-planning or deciding when good enough is enough when competing work lands. Interviewers look for researchers who create leverage through collaboration rather than working in isolation. Excludes generic prioritization frameworks, stakeholder mapping on delivery projects, earning credibility with skeptical audiences, setting research-organization strategy, the mechanics of shipping a research prototype, and user-research study design.
You have five plausible research hypotheses and bandwidth for two. How do you decide which to pursue first, and how would you defend that call to a skeptical product partner?
Sample Answer
Direct answer
I would score the five hypotheses on the same few questions, pick the best two that fit the capacity, and show the scoring to the product partner so they can challenge the inputs rather than the conclusion. The questions are: how much value if it works, how likely it is to work, how much it costs, and what we learn even if it fails. I say what would change my pick.
Structured elaboration
- Write each hypothesis as a testable claim with the metric it should move and the cheapest test that could prove it wrong.
- Estimate value (if true), probability, and cost with the partner in the room for value, and engineers and peers for cost. Probabilities are rough guesses; the point is that they are written down and arguable.
- Compute expected value (value times probability) and compare with cost. A high-upside idea with a low chance can lose to a modest, likely one.
- Check dependencies and learning. If one hypothesis unlocks the data or tooling the others need, or a failure would still teach us something, adjust.
- Pick two and park three with explicit revisit conditions.
Worked example (illustrative scores)
Value is in arbitrary "value points" agreed with the product partner, cost is researcher-weeks, and the capacity is 8 weeks.
| Hypothesis | Value if it works | Probability | Expected value | Cost (weeks) |
|---|---|---|---|---|
| A | 8 | 0.50 | 4.0 | 3 |
| B | 10 | 0.25 | 2.5 | 5 |
| C | 4 | 0.60 | 2.4 | 2 |
| D | 9 | 0.30 | 2.7 | 6 |
| E | 6 | 0.50 | 3.0 | 4 |
The best pair that fits 8 weeks is A and E: expected value 4.0 + 3.0 = 7.0 for 3 + 4 = 7 weeks. The next best pair, A and D, gives 6.7 but needs 9 weeks and does not fit.
Defending it to a skeptical partner: B has the largest upside (10), and the partner will notice. My answer: at a 0.25 chance, its expected value is 2.5, below E at 3.0. Then I invite them to challenge the input. If they have evidence that B's probability is 0.4, its expected value becomes 4.0, and A plus B (4.0 + 4.0 = 8.0, cost 3 + 5 = 8 weeks) beats A plus E. I would happily change my pick on that evidence, and I would ask what that evidence is.
Trade-offs and pitfalls
- The numbers are structured guesses, not facts. Present them that way and the debate moves to the inputs, which is where it belongs.
- Do not hide the parked ideas; state when they would come back (new data, a failed first pick).
- Avoid picking by whichever stakeholder is loudest or by what is most intellectually exciting.
Midway through a project, a competing paper appears that undercuts part of your approach. How do you decide whether to pivot, double down or merge ideas, and what evidence would change your mind?
Sample Answer
Direct answer
I would not decide on the day the paper appears. I would spend a short, fixed window (about a week) working out what exactly is undercut, then choose among three moves: pivot (change the question or method), double down (continue because your contribution still stands), or combine (incorporate their idea and keep your distinct piece). My default is combine or narrow when only part of the approach is hit, and pivot only when the core claim is gone and nothing distinctive remains.
How I decide
1. Classify the overlap. "Scooped" (someone publishes your result before you, so yours looks less new; a preprint is a paper posted publicly before peer review) is not one thing:
- Claim overlap: they show the same result you were about to show.
- Method overlap: they use a similar technique on a different problem.
- Counter-evidence: their results suggest your approach does not work, or works only in a narrower regime (a range of conditions, such as small models or one kind of data).
2. Test whether it really applies to you. Read it closely, then check setup differences (data, scale, metrics, baselines, compute). Reproduce their headline number (the single best result in their abstract) on one of your own settings if it is cheap. A competing paper often undercuts less than its abstract claims (a rule of thumb from experience, not a measured rate).
3. Ask what is still yours. Write one sentence for "what would be new after this paper exists". If you can write it, you have something to double down on or merge into.
4. Price the options in remaining time, cost already sunk (ignore it; only future cost matters) and value to the team's goals, not only to a publication.
Worked example (illustrative numbers)
A team is 5 months into a method for making fine-tuning cheaper and has 3 months left. A paper appears showing a similar technique matching full fine-tuning on one benchmark.
| Option | Remaining work | What stays distinctive | Call |
|---|---|---|---|
| Pivot | ~3 months to restart on a new question | Nothing carried over | Only if the core claim is gone |
| Double down | 3 months as planned | Your results on other tasks and a theory section | Fine if their setup is narrow |
| Merge | ~1 month to add their method as a baseline and component | Your analysis of when it fails plus combined method | Best if both ideas complement |
Suppose a one-week check (re-running their released code on your classification data, then on your long-form generation data, with your own baselines) shows their method works on classification but gives no gain on long-form generation, where you already see a gap. Decision: merge. Add their method as a baseline, narrow the claim to generation, and keep the schedule.
Be honest about how strong that week of evidence is. "No gain on long-form generation" from one run per setting could be noise, so run at least 3 seeds (different random starting points) per setting and compare the gap with the seed-to-seed spread: a difference smaller than the spread is "unproven", not "absent". The merge decision is cheap to reverse (about a month of work), which is why a modest amount of evidence is enough to act on, whereas a pivot (about 3 months, nothing carried over) deserves much stronger evidence.
Evidence that changes my mind
- Toward pivot: their result reproduces in your setting and beats yours with no regime where yours is better; a collaborator in the field confirms the community considers the question settled.
- Toward doubling down: your setting is outside their reported conditions, or their result fails to reproduce after you have matched their released configuration and ruled out a bug of your own (otherwise a failed reproduction tells you about your setup, not about their claim).
- Toward merge: ablations (experiments that remove one component at a time to see what each contributes) show their component and yours help in different places.
I would also tell my lead the decision and the evidence within the week, since a reframed project may change commitments.
Trade-offs and pitfalls
- Sunk cost: "we already spent five months" is not evidence.
- Panic pivoting: abandoning too fast wastes unique work; most overlap is partial.
- Ignoring it: reviewers will cite the paper, so not engaging with it weakens the work.
- Contacting the authors can open a collaboration, or lead to both papers citing each other as concurrent work (independent work done at the same time, so neither is treated as copying). Weigh that against revealing your unpublished details. Because concurrent work is commonly acknowledged openly, you can usually continue with a narrower claim instead of abandoning the project.
Product wants to optimize average session length. Your research suggests day-30 retention tracks long-term value better. How do you resolve the disagreement with the team, and what evidence would you bring?
Sample Answer
Direct answer
I would reframe it from "my metric versus yours" to "which number best predicts the outcome we all want", then settle it with evidence both sides agreed to in advance. Day-30 retention (the share of new users who are still active 30 days after they start) is closer to long-term value, while average session length is easy to move in ways that may hurt users. I would not argue from authority. I would propose a test where both are measured, with a decision rule we write down first.
Steps
-
Agree on the goal. Ask product what session length is a proxy for (engagement? revenue? satisfaction?). Usually the real goal is long-term value, which is what retention approximates.
-
Show how the two relate in existing data. Compare cohorts (groups of users who started in the same period): do users with longer early sessions retain more? Careful: heavier users naturally have both, so this is correlation. It shows the metrics are linked, not that raising one raises the other.
Also show the link the whole argument rests on: in past cohorts, did day-30 retention predict later value (for example revenue or active days at month 6 or 12) better than average session length did? Without that, "retention tracks long-term value" is an opinion, not evidence. Compare, say, the correlation of each metric with month-12 value across past launches.
-
Explain the failure mode. Session length can rise because the product got slower or more confusing (people take longer to find things), or because of notification nudges that bring people back once and burn them out. This is the Goodhart problem: a measure that becomes a target stops measuring what it was chosen to measure.
-
Propose an A/B test (randomised experiment) with session length as a diagnostic and day-30 retention as the deciding metric, plus guardrails (metrics that must not get worse, with a limit agreed in advance, e.g. support complaints must not rise more than 10% and uninstall rate must not rise at all beyond noise).
-
Offer a faster proxy if 30 days is too slow: check whether day-7 retention predicted day-30 in past launches, and use it as an early read.
Worked example (illustrative numbers)
A feature variant is tested on 10,000 users per arm. Average session goes from 12 to 14 minutes (about +17%). Day-30 retention goes from 40% to 38%.
The retention gap is 2 percentage points. The standard error (how much a measured gap would typically wander by chance if you reran the test) comes from adding the sampling variance of each arm, p x (1 - p) / n, then taking the square root:
SE = sqrt(0.40 x 0.60 / 10,000 + 0.38 x 0.62 / 10,000) = sqrt(0.0000476) = 0.0069, about 0.69 points
A 95% interval is the gap plus or minus 1.96 standard errors, the range that would contain the true gap in about 95 of 100 repeats of the test. Here 1.96 x 0.69 = 1.35, so the interval is 2 - 1.35 to 2 + 1.35, roughly 0.6 to 3.4 points (percentage points, meaning the 40% and 38% are subtracted directly). Zero is outside the interval, so retention probably fell, even though sessions got longer. Product's metric says ship; the retention metric says the extra minutes cost users.
The decision rule written before the test makes this concrete (illustrative): ship only if the whole 95% interval for the retention change sits above -0.5 points, meaning we can rule out a loss bigger than half a point, and every guardrail holds. The interval here is about -3.4 to -0.6 points, so the rule says do not ship. The test size was adequate for this question: detecting a 2-point difference from a 40% baseline with 80% power needs about 9,300 users per arm, and this test had 10,000. Presenting both side by side, with the interval, makes the decision a shared reading of the data, not a debate.
Trade-offs and pitfalls
- Do not just call the metric "bad". Session length is a legitimate diagnostic; the dispute is about what we optimise.
- Retention is slow and noisy, and small tests may not detect small differences; say so rather than overclaim.
- Retention can also be gamed (reminder spam), so pair it with a quality guardrail.
- What would change my mind: if the data show retention does not depend on early session length and our product is read-heavy where time-on-task is the real value, I would accept session length as a secondary goal.
What I would say to product
"Session length went up 17%, and I am glad it moved. But the people in that arm were 2 points less likely to be here at day 30, and that gap is bigger than chance. Can we agree that retention decides this, and keep session length as a signal that helps us see why?"
Tell me about a time your research came back negative or inconclusive after stakeholders had been expecting a win. What did you tell them, and how did you keep the collaboration healthy?
Sample Answer
Direct answer
I told stakeholders early, plainly, and with a recommendation attached. A negative or inconclusive result is still a result: it tells the team what not to build. The collaboration stays healthy when the person hears the news from me first, understands exactly what was and was not shown, and leaves the meeting with a next decision rather than a verdict on my work.
Structured elaboration (the story shape)
- Situation: the product team expected a win from my work, and I had committed to an evaluation plan up front.
- Task: report a result that did not meet their hopes without overselling or burying it.
- Actions: pre-wire, then present the evidence, then propose options.
- Result: the team made a clear decision and kept working with me.
- What I would do differently: set the "what would count as a win" bar even earlier.
Worked example (a story to adapt)
I was testing whether a new ranking model would improve click-through for a recommendations surface. The product manager and director had talked about it as nearly done. Our online A/B comparison (randomly splitting live users between the current model and the new one, so click-through could actually be measured for both) showed the new model was not better. Click-through (the share of shown items that get clicked) was 5.0% for the current model and 4.9% for the new one, a difference of -0.1 points. The confidence interval (the range the true difference could plausibly be in) ran from -0.5 to +0.3 points, wide enough that "slightly worse" and "slightly better" were both possible (illustrative numbers).
- I pre-wired (briefed the PM privately ahead of the wider meeting, so nobody was surprised in public). I messaged the PM privately the same day: "The result is not what we hoped. I want to walk you through it before the wider review so we can decide together how to present it."
- In the walkthrough I separated three things: what we know (no clear gain on the primary metric, the one measure we agreed in advance would decide success), what we do not know (the sample was too small to rule out a small gain, which is what the +0.3 end of the interval means), and what we tried (two feature sets, one loss function (the formula the model is trained to minimize)).
- I brought options, not apologies: run the test longer with a pre-agreed stopping rule (a written condition for when to end the test), narrow to the one user segment where the offline signal looked stronger, or stop and redirect the effort. I recommended the segment test because it was cheap and testable, and said I would drop it if it also showed nothing. Before running it I sized it, because a segment is a smaller slice of traffic and a flat reading only counts as "nothing" if the test was large enough to see a lift worth shipping. We agreed the minimum lift that would justify a build (0.3 points, illustrative) and the number of users needed to detect it, and I said I would extend the run rather than call it flat if traffic fell short.
- In the wider review I led with the decision needed, then the evidence. The PM presented the options to the director alongside me, which mattered: the news came as a shared plan.
The segment test also came back flat: its interval excluded a lift of 0.3 points or more, the bar we had agreed, so stopping was a decision backed by evidence and not a shrug (illustrative numbers). We stopped, wrote up what we learned, and the team redirected the effort. Because the criteria had been agreed before the test, nobody felt the goalposts moved, and the PM later brought me the next proposal early.
Trade-offs and pitfalls
- Softening the message until it sounds like a win damages trust far more than the bad news does.
- Do not blame the data or the stakeholders' expectations; own the framing of the plan.
- Avoid drowning people in method detail. Give the one-sentence conclusion first, the evidence second.
You are negotiating a joint research agreement with a university lab that wants to publish freely, while the company needs some control over IP and timing. What are the points you would fight for, where would you give ground, and why?
Sample Answer
Direct answer
I would treat this as a trade, not a fight over principle. The university needs to publish and keep its students free to graduate. The company needs to protect confidential material, file patents before disclosure, and be able to use what it paid for. I would fight for (1) a bounded publication review process, (2) clear ownership and licence terms (covering both what each side brings in and what is created), and (3) protection of the company's confidential data and models. I would give ground on the right to publish, authorship by contribution, and student thesis freedom. I would involve the company's legal team and the university's technology transfer office (the office that handles its IP deals) early, because I am a researcher, not the signing authority.
Points to fight for
The four terms below make up the three headline items: publication review, ownership and licensing (background and foreground IP together) and protection of confidential data and models.
- Publication review, not veto. The lab sends drafts to the company a set time before submission. The company may only (a) require removal of the company's confidential information and (b) request a short delay so a patent application can be filed. It cannot block a paper because it dislikes the result.
- Background IP stays with its owner. Background IP means what each side brings in (code, data, pretrained models). The agreement should list what is included and say the other side gets only the licence it needs for the project.
- Foreground IP (new inventions made during the project): a default rule, for example inventions made by one side's staff belong to that side, jointly made ones are jointly owned, and the company gets a licence (written permission to use the invention) or an option (the first right to negotiate such a licence) limited to the field it cares about (its field of use, e.g. fraud detection, not every possible use). Decide this up front, because it is very hard to negotiate after something valuable exists.
- Confidential data and model weights (the trained numbers that make up a model, which can embody trade secrets such as confidential methods or data) the company supplies remain restricted: named people, named systems, no onward sharing, and deletion at the end.
Where I give ground
- The lab keeps the right to publish its results, with credit by contribution.
- Students are never held back from finishing a thesis: any delay has a hard cap (in the example below, 90 days in total, and a thesis can always be submitted and examined on time with the confidential parts handled in a restricted appendix).
- Non-core outputs (evaluation code, benchmarks, general methods) can be released openly.
- The university keeps a non-exclusive right (a right to use that does not stop the company or others from using it too) to use results for its own research and teaching.
Worked example (illustrative terms)
The company proposes a 90-day publication hold. The lab says 0. A workable middle: 30 days for confidentiality review, plus an extension of up to 60 more days only if the company states in writing that it is filing a patent application. Maximum delay: 30 + 60 = 90 days; typical delay: 30 days, and none of it is a veto. Note what the middle ground actually is: the ceiling equals the company's opening 90 days, but the company only reaches it by stating in writing that a patent application is coming, so the typical delay is a third of its ask (30 of 90 days) and the lab's 0 becomes a fixed, predictable 30-day window. In return, the company accepts that if the 30-day window passes in silence, the lab may submit. The company gains time for patents; the lab gets predictable deadlines for conference submissions.
Trade-offs and pitfalls
- A flat "no publication" demand makes the best labs walk away, and the project loses the reason for partnering.
- A fully open deal can leak unprotected inventions. Public disclosure before a patent filing can ruin patent rights in many jurisdictions, which is why the delay clause matters.
- Vague terms ("reasonable review") become disputes. Put days, who approves and what happens on silence into the text.
- What would change my call: if the work is exploratory and unlikely to produce patentable output, I would loosen the delay; if it uses highly sensitive company data, I would tighten data terms rather than publication terms.
Sample wording
Clause (illustrative): "Each party may submit a draft publication to the other at least 30 days before submission. Within those 30 days the reviewing party may require removal of its Confidential Information only. If it states in writing that it intends to file a patent application, submission may be delayed by up to 60 further days. If no written response is given within 30 days, the publishing party may proceed."
Opening line in the negotiation: "We want your project to succeed and your confidential material protected. Can we agree a fixed review window so neither side is surprised?"
Unlock Full Question Bank
Get access to all 17 Research Collaboration and Stakeholder Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.