Ranking product requests by how many times they were mentioned is still the most common way I see roadmaps get built. Twenty-three mentions beats eleven, so twenty-three goes higher. The tools can do more than count, and some teams use the weighting features, but the mention count is the default view and it’s what most prioritization conversations end up leaning on. I think that habit is the cleanest example of survivorship bias in product work, because a mention count is a poor measure of demand. It mostly measures who was still around to be counted, and how loud your collection system made them.
The accounts that generate mentions are the ones with a CSM, an AE, an executive sponsor, a shared Slack channel, and a quarterly review. They can produce forty data points on one feature in a month. The account that churned in the spring produces zero, and so does the prospect who walked away, the user who never got deep enough into the product to hit the problem, the experiment that returned nothing, and the idea that died before anyone instrumented it. Some of those were never recorded anywhere. A count can’t see any of them, and a roadmap built on the count inherits every one of those blind spots while looking perfectly data-driven.
What survivorship bias looks like in product management
Survivorship bias in product management is drawing conclusions from the customers, features, experiments, or decisions still visible after some filter has run, while ignoring whatever the filter removed.
The standard example is Abraham Wald’s work with the Statistical Research Group at Columbia during the Second World War. He was estimating aircraft vulnerability from the damage on planes that came back from missions, and the problem was that the planes that didn’t come back weren’t in the data. The bomber story gets retold in a very tidy form these days, but the underlying issue is exactly as Wald described it. The dataset you have is not necessarily the population you care about.
Product teams run into this constantly, because the customers who stayed, the features people found, the users willing to talk, the experiments that ran to completion, and the projects somebody bothered to write up are the only things available to look at. Then you build a picture of reality out of what’s left.
A note on terms, since a statistician will otherwise stop reading here. Survivorship bias has neighbors. Vocal accounts producing more feedback is a response bias. Counting six tickets from one account as six independent signals is a weighting problem. Judging a decision by how it turned out is outcome bias, which gets its own section below. They travel together in product work and the fixes overlap, so I’m treating survivorship as the main thread and naming the others where they show up.
Why feature adoption doesn’t prove feature value
Feature adoption data is filtered by survival before you ever look at it, which is why an adoption-retention correlation is weak evidence that a feature worked. Say a team looks at retention and finds that accounts using Feature X retain at three times the rate of accounts that don’t. The natural read is that Feature X drives retention.
Maybe. There’s another explanation that fits the same data. Accounts that survive long enough to discover Feature X might already be the engaged ones. Bigger accounts might have had better onboarding. Power users might be more likely both to adopt the feature and to stick around regardless. And there’s a mechanical version of the problem hiding in how adoption is measured: if “adopter” means “ever used it,” an account has to survive long enough to become an adopter, so the adopter group is pre-filtered for longevity before you compare anything. The team has found an association and is treating it as a cause.
This matters for what you do next. If Feature X causes retention, redesigning onboarding to push everyone toward it makes sense. If Feature X is mostly a symptom of being an engaged customer, that redesign will do close to nothing and cost you a quarter. A successful-looking cohort is not, on its own, evidence that the decision behind it was good.
Where survivorship bias enters product decisions
Survivorship bias enters through the routine inputs, not through one obvious mistake, and it accumulates across every source the team touches.
| What you analyze | What survives into the dataset | What tends to disappear |
| Customer feedback | Active, vocal customers | Churned, disengaged, silent customers |
| Sales feedback | Open opportunities and current accounts | Lost deals and disqualified prospects |
| Product analytics | Users who reached the feature | Users who left before reaching it |
| User research | People willing to participate | People who stopped caring |
| Feature analysis | Shipped functionality | Rejected and abandoned bets |
| Experimentation | Winners and memorable tests | Null, inconclusive, and failed experiments |
| Roadmap reviews | Current priorities | Decisions deliberately not pursued |
| AI context | Data available for retrieval | Evidence that was never connected or never recorded |
The last row is the one I’d watch. AI does not fix survivorship bias on its own. If anything it scales it. A model asked to find patterns across incomplete evidence will get very good at finding patterns in the wrong population, and nothing in the output will tell you the population was wrong.
How much evidence churned accounts leave behind: a worked example
Missing populations stay abstract until you count them, so this is a hypothetical with round numbers.
Say you have 1,000 accounts over the past year: 800 retained, 200 churned. Now count how many of each have any recorded product evidence (a ticket, a call note, a survey response, a tagged conversation).
| Population | Accounts | Accounts with recorded evidence | Coverage |
| Retained | 800 | 400 | 50% |
| Churned | 200 | 20 | 10% |
Churned accounts are 20% of the customer base and under 5% of the evidence base. Whatever the churned accounts had in common, your feedback data is representing it at roughly a quarter of its real weight. The table shows unequal coverage and nothing else. It can’t tell you what the silent 180 wanted, and it can’t measure how far the roadmap was pulled off course, but it does tell you where to dig. In this example, suppose the 20 churned accounts with evidence mention setup complexity at three times the rate retained accounts do. That’s a thin sample, but it’s enough to justify going back through the 180 silent ones (renewal call notes, CSM handoffs, usage drop-off points) before the roadmap locks. Sometimes that reverses a priority. More often it changes the forecast attached to one: the advanced-reporting bet stays, but the expected impact gets smaller and an onboarding fix moves up a slot.
Your best customers can pull the roadmap in the wrong direction
Talking to great customers is useful, right up until they’re the only ones you build from, because then the product ends up serving the people it already selected for.
The shape of it usually goes like this. Your top 20 accounts keep asking for advanced reporting. Several million in ARR touches the request, sales agrees, customer success agrees, and it goes to the top of the roadmap. Now look at the population that isn’t in the room. Over the past year, lost deals kept citing implementation complexity. A chunk of accounts churned before they ever reached the advanced workflows. Smaller accounts never activated the basic reporting you already have, and new users kept dropping out during setup.
So the survivors are asking for more depth while the people who left couldn’t get through the depth that already exists, and looking only at the first group turns “our biggest customers need advanced reporting” into a truth about the market, when it may only describe the customers your current product filters for.
Building advanced reporting might still be the right call. The difference is that now it’s a tradeoff somebody chose with both populations in view.
A request count is a measurement of your collection system
Survivorship bias affects customer feedback through the collection system: not everyone produces feedback at the same rate, and the people who leave produce almost none. Back to the mention count, because it deserves more than the opening paragraph. A frustrated customer who files six tickets shows up as six data points. A customer who decides not to renew and never says why shows up as zero. A mid-sized account on a self-serve plan barely registers no matter how much pain it’s in. So a count of requests is a measurement of observed demand after your collection process has reshaped it, and the reshaping is not random. It systematically favors the accounts you already serve well.
The fix is context, attached to every request before anyone ranks it: which accounts, which segments, how much revenue, what lifecycle stage, retained or churned, won or lost, activated or not, and what other evidence points the other way. This is the problem Bagel is built around. It pulls signals from sales calls, support, CRM, Jira, and product analytics into one place and attaches reach and revenue to each opportunity, so the unit you rank is the evidence behind a decision rather than the number of times someone said it.
One caveat about that fix. Revenue weighting on its own can make the survivorship problem worse, because the accounts with the most ARR are also the ones with the most coverage. Revenue context is for making the commercial tradeoff explicit (“this serves 2% of accounts and 30% of ARR”). Population coverage is a separate check, and you need both. The same goes for lost deals: a prospect who was never in your target market shouldn’t move the roadmap the way a lost deal inside your ICP should, so define the population before you start counting who’s missing from it.
Outcome bias turns a lucky guess into a promotion
Two PMs each make a call with incomplete information. The first does careful work, has strong evidence and reasonable assumptions, and the feature fails because of a market shift nobody could have seen. The second goes mostly on gut, gets lucky, and the feature takes off. In most organizations I’ve worked in or sold to, the second one gets promoted.
Jonathan Baron and John Hershey demonstrated the underlying effect in 1988: people rated otherwise identical decisions more favorably when the outcome happened to be good. A preregistered replication with several hundred participants reproduced it. The studies don’t single out product teams, but product work is a good environment for the bias to thrive in, with long feedback loops, noisy outcomes, launches that get celebrated, and reviews that happen a year after the call was made. A feature grows revenue, so the decision was brilliant. A customer churns after you declined their request, so you should have built it.
The outcome tells you something, just not everything about the decision that produced it. Decision quality is how good the reasoning, evidence, and forecast were at the time the call was made. Outcome quality is what happened afterward. A good decision can produce a bad outcome and a bad decision can get lucky, so assess decision quality before you know the result, assess outcome quality after, and keep the two assessments separate.
Five checks you can run on a decision before the outcome is known
You measure product decision quality by scoring the decision itself, separately from its result: how much of the relevant evidence was represented, whether the missing populations were checked, what the team predicted, and how confident it was. There’s no validated decision quality score for product, and I’d be suspicious of anyone selling one, since a single number hides too much. What follows is the set of review checks we use and recommend. Nobody has validated them as a measurement standard, and each one is worth tracking on its own.
Evidence coverage
Which relevant populations and sources were actually represented? The crude version is sources consulted divided by sources available, which treats a Gong call and a duplicate Zendesk ticket as equal. A more useful version is the coverage table above: for each population that matters to the decision, what share of it left any evidence at all, and what are the known gaps. A roadmap call backed by call transcripts, tickets, usage data, lost-deal notes, churn reasons, and account revenue has different coverage from one backed by six customer interviews, and this check makes the difference visible.
Survivor coverage
Did anyone look at the non-survivors? Split it out explicitly: retained versus churned accounts, active prospects versus lost deals, activated versus failed activations, adopters versus non-adopters. If the evidence set is entirely current active users, that should trip a warning before the decision goes anywhere.
Counterevidence
Teams are good at assembling the case for what they already want to build, so record the other side. I’d avoid a ratio here, since a higher share of contradictory evidence isn’t better and any percentage target gets gamed. Instead, write down the strongest single piece of counterevidence, where you searched for it, and whether it changed anything. A decision record that can’t answer “what would make us change our mind?” hasn’t looked.
Forecast specificity
Before you build, write down what you expect, in a form someone can check later. Something like “we expect the new setup flow to raise 14-day activation for new self-serve accounts from 31% to 38% within eight weeks of full rollout, measured in Amplitude on the existing activation definition, owner: the growth PM,” rather than “improve onboarding.” Record the population, the metric and how it’s measured, the baseline, the expected change, the time horizon, your confidence, the acceptable downside, who owns the readout, and what would invalidate the whole thing. Skip this and the team will remember its original expectations differently once the numbers are in.
Calibration
Attach a probability to one resolvable event per decision. For instance: “70% that 14-day activation is up by at least five percentage points on the agreed measurement by the review date.” One event, one date, one method, so it resolves cleanly to yes or no. A single forecast like that teaches you nothing. Over dozens of comparable forecasts you start to learn something, though how much depends on how similar and how independent those decisions are. The Brier score is a widely used way to grade forecasts like this (squared difference between the predicted probability and what happened, lower is better), and it rewards both calibration and sharpness, so pair it with a plain calibration table: bucket your forecasts by stated confidence and see how often each bucket came true. If the things your team called “80% likely” happened 55% of the time, you have a systematically overconfident product org, and that’s a fixable problem once you can see it.
What a product decision record should contain
A product decision record, sometimes called a decision log, should contain twelve fields and fit on one page. Write it while the future is still uncertain and leave the PRD somewhere else.
| Field | Question it answers |
| Decision | What are we doing? |
| Problem | What are we trying to change? |
| Population | Who specifically is affected, and who is out of scope? |
| Evidence | Which signals support this? |
| Missing evidence | Which relevant populations or sources are absent? |
| Counterevidence | What’s the strongest point against it, and where did we look? |
| Alternatives | What else did we seriously consider? |
| Prediction | What do we expect to happen, measured how, by when? |
| Confidence | How sure are we? |
| Success condition | What result counts as success? |
| Failure condition | What result makes us reverse or change course? |
| Owner and review date | Who reads the result, and when does reality get a vote? |
Decision journals work because they preserve what people believed when the answer was still unknown, instead of relying on memory after the fact.
How to run a product decision postmortem
A product decision postmortem is a post-launch review that starts from the decision record rather than the result, in seven steps. Read the record before anyone asks whether the feature won.
1. Freeze what you believed. Read the original prediction, assumptions, evidence, and confidence level. Don’t edit them.
2. Measure exposure before adoption. Who realistically could have used this? Eligible, exposed, discovered, tried once, came back, never used, churned. Analyzing adopters alone selects the population after the intervention, which is the same bias again.
3. Measure the outcome you committed to up front. Activation, retention, conversion, churn, expansion, support volume, task time, win rate, implementation time, revenue, cost to serve, whichever metric you wrote down before launch, on the measurement method you wrote down with it.
4. Ask the causal question. “Did users of the feature retain better” and “did access to the feature increase retention” are different questions. The cleanest way to answer the second is a randomized rollout: assign eligible accounts to the change or the control at random, decide the duration, sample size, outcome window, and stopping rule before you start, and analyze people in the groups they were assigned to rather than the groups they ended up in. If you compare adopters to non-adopters inside the treatment group, you’ve thrown away the randomization and reintroduced the bias you were trying to remove. A staged rollout that isn’t randomized is a weaker design and should be described as one. When you can’t randomize at all, be loud about the confounders, and label correlation as correlation.
5. Go find the people who disappeared. Non-adopters: why not? Churned accounts: did this problem show up in their tickets, calls, or renewal conversations? Lost deals inside your ICP: did the decision touch the sale? Users who quit during onboarding: could they ever have reached the feature you’re celebrating? These groups make the dashboard inconvenient, and they’re usually where the interesting answer is.
6. Compare the forecast with reality. Expected +8 points of activation, observed +3 with an interval of roughly +1 to +5 from the experiment. Expected 30 days to impact, observed 75. Confidence at decision time 80%. Direction right, magnitude badly overestimated. That’s more useful than “success” because it tells you something about how the org thinks, and the interval is what keeps you from over-learning from one result.
7. Change the next decision. The learning has to modify something concrete: an assumption, the expected impact of similar bets, a segment definition, the prioritization model, how much weight a given evidence source gets, or the context an agent works from.
Post-decision analysis is where Bagel does a lot of its work. Once something ships, it keeps the exposed accounts, the churned and lost ones, and what they said connected to the original opportunity, so step 5 is a query instead of a research project. It puts the missing population next to the surviving one so the comparison is possible at all. It won’t establish causality for you, and evidence that was never collected stays uncollected.
Run a pre-mortem too
Gary Klein’s project premortem is, in my experience, the cheapest version of all this. Before anything ships, assume it has already failed and have the team write down why. It pulls out the risks and dissent that the planning meeting suppressed. For product decisions I’d add a few specific prompts: assume this decision fails, what did we misunderstand? Which customer population are we probably undercounting? What evidence will we wish we had six months from now? Those usually surface more missing information than another round of roadmap scoring.
AI makes decision memory a product requirement
Product orgs are moving from people using AI tools toward people and agents sharing the work. An agent working on your product needs to know why something was picked, what was rejected, what the evidence was, and whether the decision still holds. Without that, every new agent session can resurrect an idea the company killed six months ago with no memory of why, or open a ticket for something another agent shipped last sprint. Agents are very good at re-deriving a reasonable-sounding decision from scratch, and re-deriving is how you get the same rejected feature proposed four times by four different sessions.
Bagel’s MCP handles the first half of this. It exposes customer evidence, revenue context, and scoped decisions to Claude, Cursor, and Codex, so agents build from the same picture the product team is looking at. In our customer base, making shipped and rejected opportunities visible to agents through that shared context comes with roughly 90% fewer duplicate opportunities. We think the decision record is a big part of why, though a duplicate count on its own doesn’t prove the mechanism.
The second half is longitudinal. A product brain should be able to answer, for any decision: what did we believe before we built this, what evidence backed it, what was missing, what probability did we give it, what actually happened, which cohorts moved, were we directionally right, which assumptions broke, and do we keep overestimating a particular kind of opportunity. Then it should let that history change the decision being made today. A backlog holds none of that.
Your roadmap is a dataset
Teams spend a lot on the quality of the data going into decisions and almost nothing on a dataset of the decisions themselves. Every roadmap call is a prediction, every rejected feature is a counterfactual, and every churned customer, lost deal, and null experiment is information, and most of it evaporates because nobody wrote down what they expected before they found out.
The winners are easy to study because they’re still standing. The harder question is what left the dataset before you looked, and the quality of your decisions depends on both.
If you want to see what the missing population looks like in your own feedback data, book a walkthrough and bring a decision you’re about to make.



