The Roadmap Evidence Audit: Trace Every Claim on Your Roadmap Back to a Source
Take your next ten roadmap items and ask five questions of each. Most teams find their highest-cost commitment is backed by nothing but confident retelling. The audit takes an hour and needs no new tooling.
Take the next ten items on your roadmap. For each one, write the single sentence that explains why it is there, then ask five questions of that sentence. Most teams running this for the first time find that three or four items score zero, and that the highest-cost item on the list is among them. The audit takes about an hour, it needs no new tooling, and it is the cheapest way to find out which of your commitments are evidence-backed and which are confidently repeated opinion.
The five questions
For each roadmap item, take its justifying sentence and score one point for each question you can answer without opening a new tab.
1. Who said it? Name a person or an account, not a category. "Enterprise customers" is not an answer. "Three accounts: Northwind, Corvus, and one mid-market trial that churned in February" is an answer.
2. Did we hear it directly, or was it relayed? A transcript, recording or ticket written by the customer is direct. A colleague's summary of conversations they had is relayed, and it proves what the colleague believes rather than what customers need. This distinction is the subject of its own article on hearsay in product research, and it is the question that fails most often.
3. How many, and who did not say it? The second half matters more than the first. Five people asking for something out of forty interviewed is a different fact from five out of five. If you cannot state the denominator, you do not have a rate, you have a memory of the vivid cases.
4. When? Write the date of the evidence, not the date it was last discussed. Findings decay at very different speeds depending on type, and a fourteen-month-old finding about a workflow that has since been redesigned is not weak evidence, it is evidence about a product that no longer exists. See research refresh cadence for how fast each kind of finding goes stale.
5. What would have proved us wrong? If nobody can state a result that would have killed the item, no study was ever going to resolve it, and any research attached to it was decorative.
Score each item from 0 to 5. The bands are simple and you should not over-engineer them.
| Score | Reading | What to do |
|---|---|---|
| 4-5 | Evidence-backed | Build it. Record the evidence link in the item so the next audit is faster |
| 2-3 | Plausible, unverified | Cheap to fix. Two days of interviews usually moves it to 4-5 |
| 0-1 | Opinion in a roadmap costume | Either promote it before committing engineering time, or say plainly that it is a bet |
There is nothing wrong with a bet. Several of the best decisions any product team makes are bets, taken deliberately, by someone who owns the outcome. What is expensive is a bet that everyone in the room believes is a finding.
A worked audit
Here is a realistic result from a mid-market B2B team, lightly disguised. The point of the table is not the scores. It is that the pattern is always the same.
| Roadmap item | Justifying sentence | Q1 who | Q2 direct | Q3 denom | Q4 when | Q5 falsifier | Score |
|---|---|---|---|---|---|---|---|
| Bulk CSV export | "Customers keep asking for export" | No | No | No | No | No | 0 |
| SSO for Growth tier | "Blocked 4 named deals in Q4" | Yes | Yes | Yes | Yes | No | 4 |
| Onboarding rebuild | "Activation is our weakest funnel step" | No | Yes | Yes | Yes | Yes | 4 |
| In-app notifications | "Competitor shipped it" | No | No | No | Yes | No | 1 |
| Reporting redesign | "Every QBR turns into a reporting complaint" | No | No | No | No | No | 0 |
Two things are worth noticing.
The two zero-scoring items are the two largest in engineering cost. That is not a coincidence and it recurs across teams. Large items attract confident retelling precisely because they have been discussed for a long time; the discussion is mistaken for evidence, and the length of time the idea has been around is mistaken for validation. Small items get built before anyone bothers to justify them, so they never accumulate the same false authority.
The onboarding item scores 4 on analytics alone, with no interviews. Behavioural data answers who, how many, when and what would falsify, and it cannot answer why. That is a perfectly good state to be in, and it tells you precisely what the study should be: not "should we rebuild onboarding" but "what happens at step three". A high-scoring item can still need research; the audit tells you what kind.
The four patterns the audit reliably surfaces
The orphaned quote. One vivid customer sentence, usually from a large or beloved account, that has been repeated so often it has become the team's shorthand for a whole class of user. It fails question 3 every time.
The competitor tell. "They shipped it" answers when and nothing else. It is a legitimate reason to investigate and never a reason to build. The relevant question is not whether the competitor shipped it but whether their customers use it, which is a research question you can actually run.
The finding that outlived its product. A solid study from eighteen months ago whose conclusion is still quoted, about a flow that has been redesigned twice since. Nobody re-checks, because the original study was good and the memory of it being good persists after the thing it studied is gone.
The unfalsifiable strategic bet. "We need to be an AI-first product." Nothing could disconfirm it, so no evidence is relevant to it, so the audit's honest output is to move it out of the roadmap and into strategy, where it belongs and where it can be argued about on its merits.
The remedy: a pre-committed assumption record
The audit finds the problem. This prevents it, and it takes about ninety seconds per request.
When someone asks for work to be built or a study to be run, capture their belief in their own words, before any evidence is collected:
- Claim, verbatim
- Origin: how they came to believe it, and roughly when
- Confidence, as a number they choose
- Falsifier: the specific result that would change their mind
This is adapted from a rule in auditing. ISA 580 requires the auditor to obtain written representations from management, and it is precise about their status: written representations are necessary evidence, but the standard states they do not provide sufficient appropriate evidence on their own about any of the matters with which they deal, and that obtaining reliable representations does not reduce the other evidence the auditor must gather. Necessary, recorded, insufficient alone. That is exactly the right status for a stakeholder's belief about customers.
The falsifier field carries the weight. It is written before the results exist, which is the only moment at which anyone can answer it honestly. It converts a preference into a testable claim, and it makes the post-study conversation short: either the falsifying result appeared or it did not.
The record has a second, slower payoff. Over a year, it tells you whose relayed claims survive contact with customers and whose do not. ISA 580 has a version of this too: when representations turn out to be inconsistent with other evidence, the auditor is required to reconsider how much reliance the source deserves in general. Almost no product organisation tracks this, and it is one of the most valuable records a research function can quietly keep. Store it with the request; a standing research intake process is the natural home.
Closing a failing item in two days
For a 0-1 item worth promoting, the whole point is that the fix is now cheap. The old arithmetic, where recruiting and moderating twenty interviews took three weeks and a trained moderator, is what made "we heard it from customers" an acceptable substitute for evidence. That constraint is gone.
The path with Koji is short enough to fit in a sprint:
- Write the brief around the claim and its falsifier, not around the topic. Koji generates a research brief from the objective, and you can edit it directly before launch.
- Use structured questions to recover the denominator, which is what question 3 was asking for. All six types are available alongside conversational probing:
open_endedfor the wording,scalefor intensity,single_choiceandmultiple_choicefor frequencies,rankingfor relative priority, andyes_nofor a clean count. The structured questions guide covers how each is analysed and charted. - Let the AI follow up. This is the part a survey tool cannot do and the part that makes the finding worth having. When someone says the export is painful, Koji asks what they do instead, how often, and what happened the last time, in the same session and with no moderator scheduled. A SurveyMonkey or Typeform response ends at the first sentence, which is exactly the sentence your stakeholder already relayed to you.
- Run voice and text together so you reach both the people who will talk and the people who will only type at 11pm.
- Read the report against the falsifier, not against the claim. The question is whether the disconfirming result appeared, and the pre-committed record means nobody can move the goalposts after the fact.
Twenty AI-moderated interviews launched Monday morning are typically complete by Wednesday, with themes, quotes and a quality score of 1 to 5 per session already attached. An item can go from 0 to 4 inside a single sprint, which is the entire reason this audit is now worth running rather than merely worth agreeing with.
Cadence, and who should run it
Once per planning cycle, on the committed list only. Never on the backlog, which is where this exercise goes to die.
The person running it should not be the person who owns the roadmap. Not because roadmap owners are untrustworthy, but because they know the surrounding context too well to notice that the justifying sentence does not stand on its own. A researcher, a designer, or a PM from an adjacent squad reading the sentences cold will catch things the author cannot. Twenty minutes of preparation, forty minutes in the room, and the output is a short list of items to promote before the quarter starts. That is the whole ritual.
Frequently asked questions
Is this not just going to slow us down?
An hour a quarter, and it usually accelerates things, because the argument it replaces is the recurring one about whether a feature is really needed. Once an item's evidence is written down and scored, that debate has a factual centre and stops being re-litigated in every planning meeting. The teams that find it slow are usually the ones running it on a hundred-item backlog rather than on the ten things they have actually committed to.
What if leadership does not want their items audited?
Audit the sentence, not the person, and never present the scores as a verdict on judgement. It also helps to state up front that a low score is not an instruction to cancel anything. It is a flag that the item is a bet, and a leader taking a deliberate bet with their name on it is a perfectly healthy outcome of the exercise.
Do analytics count as evidence?
Yes, and they are strong on four of the five questions. Behavioural data gives you the population, the denominator, the date, and usually a clear falsifier. What it cannot give you is why, and a rebuild justified by a funnel drop with no qualitative work behind it tends to fix the wrong step. Combining both sources is stronger than either; see triangulation in research.
How does this differ from prioritisation frameworks like RICE?
Prioritisation frameworks assume the inputs are true and rank them. This audit checks whether the inputs are true at all. A reach estimate built on a relayed claim produces a confident score from a fictional number, which is worse than no score. Run the audit first, then prioritise; see research-driven roadmap prioritisation.
What score should we require before building?
Do not set a threshold. A hard gate turns the audit into a compliance exercise and people start writing sentences designed to score well. The useful output is visibility: everyone in the room knowing which items are evidenced and which are bets, and choosing knowingly.
We have no research function. Can we still run this?
Yes, and arguably the return is higher, because nobody has been checking. The audit needs no researcher, no repository and no tooling, only someone willing to ask five plain questions in a meeting. If it turns up items worth promoting, an AI-moderated study is now within reach of a team with no researcher on staff, which was not true two years ago.
Related Resources
- Structured Questions in AI Interviews - the six question types and how each is analysed
- Hearsay in Product Research - grading relayed versus first-hand claims
- Research Refresh Cadence - how quickly each kind of finding decays
- Research-Driven Roadmap Prioritization - what to do once the evidence is checked
- Research Request and Intake Process - where the assumption record lives
- Activating Research Insights - turning findings into decisions
Related Articles
Activating Research Insights: Turn Findings Into Product Decisions
A practical guide to insight activation — the discipline of ensuring research findings actually drive product decisions. Covers why 40-60% of insights are never used, the 4-stage activation framework, decision-ready report formats, and how AI-native research platforms close the loop in real time.
Hearsay in Product Research: Why a Relayed Customer Claim Is Not Customer Evidence
Most product decisions rest on relayed claims about what customers want. Borrow the law of evidence's hearsay rule to grade every claim, and promote the ones that matter to first-hand evidence in 48 hours.
How to Build a Research Request and Intake Process
A step-by-step guide to designing a research intake process: the request form fields that matter, how to triage and prioritize incoming requests, SLAs, and how AI-native research lets you say yes to more requests without adding headcount.
How Long Is User Research Valid? Insight Decay and When to Re-Run a Study
Research does not expire on a fixed schedule — different finding types decay at wildly different rates. A half-life table by insight class, the five decay triggers, and a refresh protocol that keeps your repository honest.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Insight Repository Methodology: How to Build, Tag, and Activate a Research Insight Library (Beyond Just Storage)
The methodology layer most repository guides skip — taxonomy design, atomic insight structure, governance, freshness/decay rules, and the insight-to-action workflow that turns a static archive into a decision engine. Includes a 2-week setup plan and how AI auto-tagging from Koji eliminates the librarian bottleneck.