Back to docs
Research Operations

Hearsay in Product Research: Why a Relayed Customer Claim Is Not Customer Evidence

Most product decisions rest on relayed claims about what customers want. Borrow the law of evidence's hearsay rule to grade every claim, and promote the ones that matter to first-hand evidence in 48 hours.

Most of what a product team calls "customer evidence" is hearsay: a claim about what customers want, offered by someone who is not the customer. It is not worthless, but it proves something different from what the team thinks it proves. A stakeholder saying "customers keep asking for bulk export" is reliable evidence that the stakeholder believes it, and weak evidence that customers want it. The fix is not to argue. It is to name the source, grade it, and convert the important claims into first-hand evidence, which platforms like Koji now make possible in about two days rather than three weeks.

The rule, borrowed from a profession that had to solve this

The law of evidence has spent four hundred years on exactly this problem, and its answer is compact. Under Federal Rule of Evidence 801(c), hearsay is a statement the declarant did not make while testifying, offered to prove the truth of the matter asserted. Rule 802 states the consequence plainly: hearsay is not admissible unless a rule or statute says otherwise.

Read that definition again with a roadmap in hand. "Customers want bulk export" asserted by a product manager in a planning meeting is a statement by someone who is not the declarant of the underlying claim, offered to prove the truth of what customers want. The customer is the declarant. The customer is not in the room. Nobody can cross-examine the customer. That is hearsay in the strict sense, and courts exclude it for a reason that transfers perfectly to product work: you cannot test the parts of the statement that would tell you whether to believe it.

The three things you lose when a claim is relayed are always the same:

  1. The wording. You get the relayer's paraphrase. "Bulk export" may have been "I need to get a year of this into a board deck once a quarter", which is a reporting problem, not an export problem.
  2. The denominator. "Customers keep asking" collapses an unknown number of people into a plural. It is frequently three, and two of them are the same account.
  3. The disconfirming detail. Every real customer conversation contains something inconvenient. Retellings shed it first, because the relayer is summarising toward a point.

The source ladder

Grade every claim before you argue about it. This ladder takes about fifteen seconds per item and settles most roadmap disputes without anyone having to be wrong out loud.

RungSourceWhat it actually provesWhat it does not prove
1Recorded interview or session with a named participant, transcript intactThis person said this, in these words, on this dateThat the population shares it
2A system record generated at the time: support ticket, in-product event, sales call recordingThe event happened and was loggedWhy it happened, or what the person wanted instead
3A researcher's contemporaneous notes, no recordingRoughly what was said, filtered onceExact wording, tone, or what was asked to prompt it
4A colleague's summary of conversations they hadThat the colleague holds this beliefAnything reliable about customers
5"Everyone knows", "the market is moving to", an unattributed slideThat the claim has survived retellingNothing

Rungs 4 and 5 are hearsay with no exception available. They are still useful, because they are excellent at generating hypotheses. They are simply not evidence, and the failure mode is treating them as though a decision can rest on them.

The three exceptions worth honouring

Rule 803 lists exceptions where hearsay is admitted anyway, because the circumstances of the statement make it trustworthy. Three of them translate cleanly and are worth building into how your team reasons.

Present sense impression (803(1)). A statement describing an event, made while or immediately after the declarant perceived it. The product analogue is anything captured at the moment of the experience: an in-app feedback note typed during the failure, a support ticket opened while the user was still stuck, a session recording. Contemporaneity is doing the work. The person had not yet had time to construct a tidier story. This is why an intercept survey fired at the moment of abandonment beats the same question asked in a quarterly relationship survey.

Records of a regularly conducted activity (803(6)). Your own systems, kept in the ordinary course, not assembled for the argument. Ticket volumes, churn reasons captured by a standing form, event logs. The trustworthiness comes from the record having been created for an operational purpose before anyone knew it would be cited. A dataset pulled together specifically to win a prioritisation debate does not qualify, and everyone in the meeting can smell it.

Statements against interest. Not a perfect analogue, but the underlying logic is the strongest single heuristic in stakeholder research: a claim that costs the speaker something is more credible than one that serves them. A sales lead saying "we lose these deals on price, not features" is testifying against the story that would get them more engineering support. Weight it accordingly. So is a customer volunteering that they did not read the onboarding email you spent a month on.

Written representations: get the claim on the record

Auditing has a parallel rule that closes the loop, and it is the practical move most product teams are missing. ISA 580 governs written representations, which are the formal statements a company's management gives its auditor. The standard is careful about their status. Written representations are necessary information and they are audit evidence; but the standard states that they do not provide sufficient appropriate evidence on their own about any of the matters with which they deal, and that receiving reliable representations does not reduce the other evidence the auditor must obtain.

That is the exact status a stakeholder claim should have on your team. Necessary, recorded, insufficient alone.

So write it down. When someone asserts a customer need, capture four fields in your intake or brief before the discussion moves on:

  • Claim, in the speaker's own words
  • Declarant: who is the customer this originates from, by name or account
  • Basis: how the speaker came to believe it, and when
  • Falsifier: what result would change the speaker's mind

The fourth field does most of the work. A claim whose owner cannot state what would falsify it is not a hypothesis, and no study will resolve it. ISA 580 also has a rule for what happens next: if representations are inconsistent with other evidence, the auditor performs procedures to resolve the matter, and if it stays unresolved, reconsiders how much weight management's statements deserve in general. Product teams almost never do the last part. Tracking which stakeholders' relayed claims survive contact with real customers is one of the highest-value records a research function can keep.

Promoting hearsay to first-hand evidence in 48 hours

The traditional reason teams run on hearsay is arithmetic. Recruiting, scheduling and moderating fifteen interviews took two to three weeks and a trained moderator, so the relayed claim won by default. AI-moderated research removes that constraint, which is why the hearsay problem is now a choice rather than a condition.

With Koji, the promotion path is short:

  1. Write the brief around the claim, not the topic. State the stakeholder's claim and its falsifier. Koji's AI turns that into a study with a research brief you can edit directly.
  2. Mix question types so you get both the wording and the denominator. Koji supports six structured question types alongside conversational probing: open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. The open_ended questions recover the wording that the retelling destroyed; the scale, single_choice, multiple_choice, ranking and yes_no questions give you the counts that "customers keep asking" was standing in for. See the structured questions guide for how each type is analysed.
  3. Let the AI probe. The value of a first-hand transcript is in the follow-up, and this is where AI-moderated interviews beat any survey tool. When a participant says "the export is painful", Koji asks why, what they do instead, and what happened last time, without a moderator on the call. A Typeform or SurveyMonkey response stops at the first sentence, which is precisely the sentence the stakeholder already relayed to you.
  4. Run text and voice in parallel. Voice interviews recover tone and hesitation; text interviews reach people who will not book a call. Both produce a full transcript, so both land on rung 1.
  5. Check quality before you cite. Every Koji interview receives a composite quality score from 1 to 5 with a breakdown across relevance, depth and coverage. Citing a 2 as though it were a 5 recreates the original problem one level down. See understanding quality scores.

Twenty conversational interviews launched on a Monday are routinely complete by Wednesday. That is the whole argument. When first-hand evidence costs two days, there is no defensible reason for a roadmap item to rest on rung 4.

The one-page hearsay audit

Run this on your current sprint or quarter. For each committed item, write the single sentence that justifies it, then mark the rung. Teams doing this for the first time typically find that between a third and a half of committed work sits on rungs 4 and 5. The point is not to cancel that work. It is to notice that the highest-cost items are often the worst-evidenced ones, because big bets attract confident retelling.

Then pick the two items with the largest cost and the weakest rung, and promote only those. Trying to promote everything is how the exercise dies.

Anti-patterns

  • Treating volume of relayed claims as corroboration. Five people repeating one account manager's story is one source, not five. This is the single most common error, and it is why "we hear this constantly" should always be met with "from whom".
  • Using the audit to embarrass colleagues. The ladder grades claims, not people. A stakeholder on rung 4 is doing their job; they are relaying signal they genuinely received.
  • Promoting a claim you were always going to build anyway. If the falsifier field is blank, the study is theatre. See confirmation bias in user research.
  • Assuming a first-hand quote settles the denominator. One transcript is rung 1 and n=1. Reliability of the source and sufficiency of the sample are separate questions; see how many user interviews you need.

Frequently asked questions

Is stakeholder input useless, then?

No, and treating it that way is a fast route to being ignored. Stakeholder input is the best available source of hypotheses, context, commercial constraints and history. What it cannot do is stand in for the customer's own account of their experience. Run stakeholder interviews deliberately, capture what people believe, and treat those beliefs as claims to be tested rather than findings to be implemented.

How is this different from just saying "talk to customers"?

"Talk to customers" is advice. The hearsay frame is a decision procedure. It tells you which specific claims to spend research capacity on, in what order, and what a claim's current evidentiary status is, without requiring anyone to agree in advance about what customers want. It also gives you a defensible answer when an executive asks why a study was needed.

Do support tickets count as first-hand evidence?

They sit on rung 2, and they are genuinely valuable because they are contemporaneous records kept in the ordinary course. The limitation is that a ticket tells you what went wrong and almost never why the person wanted the thing in the first place, and the population that files tickets is not the population that uses your product. Pair them with interviews rather than choosing between them.

What if the customer is not available to interview?

That is precisely the situation the exceptions are for. Fall back to contemporaneous records, in-product behaviour and statements against interest, and be explicit in the write-up that the finding rests on those rather than on direct testimony. Stating the rung is what keeps the conclusion honest. For hard-to-reach populations, asynchronous AI interviews often reach people a scheduled call never will, because participants complete them at 11pm in their own time.

Does an AI-moderated interview count as first-hand testimony?

Yes. The participant speaks or types in their own words, the exchange is recorded and transcribed in full, and the transcript is available for anyone to read rather than being filtered through a moderator's summary. In evidentiary terms, an AI-moderated transcript is closer to a deposition than to a note: the words are the participant's, and the follow-up questions that produced them are on the record too.

How often should we run the audit?

Once per planning cycle, on the committed list only, and it should take under an hour. Running it on the backlog is how it becomes a chore nobody repeats. Pair it with a scheduled review of whether existing findings have gone stale, which is a separate question covered in research refresh cadence.

Related Resources

Related Articles

Assumption Testing: How to Validate Product Assumptions Before You Build

Learn how to identify, prioritize, and test the assumptions behind your product decisions — before building the wrong thing. Includes the assumption mapping framework, testing methods, and how AI interviews accelerate validation.

Confirmation Bias in User Research: How to Recognize and Eliminate It

Confirmation bias quietly corrupts user research by leading teams to hear what they already believe. Learn how it shows up in interviews and analysis, and the practical tactics — and AI moderation — that neutralize it.

Research Provenance: How to Prove an Interview, a Quote, or a Report Is Genuine (2026)

When any text can be generated, a customer quote proves nothing on its own. Content Credentials, the EU AI Act marking rules, and the hash-anchored capture method that actually works for research text.

Stakeholder Interviews: How to Align Your Team Before Research Begins

A complete guide to conducting stakeholder interviews before user research — how to identify the right people, craft powerful questions, synthesize input, and build organizational alignment around your research plan.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Triangulation in Research: Combining Methods for Stronger, More Credible Insights (2026)

Triangulation is the practice of using multiple data sources, methods, researchers, or theories to validate a finding. Learn Denzin's four types, when to use each, and how AI-native research platforms make multi-method studies practical instead of aspirational.