Research Provenance: How to Prove an Interview, a Quote, or a Report Is Genuine (2026)
When any text can be generated, a customer quote proves nothing on its own. Content Credentials, the EU AI Act marking rules, and the hash-anchored capture method that actually works for research text.
Short answer: the research artefact that travels furthest - a customer quote on a slide - is also the one with the weakest claim to being real, and none of the industry provenance standards fix that. Content Credentials and watermarking were built for images, audio and video. Your most-cited artefact is forty words of plain text, and NIST has documented that text provenance techniques do not survive the ordinary things people do to text. What does work for research is older and less glamorous: anchor the record at capture, keep the linkage from every quote back to a timestamped span in a retained transcript, and be able to demonstrate that the record has not changed. That is the standard courts already use for electronic evidence, and it is achievable today. This guide covers the three separate questions provenance has to answer, what each available mechanism actually does, what the EU AI Act now requires as of 2 August 2026, and the policy to adopt.
Why this became urgent
For most of the history of user research, the authenticity of a quote was never questioned, for a simple reason: producing a convincing fake customer interview was more work than running a real one. That economics has inverted.
The environment is measurably worse. Entrust 2026 Identity Fraud Report, published November 2025 and drawn from over one billion identity verifications across 195 countries and more than 30 industries, found that deepfakes now account for one in five biometric fraud attempts, that instances of deepfaked selfies rose 58% during 2025, and that injection attacks - feeding manipulated media directly into a verification pipeline rather than presenting it to a camera - were up 40% year over year. Digital forgeries made up 35% of document fraud in 2025, against a 29% average across 2022-2024.
Those numbers describe identity verification, not research. They matter here because they establish the ambient reality: synthetic media at scale is cheap, routine and improving, and the tooling built for one adversarial context leaks into every other.
But the sharper problem for research is not adversarial at all. It is that your organisation now generates synthetic text constantly and legitimately, for good reasons, in the same documents where real customer evidence lives. A slide deck contains AI-drafted framing, AI-summarised findings and verbatim customer quotes, in the same font. Six months later, nobody can tell which is which - not because anyone lied, but because nothing in the artefact records the difference.
Provenance is not primarily a fraud-defence problem for research teams. It is a bookkeeping problem that becomes a credibility problem the first time someone asks a question you cannot answer.
Three different questions
"Is this real?" is not one question. It is three, they have different answers, and conflating them is why most provenance discussions go nowhere.
| Question | What it is really asking | Where it is answered | Covered by |
|---|---|---|---|
| Is the participant a real, eligible human? | Sample integrity | At recruitment and during the session | Fraud screening, verification, behavioural signals |
| Is this record unaltered since capture? | Record integrity | At capture and in storage | Cryptographic anchoring, access controls, retention |
| Does this quote actually appear in that record? | Linkage integrity | At extraction and every hop after | Retained pointers from quote to timestamped span |
The first question is participant fraud, and it is well covered elsewhere - see Survey Fraud and Respondent Quality for detection methods. The second is what audit trails and access controls address. The third is the one nobody owns, and it is the one that fails in practice.
Consider what happens to a single sentence a customer says. It is spoken in an interview. It is transcribed. A researcher selects it as a supporting quote. It goes into a findings deck. Someone copies it from the deck into a strategy memo. Somebody else puts it in a board slide. Marketing asks whether it can be used on the website.
Four hops from capture, and at every hop the only thing that travelled was the text. No timestamp, no interview ID, no speaker attribution beyond a first name, no way back. By the board slide, the quote is an assertion by whoever is presenting. That is not a hypothetical failure mode - it is the default behaviour of every research workflow that does not deliberately prevent it.
What Content Credentials actually do
The main industry answer to provenance is Content Credentials, the specification developed by the Coalition for Content Provenance and Authenticity (C2PA).
The model is well designed. A C2PA Manifest contains assertions - statements about the asset such as when and where it was created, what edits were made with which tools, and whether AI was involved. Assertions are bundled with additional information into a claim, which is digitally signed by the claim generator on behalf of the signer, producing a claim signature. Assertions, claim and signature together form a verifiable unit bound to the asset.
The binding comes in two forms, and the distinction is the useful part:
- A hard binding is cryptographic: a hash over the content itself. Change one pixel and the binding fails. It is strong, and it is brittle by design.
- A soft binding is a perceptual identifier - an invisible watermark or a content fingerprint - which allows the manifest to be found again in a manifest repository even if the metadata was stripped from the file. A Content Credential with one or more soft bindings is called a Durable Content Credential.
Soft bindings exist precisely because metadata gets stripped constantly and unintentionally: by screenshots, by re-encoding, by messaging apps, by every social platform. The C2PA design assumption is that provenance metadata will be separated from the asset, and that recovery has to be possible.
For an interview recording, this is genuinely useful and worth adopting when your platform supports it. For a quote, it does almost nothing.
The text problem, stated plainly
NIST addressed this directly in Reducing Risks Posed by Synthetic Content: An Overview of Technical Approaches to Digital Content Transparency (NIST AI 100-4). The report surveys provenance tracking, watermarking, metadata recording and synthetic content detection, and on text it is blunt: all provenance data tracking techniques discussed in the report, when applied to text, have limitations and can be vulnerable to tampering.
The reason is behavioural rather than cryptographic. People interact with text differently from other media. Screenshot a watermarked image and the watermark travels with the pixels. Copy a sentence out of a document, tighten it for the slide, drop a filler word - and every technique that depended on the exact token sequence is gone. Paraphrase defeats text watermarking completely, and light editing for presentation is not an attack, it is normal practice.
This has a specific consequence that research teams should internalise:
The provenance mechanisms getting the most attention are the least applicable to the artefact your organisation actually circulates. A watermarked interview recording sitting in a repository nobody opens is provenance where it is not needed. The quote in the board deck is provenance where it is needed and where no marking scheme can reach.
The answer for text is not a better watermark. It is to stop expecting the artefact to carry its own proof, and to keep the link instead.
What the law now requires
Article 50 of the EU AI Act became applicable on 2 August 2026, which changes this from good practice to an obligation for a large set of research operations.
Two paragraphs matter most.
Article 50(1) requires providers to ensure that AI systems intended to interact directly with natural persons are designed so that those persons are informed that they are interacting with an AI system, unless this is obvious from the circumstances. An AI-moderated interview is the paradigm case.
Article 50(2) requires providers of AI systems generating synthetic audio, image, video or text content to ensure that outputs are marked in a machine-readable format and detectable as artificially generated or manipulated, with technical solutions that are effective, interoperable, robust and reliable as far as this is technically feasible.
Article 50(5) sets the timing: the information must be provided in a clear and distinguishable manner at the latest at the time of the first interaction or exposure, and must meet accessibility requirements. For research, that means disclosure belongs in the invitation and the opening of the session, not in a privacy policy nobody reads.
The Commission adopted guidelines on these obligations on 20 July 2026, and the AI Office has published a voluntary Code of Practice on Transparency of AI-Generated Content offering a recognised route to demonstrating compliance with the marking and detection duties, including a set of icons for labelling AI-generated content. A limited transitional period applies to the marking obligation for generative systems already on the market, running to 2 December 2026. Non-compliance can attract fines of up to EUR 15 million or 3% of worldwide annual turnover, whichever is higher.
Note the asymmetry that follows for a research team, because it is the practically important part. Article 50(2) puts the marking duty on the provider of the generative system, but the credibility of your findings is your problem regardless of who is obliged. Marking your AI-generated report prose is compliance. Being able to prove your customer quotes are not AI-generated is a separate exercise that no regulation requires and every serious research function needs.
| Mechanism | Answers which question | Works for | Does not work for |
|---|---|---|---|
| C2PA hard binding | Record integrity | Recordings, images, exported files | Extracted text spans |
| C2PA soft binding / durable credentials | Record recovery after stripping | Media assets | Plain text |
| Text watermarking | Was this generated | Long unedited generated passages | Paraphrased or lightly edited text |
| AI Act Art 50 marking | Was this generated | Provider-side compliance | Proving something is human |
| Hash anchoring at capture | Record integrity | Any artefact including transcripts | Linkage after extraction |
| Retained quote-to-span linkage | Linkage integrity | Quotes, verbatims, evidence chains | Nothing else - but this is the gap |
The method that works: anchor and link
The legal system solved record authenticity before any of this existed, and its answer is the right template.
Federal Rule of Evidence 902(13) makes self-authenticating a record generated by an electronic process or system that produces an accurate result, shown by certification of a qualified person. Rule 902(14) does the same for data copied from an electronic device, storage medium or file, if authenticated by a process of digital identification and certified. The Advisory Committee notes explain the standard mechanism plainly: hash values are compared, and matching hashes between the original and the copy establish that they are duplicates, making forgery highly improbable.
That is the whole model, and it maps onto research cleanly:
1. Anchor at capture. When an interview completes, compute a hash over the transcript and its metadata and record it with a timestamp somewhere append-only. This costs nothing, requires no new vendor, and converts "trust me, this is the transcript" into a checkable claim. It answers record integrity permanently.
2. Keep the linkage, not just the text. Every quote that leaves the platform should carry its origin: interview ID, question ID, timestamp offset. Not as decoration - as a pointer that resolves. A quote with a resolvable pointer is evidence. A quote without one is a claim.
3. Make the pointer survive the hops. This is where most implementations fail. If the linkage lives only in the research platform and the deck contains bare text, you have solved nothing for the artefact that actually travels. The reference has to be in the deck.
4. Retain long enough to be asked. Provenance is retroactive by nature - you find out you needed it when someone challenges a finding, which is typically long after the study closed. Provenance is a retention requirement disguised as an authenticity feature. Set retention against when claims based on the research will still be in market, not against when the study ends. A quote used in advertising can be challenged years later; see Using Research Quotes in Marketing for what substantiation that actually demands.
5. Label generated content in your own artefacts. In a findings deck, distinguish verbatim participant speech from AI-generated summary from researcher interpretation. Three visual treatments, applied consistently. This is the cheapest intervention in this guide and the one with the largest effect on how your work is received, because it makes the distinction visible at the moment of reading rather than recoverable on request.
The five-minute audit
Take the most recent findings deck your team shipped and pick the three most consequential customer quotes in it.
For each one, try to answer: which interview, which participant, which question, what timestamp? Then open the transcript and confirm the quote appears as written.
Most teams cannot complete this exercise on at least one of the three. Usually it is not because the quote is wrong - it is because the linkage was never recorded, or the quote was tightened somewhere between the transcript and the slide and nobody tracked the edit. That is the gap. It is not a scandal, and it is entirely fixable, but you cannot fix what you have not measured.
Run it this week. The result tells you whether provenance is a real risk in your operation or a theoretical one.
How Koji helps
Koji is built so that the linkage exists by default rather than by researcher discipline.
Every question has a stable ID that travels. Study questions are first-class objects with IDs that follow them from interview plan, through the AI interviewer, into analysis, and out to report aggregation. A finding is therefore attributable to a question, and an answer to a specific interview - the pointer is a structural property of the data model rather than something a researcher has to remember to write down.
Extracted items are grounded in the transcript. Koji analysis extracts insights as discrete grounded items tied to the source conversation rather than producing free-floating prose, which is what makes the pointer from a claim back to its evidence resolvable rather than notional.
Full transcripts export as CSV and JSON. You hold the record, which is a precondition for anchoring it. A platform that returns only summaries has removed your ability to prove anything about the underlying evidence - a question worth asking any vendor before your quotes reach a marketing page.
Disclosure is built into the interview experience. Participants know they are speaking with an AI moderator from the invitation onward, which is what Article 50(1) and 50(5) require and what makes the resulting consent and the resulting quotes usable downstream. Retrofitting disclosure is not possible; the interview either had it or it did not.
Structured questions produce answers that need no provenance argument at all. This is an underrated property. A verbatim quote is an interpretive artefact and needs a chain. A distribution over a defined answer space is not interpretable in the same way:
| Type | What it captures | Provenance burden |
|---|---|---|
open_ended | Free-form qualitative answer with AI follow-up probing | Highest - quotes travel, need resolvable linkage |
scale | Numeric rating such as 1-10 satisfaction or NPS | Low - the number is the record |
single_choice | One option from a list | Low - answer space is defined in advance |
multiple_choice | One or more options from a list | Low |
ranking | Items ordered by preference | Low |
yes_no | Binary answer | Lowest - nothing to extract or paraphrase |
The practical design implication: carry your load-bearing claims on structured questions and use verbatims for illustration. A finding that rests on a quote needs a chain of custody. A finding that rests on 84 people choosing option B needs a sample description. Both are legitimate; only one of them can be challenged on authenticity grounds.
Traditional tooling does not help here. A survey platform has no quotes to authenticate and no conversation to prove. A recorded video call produces an asset that C2PA can bind but that nobody will ever re-open to check a slide. The AI-native position is the one that matters: when the volume of research goes up by an order of magnitude because moderation is automated, the number of quotes in circulation goes up with it, and the linkage has to be automatic or it will not exist.
A provenance policy worth adopting
- Hash and timestamp transcripts at capture. Append-only, immutable, boring.
- Never let a quote leave the platform without its interview ID, question ID and timestamp.
- Put the reference in the deck, not only in the research repository.
- Use three visual treatments in findings artefacts: verbatim, AI-generated, researcher interpretation.
- Set retention against how long the claims stay in market, not against project close.
- Confirm your platform gives you raw transcript export before you build a programme on it.
- Run the five-minute audit quarterly on a recent deck.
- Disclose AI moderation at first contact, per Article 50(5), because it cannot be added later.
None of this requires new vendors or a standards body. It requires deciding that a quote without a resolvable pointer is not finished work.
Frequently asked questions
Does C2PA or Content Credentials solve provenance for research quotes?
No, and it is important to be clear about why. C2PA binds a manifest of signed assertions to an asset using a cryptographic hard binding, optionally backed by a soft binding such as a watermark or fingerprint so the manifest can be recovered if metadata is stripped. That design works well for images, audio and video. A quote is a short span of plain text extracted from a larger record, and extraction breaks any binding. Adopt Content Credentials for recordings and exported files; solve quotes with retained linkage instead.
Can I watermark AI-generated text so people can tell it apart from real quotes?
Not reliably. NIST AI 100-4 states that all the provenance tracking techniques it surveys have limitations when applied to text and can be vulnerable to tampering. The practical failure is mundane rather than adversarial: paraphrasing strips a text watermark entirely, and light editing for presentation is standard practice. Label generated content visibly in your own artefacts rather than relying on a mark that survives nothing.
What does the EU AI Act actually require of an AI-moderated research study?
Two things became applicable on 2 August 2026. Article 50(1) requires that people are informed they are interacting with an AI system unless it is obvious. Article 50(2) requires providers of systems generating synthetic content to mark outputs in a machine-readable, detectable format using solutions that are effective, interoperable, robust and reliable as far as technically feasible. Article 50(5) requires the information at the latest at the time of first interaction or exposure. A limited transitional period for the marking obligation on systems already on the market runs to 2 December 2026, and penalties reach EUR 15 million or 3% of worldwide annual turnover.
How is this different from an audit trail?
An audit trail records who accessed or changed data inside your system, which answers governance questions about your own team. Provenance answers a question asked from outside the system, often years later, by someone who does not have access to it: can you demonstrate this quote came from a real interview and has not been altered? The two are complementary, and the failure modes are different - audit trails are usually present and provenance linkage usually is not. See Research Data Access Controls and Audit Trails for the governance side.
Is hashing a transcript really enough to prove anything?
It proves a specific and useful thing: that the record you hold now is the record that existed at the time of capture. That is exactly the standard Federal Rule of Evidence 902(14) contemplates for data copied from an electronic device, authenticated by a process of digital identification - the Advisory Committee notes describe hash comparison as making forgery highly improbable. It does not prove the participant was who they said they were, and it does not prove a quote on a slide came from that transcript. It answers one of the three questions cleanly, which is more than most research operations can currently do.
Do we need this if we only use research internally?
The internal case is where the credibility cost lands hardest, because internal findings are challenged more often than external ones and there is no legal process to fall back on. When a quote contradicts an executive intuition, the question that follows is some version of "where did that come from?" A resolvable pointer ends that conversation in ten seconds. Its absence turns a finding into a matter of opinion about the researcher.
Should participants be told their interviews are retained for provenance purposes?
Yes, and it is straightforward to do. Retention purpose and duration belong in the consent language at recruitment, alongside how the data will be used and who will see it. Keeping a record longer specifically so that claims made from it can be substantiated is a legitimate and explainable purpose, but it has to be stated at collection rather than decided afterwards. See How to Record Customer Interviews for the consent mechanics.
Related Resources
- Structured Questions in AI Interviews - the six question types and why structured answers carry a lighter provenance burden
- Survey Fraud and Respondent Quality - the participant-authenticity question this guide deliberately does not cover
- Research Data Access Controls and Audit Trails - the governance layer that sits alongside provenance
- Using Research Quotes in Marketing - FTC endorsement rules and what substantiation a public quote requires
- The EU AI Act and User Research - the broader compliance picture around AI-moderated interviews
- AI Model Cards and User Disclosure - documenting what you tell people about the system
- Customer Quotes: How to Extract, Tag, and Use the Voice of Your Customer - the extraction workflow where linkage is won or lost
- Model Version Drift - the companion problem of an instrument that changes between studies
Run your first study free. Koji gives you 10 free credits when you sign up - enough to field a real study, export the full transcripts, and see what a quote looks like when it arrives with its interview ID, question ID and timestamp attached.
Related Articles
AI Model Cards and User Disclosure: Documenting Intended Use, Limitations, and What You Tell People (2026)
A practical guide to model cards, system cards, and user-facing AI disclosure — what belongs in each section, what the EU AI Act's Article 50 has required since 2 August 2026, and how to source the Limitations section from real user research instead of guesswork.
Model Version Drift: What Happens to Your Research When the AI Changes Mid-Study (2026)
When the model behind your AI moderator or analyst is upgraded, your measuring instrument changed. The evidence, the three layers of drift, the bridge sample method, and how to make model version part of your method section.
Customer Quotes: How to Extract, Tag, and Use the Voice of Your Customer
Customer quotes are the most persuasive evidence in product, marketing, and research. This guide covers how to extract them, what makes a quote useful, and how Koji surfaces them automatically.
The EU AI Act and User Research: What AI-Moderated Interviews Actually Require (2026)
AI-moderated customer interviews sit in the EU AI Act's limited-risk transparency tier, not the high-risk tier. Here is exactly what Article 50 requires from 2 August 2026, the two things that escalate a study to high-risk, and a compliance checklist you can run this week.
Research Data Access Controls and Audit Trails: Who Can See Your Interview Data
Your vendor's SOC 2 report proves the vendor is secure. It says nothing about which colleague opened a raw transcript last Tuesday. Here is how to build access tiers, audit trails and access reviews for research data — and what an auditor will actually ask you for.
Using Research Quotes in Marketing: FTC Endorsement Guides, Material Connections, and the Reviews Rule
The moment a customer quote leaves your research repository and appears in an advertisement, it stops being data and becomes an endorsement. Three obligations attach immediately: the words must be faithful, the experience must be typical or disclosed, and any material connection must be visible.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Fraud & Respondent Quality: How to Detect Fake and Low-Effort Responses (2026)
Between 5% and 26% of survey responses are fraudulent, and AI-generated answers now pass standard quality checks. Learn the warning signs, the detection tactics that still work, and how Koji's conversational quality gate filters bad data before it reaches your report.