Someone on your team has run an AI persona study, and the output reads well: themes, quotes, a confident ranking of what users want. Your job is to decide how much of it can reach a roadmap. This playbook gives you a way to do that: turn the synthetic output into a list of testable claims, put each one to real participants in AI-moderated interviews, and score what held up. By the end you will have a protocol you can run in a week, and a record of where synthetic output has been right and wrong in your domain.
Why validating synthetic research matters now
Three things changed at once. More people are producing "insights", the tools that produce them are cheaper, and the researcher is rarely in the room when the output gets used.
Photo by Kelly Sikkema on Unsplash
On the first, Maze's 2026 report, based on a survey of nearly 500 researchers, designers and product professionals, says insights now come from product managers (39%), market researchers (35%) and marketers (23%). The same report flags that systems have not kept pace with that spread, which leaves gaps in quality, consistency and trust. Synthetic output slots straight into that gap: it arrives fast, it looks finished, and nobody has checked it.
On the second, the evidence on AI personas is mixed at best. Nielsen Norman Group reran three of its own studies with synthetic users and concluded that the responses were too shallow to be useful for many research activities. In its online training example, real learners often started courses and never finished them, while the synthetic learners reported finishing everything. A paper in Political Analysis (Bisbee et al., 2024) prompted ChatGPT with personas and compared the results with survey data. About half of the coefficients differed significantly from the survey, and roughly a third of those differing coefficients had the opposite sign. A November 2025 preprint found that even models fine-tuned on a small human sample did not reproduce the original study's regression coefficients, and the authors concluded that LLM-generated data is not yet suitable for replacing human participants in formal inferential analysis.
On the third, the best result anyone has published for simulated people started with real ones. Stanford researchers built agents from two-hour interviews with 1,052 people and found the agents replicated General Social Survey answers 85% as accurately as participants replicate their own answers two weeks later. That is a preprint, and it measures survey answers, not product decisions. Still, the lesson holds for your work: a simulation is only as good as the real conversations behind it.
None of this means you should ban synthetic tools. It means a synthetic finding is a hypothesis until someone checks it. If nobody does, you pay for it later: a feature built for needs no real person had, a price set on enthusiasm that a model produced, or a stakeholder who now distrusts all research because one confident report fell apart.
What good looks like
Treat the synthetic run the way you would treat a stakeholder's strong opinion. Useful, worth testing, not evidence yet.
Freeze it first. Save the synthetic output with a date before any real interview happens. Otherwise you will unconsciously read it into what participants say.
Work at the level of claims, not themes. "Users value flexibility" cannot be tested. "Most team leads would pay more for a plan that lets them reassign seats" can. Write each claim so a participant could contradict it.
Rank claims by the decision they drive. Validate the three to seven claims that would change a roadmap, a price or a message. Skip the colour commentary.
Know which claim types fail most. From the research above, watch priorities (models tend to care about everything equally), completion and follow-through (models say they finish), and enthusiasm for concepts (models rarely say no). Put these first.
Ask about the past, not opinions about the future. "Tell me about the last time you did this" tests a behavior claim. "Would you use this?" invites the same politeness a model shows.
Keep the synthetic output away from the interviewer. Do not feed the persona transcripts to the AI interviewer as background and do not show them to participants. You want an independent reading.
Recruit the audience the persona was supposed to be. Use the same criteria you gave the model, written as a real screener.
Score every claim. Confirmed, contradicted, reversed, or not raised. Then add a fifth column, the one most people skip: what real participants raised that the model never did.
Common mistakes:
- Validating only the claims you already believe
- Asking participants to react to the synthetic summary, which anchors them
- Sampling a handful of friendly customers and calling it confirmation
- Averaging "confirmed" and "contradicted" into a vague "mostly right"
- Never writing down the result, so the next team trusts or distrusts the tool by mood

How to set it up in Koji, step by step
You need a Koji account. Text interviews work on every plan; voice interviews need the Interviews or Enterprise plan. A text interview uses 1 credit and a voice interview 3, and only interviews that pass the quality score use credits. Insights is €29 a month and Interviews is €79 a month. Check the pricing page for what is current before you plan a budget.

- Start the study. Open the template gallery, filter by UX Researchers and pick Concept Testing Interviews, or type your research question on the home screen and let Koji draft the brief. The template's structure fits claim checking well.
- Put the claims in the brief. State the decision you are informing, then list the synthetic claims as the hypotheses under test. The research brief assistant can rewrite any part of the brief when you ask, and it names what it changed. Leave the synthetic transcripts out of the context files, since those sharpen the interviewer's follow-ups and you want it unprimed.
- Add structured questions where a number or choice is needed. Koji supports open-ended, scale, single choice, multiple choice, ranking and yes/no questions. Use ranking to test whether priorities are really equal, and a scale plus an open follow-up for intent.
- Add screening questions. Mark which answers qualify and which rule people out, so only the audience your persona represented gets through.
- Preview. Take the interview yourself, by text or voice, as a participant would. Previews do not use credits or appear in the report.
- Choose text or voice and invite people. Share the interview link, import a CSV for personal invite links to a customer list (Interviews and Enterprise plans), or recruit from the panel. Building a panel audience and seeing the cost works on any plan; launching a recruitment needs a paid plan.
- Read the report. Findings show how many participants support each one and link to the quotes behind it.

Example questions written for this use case:
- "Think about the last time you had to [do the task]. Walk me through what happened, from the start." (open-ended, tests behavior claims)
- "Put these five things in order of how much each would matter when you choose a tool for [task]." (ranking, tests the equal-priorities problem)
- "How likely are you to try [concept] in the next three months?" (scale 1 to 10, followed by an open "What made it that number?")
- "What would stop you from using this?" (open-ended, because models rarely raise blockers)
- "Have you ever started [flow] and not finished it?" (yes/no, followed by "What happened?")
- "What is the first thing you would want to know before trusting a result like this?" (open-ended, surfaces what the model missed)
A working rule for sample size: 15 to 25 real interviews usually settles claims about needs and objections within one audience. If you expect segments to differ, run that many per segment. Our guide to data saturation covers how to tell when you have enough.
What you get
By the end of the field period you have a report with themes, the number of participants behind each, and verbatim quotes for every finding. Open a finding to see its supporting responses.

Next-day uses:
- In the stakeholder meeting: a one-page claim ledger, with each synthetic claim marked confirmed, contradicted or missed, and the quotes that decided it.
- In the roadmap review: only the confirmed claims go forward. Contradicted claims become a short "we assumed, we checked" note that protects the team from relitigating them.
- In your tool decision: after three or four rounds, your own hit rate by claim type tells you where synthetic runs are cheap and safe (vocabulary, early hypotheses) and where they are not.
The same playbook elsewhere
Founders validating a pre-launch idea face the same question from the other side: how to test a pivot with real customers before committing. Product managers shipping AI features run a version of it too, checking what users did against what a launch plan assumed, as in our AI feature adoption playbook. Marketing teams do it with message testing: the model says the headline works, so someone has to ask real buyers. In each case the move is the same: write the claim down, ask real people, score the answer.
Why Koji for validating synthetic research

The reasons, each tied to something the product does today:
- Real people, structured by you. Every answer comes from a person you invited or recruited. You control who gets in with screening questions that qualify or rule people out.
- Mixed question types in one study. Ranking, scale, choice and open-ended questions sit together, so one pass can test a behavior claim, a priority claim and an intent claim.
- A quality score on every interview. Each interview is scored from 1 to 5, and only those scoring 3 or more use credits or enter the report. Since September 2026 the score is capped by how much the participant actually said, so clicking through or answering with filler no longer earns a top score.
- Evidence you can trace. Findings show how many participants support them, with the supporting quotes one click away, which is what a claim ledger needs.
- Fits where your team already works. You can start and read studies from an AI assistant through Koji's MCP integration, so the synthetic run and its check can sit side by side. See the MCP workflow guide for UX researchers.
If you want the longer discussion of when synthetic participants are safe to use, read our methodology guide to synthetic users and the comparison with synthetic user tools.
Questions you might have about AI interviews
"Isn't an AI interviewer the same problem as a synthetic user?" No. The AI is the moderator, not the participant. Every answer you analyze comes from a person. The risk to manage is different: whether those people are real and engaged. Koji scores each interview for quality, and panel recruitment lets you report a respondent you do not trust, with accepted reports refunded. Run your own checks too. A 2025 study in Sociological Methods & Research (Zhang, Xu and Alvero) found 34% of surveyed online participants used LLMs on open-ended answers, and reporting on a PNAS study by Sean Westwood says an AI agent evaded standard survey detection 99.8% of the time, though other researchers dispute how far that finding generalizes. A live conversation with follow-up questions is harder to fake than a tick box. See our post on AI-written survey responses.
"Can an AI interviewer probe like a trained moderator?" Open-ended questions get AI follow-up probing, and each question has its own follow-up setting. Preview the interview before launch and judge for yourself. See how the quality gate works.
"Where does participant data go?" Our AI interview data privacy guide covers how data is stored and who can see it.
When Koji may not be the right fit
Koji interviews capture what people say about what they did, and they do it well, but they do not record a screen or measure task time. If your claim is about whether people can complete a flow in a particular interface, pair it with a usability testing tool. And if your audience is so narrow that you cannot reach any real members of it, check the panel feasibility estimate before you promise anyone a validation round.
Metrics to track
- Claim hit rate: share of synthetic claims confirmed by real participants, by claim type
- Reversal rate: share where the real answer pointed the other way
- Misses per study: themes real participants raised that the synthetic run never did
- Time from synthetic run to validated decision
- Share of roadmap items backed by a validated claim
FAQ
Can I use synthetic users at all? Yes, for early exploration: vocabulary, draft hypotheses, rehearsing your discussion guide. Route any decision through real participants.
How many real interviews do I need to validate a synthetic study? For needs and objections within one audience, 15 to 25 is a practical starting point. Use more per segment if you expect segments to differ.
Should I show participants the synthetic findings? No. Ask the underlying question fresh and compare afterwards. Showing the summary anchors people.
Does Koji replace my synthetic tool? No. Koji runs the interviews with real people. Keep using the synthetic tool for exploration if it helps you.
Which plan do I need? Text interviews work on every plan. Voice interviews and CSV import of personal invite links need Interviews or Enterprise. Launching a panel recruitment needs a paid plan.
Start your validation round
Open the Concept Testing Interviews template, paste in the claims you want to test, and preview it before you send a single link. For more on the researcher workflow, see Koji for UX researchers and the full set of researcher playbooks.