Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies
Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.
Answer first: your data-quality controls almost certainly inspect respondents — speeders, straight-liners, duplicate IPs, fraudulent panellists. They almost certainly do not inspect the study. That asymmetry is expensive, because a leading question, a missing screener criterion, or an undefined analysis plan cannot be cleaned out of the data afterwards. The fix is a named reviewer, a fixed rubric, and a 30-minute time box between "the guide is written" and "the link is live". Studies that pass a structured design review before fieldwork produce data you can use; studies that skip it produce data you argue about.
This is the researcher-side complement to Survey Data Quality and Survey Fraud and Respondent Quality. Those articles defend against bad answers. This one defends against bad questions.
Why review beats cleanup
Questionnaire evaluation research has compared the main pretesting approaches head to head. In the classic comparison by Presser and Blair, expert review — a panel of experienced researchers reading the instrument and flagging problems — identified the largest and most consistent set of problems relative to cognitive interviews, behaviour coding, and a conventional pretest, and it did so at low cost. Different methods surface different problem types, which is why expert review complements rather than replaces piloting. But if you can only afford one control, a structured read by an experienced colleague is the highest-yield one available.
The economics are simple. A defect caught before fieldwork costs an edit. The same defect caught after fieldwork costs the whole study, because the only remedy is to re-field. And a defect never caught at all costs more than that: a decision made on data that measured the wrong thing.
Note what a review gate is not. It is not a pilot — a pilot study tests the instrument on real respondents, and it comes after the review. It is not approval to spend, and it is not a manager signing off on the conclusions. It is one qualified person checking that the design can answer the question it claims to answer.
The two-tier model
A universal gate on every study will be resented and then quietly ignored. Tier it.
Tier 1 — self-review (every study, 10 minutes). The author works through the rubric alone and records the answers in the brief. No second person, no waiting. This alone catches a surprising share of defects, because a checklist forces attention onto the parts of a design that feel finished.
Tier 2 — peer review (triggered, 30–45 minutes). A second qualified reviewer reads the design and returns a decision. Trigger it on any of:
| Trigger | Why |
|---|---|
| The result will inform an irreversible decision (pricing, launch, sunset) | Cost of error is high |
| Findings will be published externally or used in marketing | See claim substantiation |
| The study is run by someone outside the research team | The core case for a gate under democratized research |
| Participants are in a regulated or vulnerable population | Legal and ethical exposure |
| Sample cost exceeds a threshold you set | Re-fielding is expensive |
| The instrument is new rather than a template reuse | Templates have already been reviewed |
Everything else gets Tier 1 only. Roughly 20–30% of studies hitting Tier 2 is a healthy steady state in most teams.
The review rubric
Seven sections. A reviewer works top to bottom; the order matters, because a failure at section 1 makes sections 3–7 irrelevant.
1. The decision the study serves
- Is there a named decision, a named owner, and a date by which the answer is needed?
- Would a plausible finding actually change what the owner does? If every possible result leads to the same action, stop — the study is theatre.
- Has this question already been answered in the repository? A five-minute search prevents a five-week study.
2. Method fit
- Does the method match the question type? Attitudes and reasons need conversation; incidence and preference need structured measurement.
- If the claim is quantitative, is there a structured item producing a countable answer, or only open text?
- If the question is about behaviour, does the design capture behaviour rather than a self-report about behaviour?
3. Sample and universe
- Who is the finding meant to generalise to, and does the recruitment reach them?
- Does the screener exclude the people whose answer would be inconvenient? Selecting on the outcome is the single most damaging design error — see Sampling Bias and Survivorship Bias.
- Is the target base size sufficient for the cuts the analysis plan promises? Subgroup analysis on n=8 is not analysis.
- Are quotas defined, and is there a plan for what to do if one fills late? See Quota Sampling.
4. The instrument
- Read every question aloud. Any item you stumble over will confuse a participant.
- Flag leading, loaded, and double-barrelled items — the wording rules are the checklist here.
- Check sequence for question order effects: general before specific, unaided before aided, sensitive items late, sponsor identity last.
- Check scale symmetry, labelled endpoints, and that no scale silently changes direction mid-instrument.
- Check length against modality. An instrument that takes 25 minutes will be abandoned or straight-lined, and you will blame the respondents.
- For AI-moderated studies: are probing instructions specific enough to be useful and bounded enough to avoid an endless interview?
5. The analysis plan, pre-specified
- Write the headline chart before fieldwork. If you cannot sketch it, the instrument is not ready.
- State which comparisons will be tested and which subgroups will be reported. Pre-specifying is what separates a finding from a fishing expedition — see Statistical Significance in Survey Research.
- For qualitative work, name the coding approach and who will code. If two people will code, define the agreement check up front — see Inter-Rater Reliability.
6. Ethics, consent, and legal
- Is consent appropriate to the data being collected, and is recording handled lawfully? See Interview Recording Consent Laws.
- Does anything in the study touch a vulnerable population, a health context, or children? Route accordingly.
- Is the incentive proportionate — enough to be fair, not so much that it becomes undue influence?
- Is retention specified, and does anyone need to approve it? See Research Data Retention and Deletion.
7. Delivery
- Who receives the result, in what format, by when?
- Is there a plan for the finding to reach the repository in a reusable form rather than dying in a slide?
Making the decision explicit
A review that ends in a conversation ends in nothing. Force one of three outcomes:
| Decision | Meaning | Typical share |
|---|---|---|
| Approve | Launch as written | 30–40% |
| Approve with changes | Author applies listed changes and launches; no re-review | 45–55% |
| Rework | Design has a section 1–3 failure; resubmit | 10–20% |
"Approve with changes" is the workhorse. Requiring a second review round for wording fixes is what makes teams abandon gates. Reserve "rework" for failures of purpose, method, or sample — the three things that cannot be patched.
Two staffing rules keep the gate honest. The reviewer should not be the author line manager, because seniority converts feedback into instruction. And reviewers should rotate: being the reviewer is how less experienced researchers learn design, and it distributes the load so the gate does not depend on one person availability.
Track the defects, then fix the templates
Log every defect the gate catches, tagged by rubric section. After thirty or so reviews the pattern is unmistakable, and it tells you what to fix systemically rather than one study at a time.
| Defect category | Typical systemic fix |
|---|---|
| No decision the study would change | Require a decision field in the research brief |
| Screener selects on the outcome | Add a worked counter-example to the screener guide |
| Leading or double-barrelled wording | Add reviewed question banks per study type |
| Base too small for promised cuts | Add a base-size calculator step to the brief |
| No analysis plan | Make the headline-chart sketch a required field |
| Consent mismatch | Ship consent templates by study type |
A gate that only catches defects is a tax. A gate that feeds template improvements pays for itself, because next quarter fewer studies arrive with the same problem.
The 10-minute smoke test
After review and before full fieldwork, run three to five real responses and read them completely. Not the summary — the actual sessions. You are checking four things:
- Did people understand the question the way you meant it?
- Did anyone finish suspiciously fast, or abandon at a particular item?
- Do the answers to your key item actually vary? Zero variance means the question cannot discriminate.
- Can you compute your headline number from these responses? If the pipeline breaks on five, it breaks on five hundred.
This is where AI-moderated research changes the economics. With a traditional panel study, a soft launch means booking fieldwork twice. With Koji, you publish the study, watch the first interviews arrive in real-time reports, and either continue or pull the link and fix the guide — usually within an hour of launch. The quality gate means low-scoring conversations do not consume credits, so an aborted smoke test costs almost nothing.
What Koji reviews faster than a document
A study design in a Google Doc has to be read imaginatively — the reviewer simulates what a participant would experience. A study built as a structured artefact can be inspected directly:
- Six structured question types (
open_ended,scale,single_choice,multiple_choice,ranking,yes_no) make section 2 and section 5 of the rubric mechanical. The reviewer can see at a glance whether a quantitative claim has a countable item behind it, or whether the study is entirely open text and therefore cannot produce a number. - Explicit probing configuration per question shows the reviewer exactly how deep the AI will go, rather than leaving depth to a moderator improvisation on the day.
- Consistent execution removes an entire class of defect that human-moderated research reintroduces at every session: the AI asks what was approved, in the approved order, to every participant. Nothing drifts between session one and session twenty — the variance problem described in Interviewer Bias.
- A shareable test link lets the reviewer take the interview themselves in five minutes. Reviewers who experience the instrument catch things that reviewers who read it never do.
- AI-drafted guides still need the gate. Generating a first draft in seconds is a real speed advantage, and it makes the review step more important, not less: the fastest way to ship a beautifully worded study that measures the wrong thing is to accept a generated draft without asking whether it serves the decision.
Start with 10 free credits, build the study, and send the test link to your reviewer before you send the real one to anybody else.
Frequently asked questions
Is peer review the same as a pilot study?
No, and the order matters. Peer review is an expert reading the design before anyone is contacted; a pilot is a small live run that tests the instrument on real respondents. Review first — piloting a design with a purpose failure just wastes participants.
Who should review when there is only one researcher?
Use the Tier 1 self-review rubric on every study, and trade reviews with a researcher at another company or in a professional community for Tier 2 triggers. Failing that, ask a smart colleague outside research to take the interview and tell you where they hesitated. An imperfect reviewer beats no reviewer substantially.
How do we stop the gate from slowing everything down?
Time-box it to 30–45 minutes, allow "approve with changes" to close a review without a second round, and only trigger Tier 2 on defined criteria. Gates fail when they become open-ended conversations or when every study needs one.
Does this apply to studies run by non-researchers?
That is the strongest case for it. Democratizing research without a design gate produces volume and reduces reliability. A lightweight, fast review is what makes self-service research trustworthy — pair this with the self-service research programme model.
What if the reviewer and the author disagree?
Distinguish preference from defect. A reviewer may write a question differently; that is not a finding. Escalate only failures against the rubric — no decision served, method mismatch, sample selected on the outcome, leading wording, missing analysis plan. Everything else is advice the author may decline.
Should AI-generated study designs be reviewed differently?
Same rubric, extra attention to sections 1 and 3. Generated drafts are usually well-formed and fluent, which makes purpose and sample errors easier to miss — the prose is not the part that fails. Read the decision statement and the screener first, before you are charmed by the questions.
Related resources
- Structured Questions Guide — the six question types that make a design inspectable
- Survey Data Quality — the respondent-side half of a quality programme
- Pilot Study Guide — what to run after the design passes review
- How to Write Unbiased Survey Questions — the wording checklist for rubric section 4
- Research Democratization Playbook — scaling research without lowering the floor
- Research Brief Template — where the review answers should live
Related Articles
Pilot Study in User Research: How to Pre-Test Your Methodology Before Going Live (2026)
A pilot study is a small-scale rehearsal of your full research project that catches broken questions, biased prompts, and recruiting issues before they invalidate your real data. Learn when to run one, how many participants you need, what to test, and how AI-moderated platforms compress the pilot loop from weeks to hours.
Research Brief Template: How to Define Your Research Before You Start
A complete research brief template with sections for problem context, participant profile, methodology, and success criteria — the foundation of any effective user research project.
Research Democratization: The 2026 Playbook for Scaling User Research Across Your Whole Organization
A complete playbook for democratizing user research without sacrificing rigor — what democratization actually means in 2026, what fails, what works, and how AI-native platforms like Koji make scale safe.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Data Quality: How to Detect and Prevent Bad Responses (2026)
The threats that corrupt survey data — straightlining, speeding, bots, fraud, and inattentive respondents — how to detect and prevent each, and why conversational AI interviews are structurally resistant to the junk that plagues panel surveys.
How to Write Unbiased Survey Questions: Avoiding Leading, Loaded & Double-Barreled Questions
A practical guide to question wording — the biggest hidden source of bad data. Learn to spot and fix leading, loaded, double-barreled, and assumptive questions, with real research examples and a pre-launch checklist.