Back to docs
Research Methods

Intention to Treat vs Per Protocol: Which Responses Belong in Your Analysis

Almost every research tool reports on completed responses only. That is a per-protocol analysis, it is the optimistic one, and nobody decided to run it. Here is how to choose an analysis population on purpose.

Answer first: reporting only on the people who finished is a per-protocol analysis, and the per-protocol analysis is systematically the flattering one. In clinical trials this is not a matter of opinion. The ICH E9 guideline Statistical Principles for Clinical Trials states that in superiority trials the full analysis set is used for the primary analysis "because it tends to avoid over-optimistic estimates of efficacy resulting from a per protocol analysis, since the non-compliers included in the full analysis set will generally diminish the estimated treatment effect." Customer research runs the per-protocol analysis by default, without naming it, because the analysis population is whatever the tool happens to put in the report.

This guide covers the two analysis principles, the four analysis sets that exist in a typical research study, how to choose between them before you know who dropped out, and how this differs from the three adjacent problems it gets confused with.

The two principles, defined properly

Intention to treat. ICH E9 defines it as "the principle that asserts that the effect of a treatment policy can be best assessed by evaluating on the basis of the intention to treat a subject (i.e. the planned treatment regimen) rather than the actual treatment given. It has the consequence that subjects allocated to a treatment group should be followed up, assessed and analysed as members of that group irrespective of their compliance to the planned course of treatment."

The equivalent in research is: analyse everyone you invited into the study as a member of the group you invited them into, regardless of how much of the study they completed.

Per protocol. ICH E9 describes the per protocol set - "sometimes described as the valid cases, the efficacy sample or the evaluable subjects sample" - as the subset of the full analysis set who are more compliant with the protocol, characterised by criteria such as completion of a pre-specified minimal exposure, availability of measurements of the primary variable, and the absence of major protocol violations including violation of entry criteria.

The guideline adds a procedural requirement that is the whole ballgame: "The precise reasons for excluding subjects from the per protocol set should be fully defined and documented before breaking the blind." Before you see the results. Not after.

Neither principle is right in the abstract. They answer different questions. Intention to treat estimates the effect of offering the thing. Per protocol estimates the effect of receiving it as designed. In a product research context: "what will happen if we ship this to everybody" versus "what happens for the people who engage with it properly." Both are legitimate questions and almost nobody says which one they asked.

The four analysis sets in a typical study

Every research study has these four populations, whether or not anyone names them. Numbers are illustrative of a common shape.

Analysis setDefinitionTypical nWhat it answers
InvitedEveryone who received the invitation1,000Effect of the offer, including reach failures
StartedEveryone who opened and answered at least one question340Effect among people willing to engage at all
CompletedEveryone who reached the end210The default report in most tools
Quality-passedCompleted and meeting a quality bar185The cleanest data, and the most selected group

The report you read almost certainly describes the fourth row, and the headline it produces is almost certainly better than the first row would produce, for a structural reason: the criteria that push people out of the later sets are correlated with the outcome being measured. People who find the product confusing abandon the interview about the product. People with nothing to say about a feature do not finish the feature survey. People with low-bandwidth connections drop out of voice interviews. Every one of those exclusions filters toward the engaged, the fluent and the satisfied.

That is the same structure as survivorship bias, with one important difference worth keeping straight: survivorship bias is about a sampling frame defined by an outcome, whereas this is about an analysis rule applied after the frame is fixed. Two other neighbours are also distinct. Survey universe definition is about who is eligible before fielding. Nonresponse bias is about the people who never started. This guide is about the people who started and did not finish - the ones you have partial data on and therefore a genuine choice about.

How badly this goes even in medicine

Hollis and Campbell surveyed every randomized controlled trial published in 1997 in the BMJ, the Lancet, JAMA and the New England Journal of Medicine (BMJ, 1999, 319(7211):670-674). Only 119 (48%) of the reports mentioned intention to treat analysis at all. Of those, 12 excluded any patients who did not start the allocated intervention and three did not analyse all randomised subjects as allocated - so a stated intention-to-treat analysis was frequently not one. Of the 99 that did appear to analyse according to allocation, only 34 explicitly stated it. 89 (75%) had some missing data on the primary outcome variable, and 29 (24%) had more than 10% of responses missing. The authors concluded that the approach "is often inadequately described and inadequately applied."

If the field with a formal guideline, a regulator and a reporting standard gets it wrong half the time, a product team with a dashboard has no chance unless it decides on purpose.

The modern framing: estimands and intercurrent events

The ICH E9(R1) addendum reframed the whole problem in a way that is unusually useful for product research. An estimand is "a precise description of the treatment effect reflecting the clinical question posed by a given clinical trial objective. It summarises at a population level what the outcomes would be in the same patients under different treatment conditions being compared. The targets of estimation are to be defined in advance."

The key move is the concept of an intercurrent event - something that happens after the study starts and affects the interpretation or the existence of the measurement. Discontinuing the assigned treatment is one. Switching to something else is one. The addendum is emphatic that these "are not to be thought of as a drawback to be avoided" - they happen in real life as they do in trials, "and their occurrence needs to be considered explicitly when defining the clinical question of interest."

Translated: abandoning your survey halfway is not a data-quality problem to be cleaned away. It is a finding, and your analysis has to state what it did with it. The estimand framework forces four strategies into the open, and each answers a different question:

Strategy for the drop-outThe question it answersWhen to use it in research
Treatment policy (count them regardless)What happens if we ship this to everyoneDefault for go/no-go decisions
While on treatment (use data up to the drop-out point)What happens while people are still engagedOnboarding and funnel studies
Principal stratum (restrict to those who would engage)What happens for the committed segmentPower-user and pricing work
Composite (treat dropping out as an outcome)Is dropping out itself the resultUsability and friction studies

The composite strategy is the underused one. If a third of participants abandon a concept test at the pricing question, the abandonment is the result, and reporting only on the two-thirds who pushed through actively destroys the finding.

Choosing your analysis set before you see the data

Three lines in the study brief. Write them at design time, alongside the sample size and the primary question.

  1. Primary analysis set. Name it: invited, started, completed, or quality-passed. For decision-driving studies, default to started - it is the closest practical equivalent to intention to treat, and it is honest about the fact that people who bounce off your questions are people who would bounce off your product.
  2. Exclusion criteria, fully specified. What makes a response invalid, written before fielding. Straight-lining on scale items, sub-threshold completion time, failed attention checks, off-topic open_ended responses. ICH E9 requires these be defined before unblinding; you should require them before fielding.
  3. The sensitivity analysis. Report the headline number on both the primary set and the completed set. If they differ materially, that difference is a finding about who abandons, and it belongs in the readout rather than in a footnote.

That third line is the cheapest credibility purchase in research. A readout that says "on completers the score is 7.4; on everyone who started it is 6.1, and the gap is driven by people who stopped at the pricing question" is a far stronger document than one reporting 7.4 alone - and it is the same data.

How Koji makes this a real choice instead of a default

Most tools discard or hide partial data, which forecloses the decision. Koji does not.

Partial and abandoned conversations are first-class. A Koji conversation carries an explicit status - active, completed, partial or abandoned - so the responses of someone who answered six questions and left are retained and attributable, not silently dropped from the denominator. That is the precondition for any analysis set other than completed, and it is what makes an intention-to-treat-style readout possible at all.

Per-question drop-off is visible. Because the interview is structured, you can see the exact question at which people stopped, which turns the composite strategy from theory into a chart. If abandonment clusters on one ranking item or one sensitive single_choice question, that is a specific, fixable finding.

A quality score exists as a separate axis from completion. Koji scores conversation quality on a 1-to-5 scale, and only conversations meeting the quality threshold consume credits. That gives you a documented, consistent exclusion rule rather than an ad hoc one invented after someone reads the results - which is exactly what ICH E9 asks for and exactly what hand-cleaning a spreadsheet cannot provide.

Structured questions make the primary variable unambiguous. The per-protocol criterion "availability of measurements of the primary variable" only means something if the primary variable is identified. With six question types available - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - you can designate a specific item as primary in the brief. See the structured questions guide.

Traditional survey platforms such as SurveyMonkey, Typeform and Qualtrics report completes by default and treat partials as a data-hygiene setting. Platforms like Koji keep the partials as evidence, which is the difference between choosing your analysis population and inheriting it.

Frequently asked questions

Which analysis set should be the default for product research?

Everyone who started. It is the closest practical analogue to intention to treat, it includes the people whose experience of your questions was bad enough to make them leave, and it produces the more conservative number - which is the right default when the readout drives a shipping decision. Report the completers number alongside it as a sensitivity analysis.

Is it ever right to report only on completers?

Yes, when the question is explicitly about the people who complete. Studies of power users, of a workflow that only exists after onboarding, or of a segment defined by sustained engagement all legitimately restrict to a compliant population. The requirement is that you say so, and that the exclusion rule was written before fielding rather than after.

How is this different from survey weighting?

Weighting corrects the composition of the people who did answer on the variables you weight on. Choosing an analysis set decides who is in the denominator at all. They compose: you pick a population, then weight within it. Neither can recover information from a group you never observed - see survey weighting.

What counts as an intercurrent event in a research study?

Anything after the study starts that changes the interpretation of the measurement. Abandoning the interview, switching device or modality mid-session, answering after an unusually long gap, or being exposed to the thing you are measuring by another route during fielding. The estimand framework asks you to name in advance how each is handled, rather than deciding once you can see which choice helps.

Does the completers-only bias have a predictable direction?

In superiority comparisons it usually flatters the thing being tested, which is exactly why ICH E9 makes the full analysis set the default for superiority trials. In equivalence or non-inferiority work the direction reverses and the full analysis set stops being the conservative choice - E9 says its role "should be considered very carefully" there. If you are trying to show two things are the same, see equivalence testing.

How much partial data do you need before including someone?

Set a minimum in the brief - typically the primary variable plus any variable used for segmentation. Someone who answered the primary scale question and left is analysable for the primary outcome. Someone who abandoned on the first screen contributes to the started count and to the drop-off finding but has no outcome data. Both facts belong in the readout.

Related Resources

Related Articles

Nonresponse Bias: How Missing Respondents Skew Your Data

Nonresponse bias occurs when the people who do not answer your survey differ systematically from those who do. Learn why a low response rate is not the same as bias, how to detect it, and how to reduce it.

Survey Completion Rate: How to Stop People Abandoning Your Survey Halfway

How to calculate survey completion rate, why it differs from response rate, what makes people quit mid-survey, and how to fix it — including why conversational formats finish stronger.

Survey Universe: How to Define Who Counts Before You Collect a Single Answer (2026)

The universe is the population whose opinion is actually relevant to your claim. Get it wrong and no sample size, weighting or analysis can rescue the study. A protocol, four documented failures, and how to enforce it at the door.

Survey Weighting: How to Correct a Skewed Sample

A practical guide to survey weighting — post-stratification, raking, and propensity weighting — plus how to calculate design effect and effective sample size, and when weighting cannot save your data.

Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story

Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.