Agreed-Upon Procedures: The Research Deliverable That Deliberately Has No Recommendation
Sometimes the most useful research output is a list of pre-agreed procedures, the factual findings each produced, and no conclusion at all. How to run an agreed-upon procedures study, and when it is the wrong choice.
Some research questions should be answered with facts and no recommendation. When the decision depends on commercial context the researcher does not hold, the most useful and most defensible deliverable is a list of pre-agreed procedures, the factual findings each one produced, and nothing else. The accounting profession formalised this decades ago and calls it an agreed-upon procedures engagement. Product research has no equivalent, which is why researchers keep getting drawn into recommendations they are not positioned to make, and then blamed for the outcome.
The short version
An agreed-upon procedures study is one where the requester and the researcher agree in advance on exactly what will be done, the researcher does it, and the report states what was found. No conclusion. No recommendation. No "therefore we should".
This runs directly against the standard advice, including most of the advice on this site, which is that a study that does not change a decision was a waste of money. That advice is right most of the time and wrong in a specific, common case: when the researcher would have to supply business judgement they do not have in order to reach a conclusion. In that case a recommendation is not added value, it is added noise wearing a lab coat.
What the standard actually says
ISRS 4400 (Revised), the International Standard on Related Services covering agreed-upon procedures engagements, has been in effect for engagements agreed on or after 1 January 2022. Two of its provisions are worth reading closely, because they are more precise than anything product research has written on the subject.
First, on what the engagement is not. The standard states that an agreed-upon procedures engagement is not an audit, review or other assurance engagement, and does not involve obtaining evidence for the purpose of expressing an opinion or an assurance conclusion in any form.
Second, on what a finding is. The standard defines findings as the factual results of the procedures performed, states that findings are capable of being objectively verified, and says explicitly that references to findings exclude opinions or conclusions in any form as well as any recommendations the practitioner may make.
The operational test is in the application material: findings are capable of being objectively verified, which means that different practitioners performing the same procedures are expected to arrive at equivalent results.
That sentence is the whole discipline. If two competent researchers running your stated procedures on your stated sample would produce different write-ups, you have not produced findings. You have produced interpretation, and interpretation needs a different kind of report, a named author, and a stated basis.
Why product research needs this mode
Four situations come up constantly where the recommendation-shaped report actively damages the decision.
The decision hinges on numbers you do not have. "Should we build the enterprise SSO tier?" depends on pipeline, contract values, the cost of the build and the opportunity cost of the alternative. A researcher can establish how many interviewed buyers named SSO as a blocker, in what words, at what stage. Turning that into "we should build it" requires the commercial half of the equation. Report the half you own.
The requester is more senior and better informed about the trade-off. An executive asking a specific factual question usually already has the surrounding context. Supplying a recommendation on top invites them to argue with the recommendation and ignore the facts, which is the worst of both outcomes.
The finding will be contested. When research feeds a negotiation, a pricing change, a regulatory filing or a partner discussion, a factual-findings report survives cross-examination and an interpretive one does not. This is the same instinct behind survey evidence that holds up in court: the more the output must withstand an adversary, the more it should confine itself to what was done and what was found.
Multiple teams will use the same data differently. One study, four consumers, four legitimate conclusions. Publishing one team's conclusion as "the finding" quietly disenfranchises the other three.
How to run one
Agree the procedures in writing, before fieldwork. This is the defining step and the one people skip. The requester must acknowledge that the procedures are appropriate for their purpose; ISRS 4400 makes that acknowledgement a condition of accepting the engagement, and it protects both sides. A procedure is a sentence with no adjectives in it: "Ask each participant to rank the five proposed capabilities in order of importance to their renewal decision, using a ranking question, and report the mean position of each."
Define the population and how it was drawn. "Twenty current customers on the Growth plan, with more than ninety days tenure, invited by email in the order returned by the export, until twenty completed."
Perform exactly those procedures. Deviations are permitted, but they get reported as deviations. Quietly adding a question because it seemed interesting breaks the objective-verifiability property that gives the report its authority.
Report findings and exceptions, including the uncomfortable ones. ISRS 4400 requires the report to include the findings from each procedure performed, including details of exceptions found. An exception is any case where the procedure could not be performed as agreed: four participants declined to rank, two sessions ended early, one respondent misread the scale. Exceptions belong in the report at the same prominence as the results.
Stop. No summary paragraph beginning "This suggests". No slide titled "So what". If the requester wants your view, they can ask for it in the meeting, on the record, as your view.
The report template
Agreed-Upon Procedures Report: Renewal Blockers, Growth Plan
Engaging party: VP Customer Success
Date agreed: 3 March 2026
Fieldwork: 10-12 March 2026
Procedures agreed
P1 Interview 20 Growth-plan customers with >90 days tenure...
P2 Ask each to rank five capabilities by importance to renewal...
P3 Ask each a yes_no question on whether they have evaluated an alternative...
Findings
P1 22 interviews invited, 20 completed, 2 abandoned mid-session.
P2 Mean rank position: Reporting 1.8; SSO 2.4; API 3.1; ...
P3 9 of 20 answered yes. Of those 9, 6 named the same competitor...
Exceptions
P2 3 participants ranked only their top three and declined the rest.
Their partial rankings are excluded from the means above.
This report contains no opinion, conclusion or recommendation.
That last line is not decoration. It is the sentence that stops a reader from importing a conclusion you did not make, and it is the sentence that protects you when the decision goes badly.
Structured questions make this mode practical
The reason agreed-upon procedures have been impractical in product research is that qualitative research does not naturally produce objectively verifiable results. Two analysts coding the same twenty transcripts will not produce equivalent write-ups, which fails the standard's core test.
Koji's structured questions solve this directly, and it is one of the clearest cases where the platform's design maps onto a governance need. Five of the six question types produce results that are objectively verifiable by construction: scale yields a distribution, single_choice and multiple_choice yield frequencies, ranking yields mean positions, and yes_no yields a count. Anyone re-running the same procedure on the same responses gets the same numbers. The sixth type, open_ended, is the one that does not, which is exactly why it belongs in the interpretive part of your work rather than in the findings section of an agreed-upon procedures report.
This is also where AI-moderated interviews earn their place over a conventional survey tool. A survey platform can give you a clean ranking, but it cannot ask the participant why they put reporting first, and it cannot tell whether the person was engaged or clicking through. Koji runs the structured question and then probes conversationally on the answer, so the same session produces both the objectively verifiable finding and the verbatim material you will need later when someone asks what the number means. Every session carries a quality score from 1 to 5, which is itself an objectively verifiable input to whether a response should be included. See the structured questions guide for how each type is presented and analysed.
Practically: put the agreed procedures on the structured questions, and let the conversational layer collect the colour separately. Report the first. Offer the second on request.
When this mode is the wrong choice
It is the wrong choice more often than it is the right one, and using it as a hiding place is a real risk.
- When the requester wants your judgement and is entitled to it. If you have run forty studies in this domain and they have run none, withholding a view is not neutrality, it is abdication.
- When the question is genuinely exploratory. You cannot pre-agree procedures for "why are trials not converting" because you do not yet know what to ask. Discovery work needs the opposite posture; see continuous discovery.
- When it is being used to avoid delivering bad news. A factual-findings report is not a way to avoid saying the thing. If the finding is that the feature nobody wants is the one already being built, the finding says so, plainly, in the findings section.
The honest test: are you declining to conclude because the conclusion requires context you lack, or because concluding would be uncomfortable? Only the first is a legitimate use.
The failure mode: the smuggled conclusion
The most common way this goes wrong is that a conclusion is smuggled in through presentation. Ordering the findings so the intended one comes first. Choosing a chart that makes one bar look decisive. Bolding one number. Titling a section "The reporting problem". Each of these is an opinion expressed without accountability, and it is worse than an openly stated recommendation because nobody can argue with it.
The remedy is a review pass by someone who was not in the study, looking only for implied conclusions. This slots naturally into an existing peer review gate. Give the reviewer one instruction: mark every place where the report tells the reader what to think.
Frequently asked questions
Is this just a fancy way of saying "here is the data"?
No. "Here is the data" is a dump with no defined procedures, no stated population, no exceptions and no prior agreement about what would be done. The agreed-upon procedures form is a commitment made before fieldwork about exactly what will be performed, which is what makes the findings verifiable and what stops the scope shifting once early results appear.
Does this let researchers off the hook?
The opposite. A recommendation can be defended with judgement and seniority. A findings-only report is checkable line by line, so any sloppiness in sampling, question wording or execution is exposed. Teams generally find this mode harder to produce, not easier.
How do I stop stakeholders from just asking for the recommendation anyway?
Give it to them verbally, in the meeting, clearly labelled as your view rather than as a finding. The distinction that matters is between what the study established and what you personally think, and keeping those in different containers is the entire point. Many researchers find the conversation goes better once the two are separated, because the facts stop being negotiable.
Can a study be part agreed-upon procedures and part interpretive?
Yes, and this is the usual arrangement in practice. Keep them in separate, clearly titled sections, in that order, and never let interpretation appear inside the findings section. Readers will honour the boundary if you do.
What about the participants' quotes? Are those findings?
A verbatim quote with its source and date is a factual result, so it can sit in the findings. A selection of quotes chosen to illustrate a theme is interpretation, because a different researcher would choose differently. If you include quotes in the findings section, state the selection rule you applied, such as every mention of pricing in order of occurrence.
Does an AI-moderated study make findings more or less verifiable?
More, on the mechanical dimension, and it is the main practical argument for running this mode on a platform like Koji. The questions asked are identical across participants unless the AI probes, the full transcript is retained, the probes themselves are on the record, and the structured question results are computed the same way every time. A human moderator introduces unrecorded variation in wording and emphasis that makes exact re-performance impossible.
Related Resources
- Structured Questions in AI Interviews - the six question types, and which ones produce verifiable findings
- Research Peer Review - the QA gate that catches smuggled conclusions
- Presenting Research Findings to Stakeholders - when interpretation is the deliverable
- Survey Evidence in Court - research built to survive cross-examination
- How to Write a Research Brief - where the agreed procedures get recorded
- Research Debrief - separating what was found from what to do about it
Related Articles
How to Write a Research Brief: Templates, Examples, and AI-Assisted Generation
A step-by-step guide to writing an effective user research brief. Covers the 7 essential components, participant targeting, methodology selection, and how Koji's AI generates briefs automatically from a plain-language goal.
Presenting Research Findings to Stakeholders
Learn how to present qualitative research findings effectively — from storytelling with data and using participant quotes to structuring reports for executives, product teams, and designers.
Research Debrief: How to Synthesize and Share Findings After Every Study
A complete guide to running research debriefs — turning raw interview data into stakeholder-ready insights. Includes a debrief template, synthesis techniques, and how Koji automates the hardest parts.
Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies
Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Evidence in Court: Daubert, FRE 702, and Research That Survives Cross-Examination
The standard courts apply to survey evidence is a free quality bar for ordinary product research. Here is the checklist, the attacks it defeats, and why the control group is a legal instrument as much as a statistical one.