Back to docs
Research Operations

The Research Expectation Gap: Why Stakeholders Are Disappointed by Studies That Were Done Right

Auditing measured its own credibility gap and found only 16% was sub-standard work. Half was scope. Here is how that decomposition changes what research teams should fix.

Answer first: when stakeholders are disappointed by research, the reflex is to assume the work was not good enough. Auditing tested that assumption empirically and found it mostly false. Porter's study of the audit expectation-performance gap decomposed it into three parts: 34% unreasonable expectations, 50% deficient standards, and 16% sub-standard performance. Only one part in six was practitioners doing their job badly. Half was the profession's own definition of the job being narrower than what a reasonable person expected it to cover. If research follows the same distribution - and the structural argument says it should - then raising rigour addresses the smallest component, and the other 84% is a conversation nobody is having.

This is the article that undercuts the three that precede it, and it is the reason they need it.

The gap, decomposed

Brenda Porter's 1993 study in Accounting and Business Research did something the endless commentary on the "audit expectation gap" had never done: it measured the parts. Rather than treat the gap as one undifferentiated complaint, she separated society's expectations of auditors from what auditors can reasonably deliver, and then separated what is reasonably expected from what auditors are actually required and perceived to do.

ComponentShareWhat it meansWho can fix it
Unreasonable expectations34%The public expects things no auditor could deliver at any priceNobody - it can only be communicated
Deficient standards50%Things reasonably expected that the profession's own standards do not requireThe profession, by widening scope
Sub-standard performance16%Auditors failing to meet the standards that do existIndividual practitioners

Porter also argued the label itself was wrong, and proposed "audit expectation-performance gap" - the gap between society's expectations of auditors and society's perceptions of auditors' performance. The renaming matters, because it locates the gap between two sets of beliefs rather than inside the work.

Reading the three components into research

Unreasonable expectations (roughly a third). Research cannot tell you what people will do in a market that does not exist yet. It cannot deliver certainty from twelve interviews. It cannot resolve a disagreement between two executives who want different things, because that is a values conflict wearing an evidence costume. No budget closes this component. The only available move is to say it out loud, early, in the brief - which is what an explicit assurance level does, as covered in levels of assurance in research.

Deficient standards (roughly a half, and the one nobody works on). This is the component that should reorganise a research team's priorities. It is not that the study was done badly. It is that the standard scope of the team's standard study is narrower than what a reasonable stakeholder assumed it covered.

Concrete examples, all common:

  • The study covered current users; the stakeholder assumed it spoke for prospects
  • The study covered the happy path; the stakeholder assumed it covered the error states
  • The study measured comprehension; the stakeholder heard willingness to pay
  • The study ran in one market; the stakeholder read it as global
  • The study tested the prototype; the stakeholder took it as a verdict on the strategy

Every one of those is a scope question that could have been settled in the brief and was not - and note that in each case the researcher did nothing wrong by the standards of the work as scoped. That is precisely Porter's point. The gap is real, the disappointment is legitimate, and the remedy is not more rigour.

Sub-standard performance (roughly a sixth). Leading questions, thin samples, unrepresentative recruits, analysis that skipped the disconfirming cases. This is the component the entire research quality literature addresses, including the three articles this one sits alongside. It is worth fixing. It is not where most of the disappointment comes from.

Why rigour cannot close a gap made of beliefs

The three preceding articles in this cluster - on assurance levels, corroboration and source independence, and research independence - all rest on one shared assumption: that the value of research is determined by the quality of the evidence behind it. Porter's decomposition says that assumption is true for about a sixth of the problem.

The reason is that the reader's expectation is formed somewhere else entirely. It is formed by the roadmap they have already committed to, by the last research readout they half-remember, by what the word "research" means in their previous company, and by how much of the quarter is left. None of those inputs is affected by your sample size. A methodologically flawless study lands in a mind that has already decided what the study is for, and if those two objects do not match, the study is experienced as a failure regardless of its internal quality.

This is also why the standard defensive response makes things worse. When a researcher answers disappointment by explaining the methodology, they are addressing the 16% component to an audience whose complaint lives in the 50% component. The stakeholder is not saying the work was sloppy. They are saying it did not cover what they needed covered - and they are usually right about that.

What auditing actually did about it

The profession's response is instructive because it was not "try harder". It was to fix the wording, and then to widen the scope.

Audit reports are now built from prescribed components: a scope paragraph describing what was examined, a basis-for-opinion paragraph, a statement of respective responsibilities separating management's job from the auditor's, and an explicit statement that reasonable assurance is a high level of assurance but is not a guarantee that an audit will always detect a material misstatement. That last clause exists solely to attack the unreasonable-expectations component, in fixed language, in every report, whether or not anyone complained.

And on the deficient-standards half, the profession repeatedly widened what an audit must cover - going-concern assessment, fraud-risk procedures, internal control over financial reporting, and in recent years the expanded reporting of key audit matters. Those changes came from accepting that the public's expectation was reasonable and the standard was too narrow.

Both moves are available to a research team, and neither requires a regulator.

A four-line preamble that does most of the work

The cheapest available version of the fixed-wording response is four lines at the top of every readout. It is not a limitations section - those go at the end, where nobody reads them, and by then the headline has already landed.

LineExample
What this study covered18 interviews with existing admin users in the UK and Germany, February 2026
What it did not coverProspects, non-admin roles, and any market outside the EU
What it is entitled to concludeLimited assurance: nothing surfaced that contradicts the proposition, and three risks did
What would change the answerAny of the three named risks appearing in a prospect sample

The fourth line is the one that changes the conversation, because it converts a verdict into a live hypothesis and gives the stakeholder something to commission next rather than something to argue with now.

Renegotiating standards instead of defending them

The larger move is to treat the deficient-standards component as a backlog rather than an insult. Once a quarter, list the questions stakeholders keep asking that your standard study does not answer, and decide explicitly for each one: widen the standard scope to include it, or declare it permanently out of scope and say why. Both are acceptable. Leaving it undeclared is what produces the gap, because an undeclared boundary is assumed by the reader to be inside the scope.

Two practical notes. First, the register of what you decided not to cover is as valuable as the studies themselves, and it is the natural companion to the study register described in publication bias in product research. Second, widening scope has a cost, and the honest response to "we want prospects covered too" is a number, not a yes. That is the conversation the decomposition unlocks: it moves the argument from the researcher's competence to the programme's coverage, where it can actually be resolved with budget.

How Koji shrinks the deficient-standards half

The 50% component is a coverage problem, and coverage is mostly a cost problem. That is where an AI-native platform changes the arithmetic rather than the rhetoric.

Scope that used to be unaffordable becomes routine. The reason a study covered only current users, only the happy path, or only one market is almost never that the researcher thought those boundaries were correct. It is that each additional cell meant more recruiting, more moderator hours, and more analysis time. When interviews run 24/7 without a moderator and analysis arrives as a by-product, adding the prospect cell or the second market stops being a trade against the deadline. Most deficient-standards complaints are, underneath, budget constraints that were never surfaced as choices.

The scope statement can be generated from the study rather than remembered. Because the brief, the screener, the structured questions, and the completed interviews are all one object, the "what this covered / what it did not" preamble is a description of the study's actual configuration rather than a sentence someone has to remember to write honestly at 6pm.

Structured questions make the entitlement checkable. With six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - a stakeholder can see whether the study measured comprehension or measured willingness, rather than inferring it from a narrative summary. Most "the study measured X and I heard Y" failures are ambiguity about what was actually asked, and a structured instrument removes the ambiguity. See the structured questions guide.

Real-time results let scope be renegotiated while it still costs nothing. When a stakeholder sees themes emerging mid-field and realises the study is not going to answer their real question, the cheap fix is available - a second cell, another segment - rather than the expensive one, which is a rerun next quarter. The rules for doing that without corrupting the study are in changing a study mid-field. Compare that to a survey platform, where scope is frozen at launch and the gap only becomes visible in the export.

Honest objections

Sometimes the researcher genuinely did the work badly. Sometimes, and that is the 16%. It is a real component with a real remedy, and this article is not an argument for ignoring it. It is an argument against treating it as the whole problem when it is the smallest of three.

Naming limits up front can become a way to dodge accountability. It can, and the guard is that the preamble is written before fielding, not after the results disappoint. A limitation declared in the brief is scoping. The same sentence added to the readout after a hostile question is an excuse, and everyone can tell the difference.

Porter's numbers come from auditing in New Zealand in 1989, not from product research. True, and the specific percentages should be treated as illustrative rather than as measurements of your team. What transfers is the decomposition itself - that a credibility gap has a reasonableness component, a scope component, and a performance component, and that they have different owners and different remedies. Any research team can estimate its own split by asking, for the last ten disappointing readouts, which of the three each one was. Most teams find the same shape, and the exercise takes an hour.

Frequently asked questions

What is the expectation gap?

It is the gap between what an audience expects a professional service to deliver and what they perceive it actually delivered. Porter reframed it as the expectation-performance gap and split it into three parts: unreasonable expectations (34%), deficient standards (50%), and sub-standard performance (16%). The point of the split is that the three parts have different causes and different owners.

Why does better research not close the gap?

Because only about a sixth of the gap is caused by work falling short of existing standards. Half is caused by the standard scope of the work being narrower than what a reasonable stakeholder expects, and a third by expectations no study could satisfy. Improving rigour addresses the smallest component, which is why teams that keep raising quality still feel under-appreciated.

How do we tell which component a specific complaint belongs to?

Ask what the stakeholder wanted the study to cover. If no study could have delivered it, the complaint is unreasonable expectations. If a study could have delivered it but yours was not scoped to, it is deficient standards. If your study was scoped to deliver it and did not, it is sub-standard performance. Only the third is a quality problem.

Where should limitations be stated?

At the top, before the findings, and written before fielding. A limitations section at the end of a deck is read after the headline has already been believed, and a limitation added after a hostile question reads as an excuse. Four lines - what was covered, what was not, what the study is entitled to conclude, and what would change the answer - do most of the work.

Is this an argument for lowering expectations?

No. It is an argument for locating them accurately. Two of the three components are addressed by communication and by explicitly widening scope, and widening scope raises what research delivers rather than lowering what is expected. The only component that calls for expectations to come down is the one where the expectation was genuinely impossible.

How often should research scope be renegotiated?

Roughly quarterly. List the questions stakeholders keep asking that your standard study does not answer, then decide for each whether to widen the standard scope or declare it permanently out of scope with a reason. An undeclared boundary is assumed by readers to be inside the scope, which is what generates the deficient-standards half of the gap.

Related Resources

Related Articles

Activating Research Insights: Turn Findings Into Product Decisions

A practical guide to insight activation — the discipline of ensuring research findings actually drive product decisions. Covers why 40-60% of insights are never used, the 4-stage activation framework, decision-ready report formats, and how AI-native research platforms close the loop in real time.

Changing a Study While It Is Running: Pre-Planned Adaptations vs Protocol Amendments

Every other discipline in research assumes the instrument holds still. In practice teams rewrite questions, add segments and drop arms while fielding. Here is which mid-study changes are free, which are expensive, and which destroy the study.

Corroboration in Research: Why Three Sources Saying the Same Thing Can Be One Source

Evidence from multiple sources only multiplies confidence when the sources are independent. Four ways research sources secretly share an origin, and a ten-minute test for catching it.

Publication Bias and the File-Drawer Problem in Product Research: Why Your Evidence Base Only Remembers the Studies That Worked (2026)

Publication bias is not an academic curiosity. In product research it is worse, because nobody rejects your null study - you simply never write it up. Learn how big the file drawer is, what it does to your confidence, and how to build a study register that closes it.

Levels of Assurance in Research: How Much Confidence a Study Can Honestly Support

Auditing defines three levels of assurance - reasonable, limited, and none. Research reports use one voice for all three. Here is how to pick and state the level before you field a study.

Research Independence: Why the Team That Built the Feature Should Not Grade It

The five threats to independence from professional ethics codes, applied to product research - and why structural independence is a different problem from cognitive bias.

Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies

Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.