Key Informant Interviews: Why One Person Cannot Speak for a Whole Company (2026)
A key informant's report correlates only .612 with an independent source - and just .502 on internal process questions. What the measurement literature says about interviewing one person about a whole organization.
When you interview one person about their company and write down what they say as a fact about the company, you are running a key informant design. It is a named method with a measured error rate, and the error rate is worse than almost anyone using it assumes. Across a meta-analysis of 127 organizational studies, a key informant's report correlated .612 with an independent second source. For questions about what happens inside the organization - how work actually flows, which priorities really won, who uses what - the correlation fell to .502, and roughly a third of the variance in the answer was attributable to which person you asked rather than to anything about the company.
That is the whole problem in one paragraph. A single informant is not a window onto an organization. It is one measurement, taken with an instrument whose calibration you have not checked.
What a key informant design actually is
A key informant is someone who reports on a collective rather than on themselves. The distinction is sharp and it changes the error structure completely:
- "How often do you export a report?" - a self-report. The respondent is the unit being measured.
- "How often does your team export reports?" - a key informant report. The respondent is an instrument pointed at a unit they are not.
Nearly every B2B research program runs on the second kind of question without ever naming it. You ask the VP of Operations how the rollout went. You ask the admin whether the team finds the new workflow confusing. You ask the champion why the company chose you. In each case one person is being asked to summarize a distributed reality they observed partially, from a fixed position, with their own stake in the answer.
The method is old and respectable - it came out of community sociology and organizational research, and it is still how most survey work on firms gets done, because the alternative (census every employee) is usually impossible. The problem is not the method. The problem is that product and research teams use it without inheriting any of the discipline the methodological literature attached to it.
The evidence: how accurate is one informant, really?
The most useful single source here is Homburg, Klarmann, Reimann and Schilke's 2012 study in the Journal of Marketing Research ("What Drives Key Informant Accuracy?", 49(4), 594-608). They did two things: a meta-analysis of 127 studies from six major marketing and management journals, covering 431 constructs and 9,814 observations, plus a second study on eight multi-informant datasets.
The headline reliability numbers - the correlation between what a key informant said and what an independent second source said - break down like this:
| What the question is about | Studies | Constructs | Mean correlation |
|---|---|---|---|
| Overall | 105 | 320 | .612 |
| Performance (revenue, growth, hard outcomes) | 50 | 84 | .764 |
| Intraorganizational (internal process, structure, priorities) | 65 | 174 | .502 |
| Interorganizational (relationships with partners, vendors) | 15 | 39 | .649 |
| Extraorganizational (market, competitors, environment) | 16 | 23 | .612 |
Read the third row again. The category product teams care most about - what actually happens inside the company - is the category informants are least reliable about. Nobody in a discovery call is asking the VP for the company's revenue growth. They are asking exactly the intraorganizational questions where a second informant would only have agreed with the first at r = .50.
The second study puts a number on the same problem from the other direction: how much of the variance in an answer is explained not by the company but by the type of person you happened to ask.
| What the question is about | Share of variance due to informant type (construct level) |
|---|---|
| Overall | .264 |
| Performance | .123 |
| Intraorganizational | .337 |
| Interorganizational | .332 |
| Extraorganizational | .173 |
For internal-process questions, 33.7% of the variance in the construct is attributable to which kind of informant you selected. That is not a rounding error you can average away with a bigger sample of accounts. If you always interview the champion, you have not sampled companies - you have sampled champions, and a third of your signal is the shape of that role.
To their credit, the authors checked whether this was publication bias. It was not: 24 unpublished null studies would be needed to drag the mean correlation from .612 down to .500, and after a trim-and-fill correction the mean agreement moved only from .876 to .841.
What actually predicts a good informant (and what does not)
This is where the paper stops being interesting and starts being useful. The same meta-analysis tested thirteen candidate drivers of informant reliability. Several results contradict how research teams actually pick people.
| Factor | Does it improve reliability? | Detail |
|---|---|---|
| Higher hierarchical position | Yes - the strongest informant-level effect | b = .319, p < .01 |
| Specialist vs generalist function | No effect at all | b = -.273, not significant |
| Longer tenure in the unit | Partly - improves agreement, not correlation | b = .021 on agreement, p < .05 |
| Question is about the present, not the past | Yes | b = .094, p < .01 |
| Question is objective, not a subjective judgement | Yes | b = .072, p < .01 |
| Question is about a salient event, not a routine one | Yes - the largest construct effect | b = .129, p < .01 |
| Question is about a specific person vs an entity | No effect | not significant |
| Larger organization | Reduces reliability | b = -.131, p < .05 |
| Older organization | No effect | not significant |
Three of these are worth stopping on.
Seniority helps; specialism does not. The instinct in product research is to skip the executive and go to the practitioner "closest to the work." On reliability grounds that instinct is wrong: hierarchical position was the strongest informant characteristic tested, and functional specialization had no measurable effect at all. Senior people are continuously and formally informed about what the organization is doing; the specialist knows their own slice extremely well and infers the rest.
Bigger company, worse informant. Reliability fell as organization size rose. Asking one person to characterize a 400-person business unit is a harder measurement problem than asking one person to characterize a 12-person team, and the data show it. Your enterprise accounts are exactly where a single informant is least trustworthy - and exactly where teams most often settle for one, because access is hard.
Salience beats everything else about the question. The largest construct-level effect was whether the question referred to a conspicuous event rather than a routine process. People recall landmark events; they reconstruct routines. This is the single most actionable finding in the paper, because it is under your control: you cannot change your informant's seniority, but you can change "how does your team usually handle renewals?" into "walk me through what happened at your last renewal."
So the rule is not find a better informant. The rule is match the question to what this informant can actually be a reliable instrument for, and then use different informants for different questions.
Choosing an informant: a competence screen, not an availability screen
Most teams select informants by who replies to the email. The methodological literature - Kumar, Stern and Anderson's 1993 paper in the Academy of Management Journal (36(6), 1633-1651) is the standard reference - says to select on competence: proximity to the phenomenon, involvement in it, and confidence in reporting on it.
Operationally, that means putting the screen inside the interview instead of assuming it:
- Proximity. "In the last 90 days, how often were you personally in the room when this was decided or done?" Never / once / monthly / weekly.
- Role in the phenomenon. "Were you a decider, a doer, an approver, or an observer for this?"
- Self-rated confidence, asked per topic, not once. "How confident are you describing how other teams use this, on a 1-5 scale?" A 2 is not a failure - it is the datum that tells you to route that topic to somebody else.
- Scope honesty. "Are you answering for your team, your department, or the whole company?" Most people will answer for a much smaller unit than the question implied, which is exactly what you need to know.
Informants who score low should not be discarded. They should be re-scoped: their answers are still good evidence about their own team, and only weak evidence about the organization. The failure mode is not interviewing the wrong person, it is recording a narrow-scope answer as a wide-scope fact.
Multiple informants: necessary, and not free
The standard prescription is more informants per organization, and it is the right prescription. But it comes with two conditions that teams routinely skip.
Condition one: disagreement has to mean something before you collect it. If two informants disagree, you need to have decided in advance whether that is measurement error to be averaged away or a real difference in perspective to be reported. Kumar, Stern and Anderson's core contribution was procedures for exactly this problem - obtaining and interpreting perceptual agreement among multiple informants rather than silently averaging them.
Condition two: informants from the same role are barely independent. In the Homburg meta-analysis, triangulating against a second informant with the same functional background produced higher apparent reliability estimates (b = .101, p < .05) - which is what you would expect if two people in the same role share the same blind spots, not if they were independently right. Three champions is not three informants. It is roughly one informant with better error bars around the champion's view.
A workable target for a serious account study: one senior informant, one practitioner, and one person from a function that does not report to either. Then measure agreement explicitly rather than merging their answers into a single narrative.
What this is not
Three neighbouring problems get confused with this one, and the distinction matters because the fixes are different:
- This is not interviewer bias. Interviewer bias is variance introduced by the person asking. Key informant error is introduced by the person answering, about a unit they are not. Removing moderator variance - which the AI interviewer house effect guide covers in depth - does nothing to fix it.
- This is not the buying-committee playbook. Buying committee interviews tells you which roles to include and how to get them scheduled. This guide is about the measurement error in what any one of them tells you, whether or not you got all seven.
- This is not method triangulation. Triangulation is about combining methods. Multiple informants is triangulation across sources within one method - and as the same-function result above shows, it can be done in a way that produces false confidence.
The modern approach: how Koji makes multi-informant research affordable
The reason single-informant research persists is not that anyone thinks it is rigorous. It is that three informants per account, scheduled across three calendars, is three times the coordination cost - so teams take the one who answers.
AI-moderated interviews remove the coordination cost, which is the actual binding constraint:
- Every informant gets the same instrument. A study run by several human moderators carries the moderator's variance on top of the informant's; a single AI interviewer with a pinned configuration means the only variance left is the one you are trying to measure. Three informants become genuinely comparable measurements.
- Asynchronous means parallel. Informants complete on their own time, so a three-informant account study takes as long as the slowest single participant rather than as long as the hardest calendar. Enterprise accounts - where the Homburg data says single informants are weakest - become the ones you can actually cover properly.
- Structured questions give you the agreement statistic for free. Koji supports six types:
open_ended,scale,single_choice,multiple_choice,rankingandyes_no. Ask the samescalequestion ("how confident are you describing how other teams use this?") and the samerankingquestion ("rank these five factors by influence on the decision") of every informant in an account, and inter-informant agreement drops out as a number instead of an impression. That is the difference between "our informants broadly agreed" and "our three informants ranked the top factor identically and the second factor differently." - Salience is enforceable in the guide. Because the AI probes adaptively, you can write the guide around specific dated events - "the most recent renewal", "the last time this broke" - and have it push back when an informant drifts into generalities. The largest single driver of reliability in the meta-analysis is a property of the question, and it is the one property software can hold constant across dozens of interviews.
- Competence screening runs automatically. Proximity and confidence screens can sit at the top of every interview as structured questions, so every transcript arrives tagged with how much weight it deserves - rather than that judgement being made from memory during synthesis.
Traditional survey tools like SurveyMonkey or Qualtrics can field the same items, but they cannot probe when an informant answers for a narrower unit than you asked about - which is the exact moment a key informant report goes wrong.
A five-step key informant protocol
- Name the unit. Write down, before fielding, whether each question is about the individual, the team, the department, or the company. Questions that cross units get split.
- Sort your questions by informant-reliability class. Objective, present-tense, event-anchored questions can survive a single informant. Subjective, past-tense, routine-process questions need more than one.
- Recruit at least one senior informant per account, and at least one person outside the champion's reporting line.
- Screen for competence inside the interview, per topic, using scale questions - not once at recruitment.
- Report agreement, not just themes. Every organization-level claim in your readout should carry how many informants supported it and whether any contradicted it.
Frequently asked questions
What is a key informant interview?
A key informant interview is one in which a person reports on a group, organization or system rather than on themselves - for example asking a VP how their department handles renewals. The respondent acts as an instrument pointed at a unit they are only part of, which introduces a specific kind of measurement error distinct from ordinary self-report bias.
How accurate are key informant reports?
Across a meta-analysis of 127 organizational studies (Homburg, Klarmann, Reimann and Schilke, Journal of Marketing Research, 2012), key informant reports correlated .612 with independent second sources on average. Accuracy varied sharply by topic: .764 for hard performance figures but only .502 for questions about internal organizational processes, where 33.7% of construct-level variance was attributable to which type of informant was asked.
How many informants do I need per company?
There is no universal number, but one is defensible only for objective, present-tense, event-anchored questions. For anything about internal process, priorities or perception, plan for at least two informants from different levels and different reporting lines, and treat two informants in the same role as close to a single source rather than two independent ones.
Should I interview the executive or the person closest to the work?
For reliability, seniority helped and functional specialization did not - hierarchical position was the strongest informant-level predictor in the meta-analysis, while specialist-versus-generalist had no significant effect. In practice, interview both and route questions accordingly: the executive for organization-level facts and decisions, the practitioner for their own workflow, where they are a self-reporter rather than an informant.
Is a single-informant study worthless?
No. It is a valid measurement with a known error rate, and for salient, recent, objective questions that error rate is acceptable. It becomes a problem when a narrow-scope answer is written up as a company-level fact without any note of who said it or how confident they were. The cheapest possible improvement is not more interviews - it is attributing every organization-level claim to the informant who made it.
How is this different from ordinary response bias?
Ordinary response biases like acquiescence or social desirability distort what a person says about themselves. Key informant error is about the mismatch between the respondent and the unit being measured: a perfectly honest, perfectly calibrated person can still be wrong about what their own company does, because they only ever saw part of it.
Related Resources
- Structured Questions in AI Interviews - the six question types, including the scale and ranking questions that turn informant agreement into a number
- Proxy Response Bias - what happens to the error when a respondent answers for another person rather than for an organization
- Nobody Made the Decision - why better informant selection still will not recover why an account bought
- Buying Committee Interviews - which stakeholder roles to include in a multi-stakeholder B2B study
- The AI Interviewer House Effect - removing the moderator's variance so the informant's is the only one left
- Proxy Research - what to do when you cannot reach real users at all
- Triangulation in Research - combining methods, and how it differs from combining sources
- B2B User Research - the wider playbook for enterprise research programs
Related Articles
The AI Interviewer House Effect: When One Interviewer Turns Variance Into Bias
An AI interviewer removes interviewer variance and converts what remains into bias. How to measure your house effect with an interviewer A/B.
Buying Committee Interviews: Multi-Stakeholder B2B Research Without the Scheduling Nightmare
Run research interviews with every member of a B2B buying committee — economic buyers, end users, IT, security, finance, legal — without coordinating six calendars. Use Koji's personalized AI interview links to capture role-specific perspectives at each stakeholder's convenience, then synthesize the full account view in one report.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Triangulation in Research: Combining Methods for Stronger, More Credible Insights (2026)
Triangulation is the practice of using multiple data sources, methods, researchers, or theories to validate a finding. Learn Denzin's four types, when to use each, and how AI-native research platforms make multi-method studies practical instead of aspirational.
5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)
A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.
The 5-Second Test: How to Measure First Impressions and Visual Hierarchy (2026 Guide)
A complete guide to the 5-second test — the lightweight UX research method that measures gut reactions, message clarity, and visual hierarchy. Learn how to design questions, recruit participants, analyze results, and combine 5-second tests with AI interviews.