Proxy Response Bias: Why the Question, Not the Person, Decides How Wrong a Proxy Is (2026)
Observability and interaction explain over 60% of the gap between self-reports and proxy reports. Proxy error is systematic and directional, which is why adding more proxies makes you more confident and no more correct.
The standard fix for a shaky single source is to add more sources. When the extra sources are proxies - people answering on behalf of someone else - that fix quietly stops working, because proxy error is not noise. It has a direction, and the direction is set by the question, not by the person.
In the largest survey study of the effect, researchers regressed the gap between self-reports and proxy reports on independent ratings of how observable each item was. Observability and the "interactional" nature of the item accounted for more than 60% of the variance in self/proxy differences (Todorov and Kirchner, American Journal of Public Health, 2000, 90(8), 1248-53). Proxies were not randomly worse. They were accurate about what could be seen and systematically wrong about what could not.
The practical consequence inverts the usual advice: adding proxies narrows your confidence interval around a number that is biased, which makes you more certain and no more right.
The sign inversion, stated plainly
This is the second article in a chain, and it reverses the first one's lever.
Key informant research says: one person reporting for a whole organization is one measurement with an unchecked calibration, so add informants and measure their agreement. Adding sources reduces variance. That is correct.
This article says: when the added source is reporting on another person's internal state, adding sources shifts the estimate as well as tightening it. Variance goes down; bias does not. Two proxies who both cannot see the thing you are asking about will agree with each other beautifully, and their agreement is evidence of nothing except that they are both inferring from the same visible cues.
Same lever, opposite sign. This is why "we talked to three people at the account" is not automatically stronger evidence than talking to one - it depends entirely on whether you asked them about things they could observe.
What "proxy" means here, and what it does not
Koji's proxy research guide covers the substitution problem: you cannot get to real users, so you rank the internal roles who can stand in for them - support agents, CS, solution engineers, SMEs - by evidence quality. That is a question about who to use.
This guide is about the layer underneath, and it points at a different defendant. Once you have a proxy - any proxy, high or low on that ladder - the size and direction of their error is governed mostly by which question you ask them, not by their job title. A support agent is an excellent instrument for "what error message did they hit" and a poor one for "how frustrated were they". Moving up the ladder does not fix the second question. Re-sorting your question list does.
The proxy situations product teams are actually in:
- The admin or IT contact answering for 400 seats they provision but do not use
- The manager answering for how their team feels about a workflow
- The CSM or AE answering for what the customer wants, in the customer's absence
- The champion answering for the CFO who overruled them
- A parent, teacher or carer answering for the person the product is actually for
- Synthetic users - the limit case, a proxy with no observation of the person at all
The mechanism: proxies infer, self-reporters retrieve
The reason proxy error is predictable rather than random was pinned down in a follow-up study (Todorov, Applied Cognitive Psychology, 2003, 17, 215-224). Its summary names the source in one line: "proxies' higher reliance on inferences."
A self-respondent asked whether they have difficulty learning something searches memory. A proxy asked the same question about someone else cannot search memory, because they were never inside it. So they infer - from what they have seen, from what usually goes with what, from what would be true of a person like this.
The study tested that directly. Independently collected conditional likelihood judgements - how likely is a person to have difficulty X given that they have difficulty Y - predicted the pattern of proxy reports but not the pattern of self-reports. A model of self/proxy differences estimated on one year of national survey data and tested against the next year produced a correlation between predicted and actual differences of 0.76, and between predicted and actual proxy reports of 0.95.
That is the single most important fact in this article, and it cuts both ways:
- Proxy error is systematic, so it does not average out no matter how many proxies you add.
- Proxy error is modellable, so it can be anticipated, bounded, and in some cases corrected - which is more than you can say for most biases in qualitative research.
The size and direction of the error
Three findings give you the shape of it.
Direction depends on who the proxy is talking about. In the National Health Interview Survey on Disability, proxies under-reported conditions for people aged 18 to 64 and over-reported them for people aged 65 or older (Todorov and Kirchner, 2000). The authors' conclusion was blunt: use of proxies in representative surveys "introduces systematic biases, affecting national disability estimates."
Subjective domains are where proxies fall apart. A systematic review and meta-analysis of nine studies covering 1,980 person-proxy dyads across 10 countries (Crocker, Smith and Skevington, Journal of Clinical Epidemiology, 2015, 68(5), 584-95) found person-proxy correlations ranging from 0.28 for social quality of life to 0.44 for physical quality of life - and proxies significantly underestimated across the board: social (mean difference 4.7, 95% CI 1.8-7.6), psychological (3.7, CI 0.6-6.8) and physical (3.1, CI 0.6-5.6). The conclusion: "Proxies tend to be imprecise, underestimating."
Observability explains most of the gap. This is the Todorov and Kirchner >60%-of-variance result, and it is the one that turns into a working rule.
| Kind of question | Observable to a proxy? | Expected proxy error |
|---|---|---|
| Which feature did they open, how often, on what plan | Directly observable | Low - proxies are often better than self-report here |
| What error or blocker did they hit | Observable if it was escalated | Low to moderate, skewed toward escalated cases |
| What workaround did they build | Observable if visible in shared tooling | Moderate |
| How difficult did they find it | Inferred from behaviour | High, and directional |
| How frustrated, anxious or confident did they feel | Not observable | Very high - the 0.28 end of the range |
| What they would have done instead | Not observable, and not reliably known even to the person | Do not ask a proxy at all |
The honest counterweight
It would be too convenient to conclude that proxies are simply bad. Two things stop that.
First, Robert Moore's review of the literature (Journal of Official Statistics, 1988, 4(2), 155-172) concluded that there was not substantial evidence that data quality is altered by reporting status - the differences are real but far more topic-dependent than a blanket "proxies are worse" implies. On some factual items proxies do fine, and occasionally better, because they are less motivated to manage the impression.
Second, the comparison is structurally confounded, and honest researchers say so. As Todorov (2003) puts it, the main problem in studying self/proxy differences is that respondents are not randomly allocated to self- or proxy-status. In household surveys, the people who need a proxy differ from the people who do not - they are more likely to be absent, busier, sicker, or less able to respond. Some of what looks like proxy bias is selection.
Both caveats point the same way: the useful question is never "should I use a proxy?" It is "for which items is this specific proxy a defensible instrument?"
The observability triage: a two-question test
Before any question goes to a proxy, run it through two checks:
- Could this person have observed the answer? Not inferred it - observed it. If the answer lives only inside someone's head, a proxy is guessing, and their guess will be systematically shaded by what they can see.
- Is the thing being asked about interactional? Todorov and Kirchner's second predictor: items involving interaction with others are more visible to a proxy than items that are purely internal. "Does she ask colleagues for help with this?" is far more answerable by a proxy than "does she find this confusing?"
Questions that fail both checks should not simply be dropped. Rewrite them into their observable counterpart:
| Do not ask a proxy | Ask instead |
|---|---|
| "How frustrated are your users with the import flow?" | "In the last month, how many times did someone come to you about the import flow?" |
| "Do your team members find the new dashboard useful?" | "Which of these five reports have you personally seen someone open?" |
| "Why did your team stop using it?" | "What changed in your team's process in the month it stopped being used?" |
| "Would your users pay for this?" | Do not ask. Route to the users. |
The rewrite is not a workaround; it is the correct measurement. You are asking the proxy for what they are an instrument for, and refusing to accept an inference dressed as an observation.
Bounding the error when you cannot avoid it
Sometimes the person is genuinely unreachable and the subjective question genuinely matters. Then the move is to bound the answer rather than believe it:
- Ask the proxy for their confidence, per item. A
scalequestion - "how confident are you answering this on their behalf?" - converts an invisible inference into a visible one, and low-confidence items get flagged rather than averaged in. - Ask what they are basing it on. "What did you see that makes you say that?" separates observation from inference inside the transcript itself.
- Ask for the direction, not the level. Proxies are more reliable on change ("is this getting better or worse?") than on absolute states, because change is more often marked by visible events.
- Collect one matched pair. If you can reach even five real users alongside twenty proxies, you can measure the gap in your own setting and apply it as a correction - which is exactly the strategy the survey-methodology literature validated when it predicted proxy reports at r = 0.95.
The modern approach: how Koji handles proxy questions
Proxy bias survives in practice because splitting a discussion guide by observability, running two different instruments, and comparing them is more work than most teams can justify. AI moderation changes that arithmetic.
- Route by observability, not by role. Build one study where observable items go to the admin, the CSM or the manager, and the subjective items go to the end users - then field both at once. Because Koji interviews are asynchronous and self-serve, running two instruments costs roughly what running one costs.
- Measure the gap instead of assuming it. Ask the same
scaleitem of both the proxy and the person - "how difficult was this, 1-5" - and the proxy gap becomes a number in your report rather than a caveat in your appendix. - Confidence screens on every proxy item. Koji's structured questions cover six types -
open_ended,scale,single_choice,multiple_choice,rankingandyes_no- so a per-item confidencescaleand an observation-sourceopen_endedfollow-up can sit behind every question a proxy answers, automatically, without doubling the guide's length for the humans reading it. - Adaptive probing catches inference in the act. When a proxy answers a subjective question, the AI can ask "what did you see that told you that?" every single time. A human moderator does this when they remember to; an AI interviewer does it on turn after turn, across a hundred interviews, identically.
- Reaching the real person gets cheap. The deepest fix for proxy bias is not to need a proxy. Async AI interviews remove the scheduling cost that pushed teams toward the admin in the first place - which is why proxy use should shrink as a share of your evidence base rather than being optimised forever.
Legacy survey platforms will happily field a proxy questionnaire, but they cannot ask the follow-up that separates what this person saw from what this person assumes - and that distinction is the entire difference between a usable proxy answer and a confident fabrication.
Frequently asked questions
What is proxy response bias?
Proxy response bias is the systematic difference between what a person reports about themselves and what someone else reports on their behalf. It is not random error: research on national survey data found that how observable an item is, plus whether it involves interaction with others, explained more than 60% of the variance in self/proxy differences.
Are proxy respondents always less accurate than the person themselves?
No. Robert Moore's 1988 review in the Journal of Official Statistics found no substantial evidence that data quality is uniformly altered by reporting status. Proxies can be as good or better on observable, factual items, and occasionally less prone to impression management. They degrade sharply on subjective and internal states, where person-proxy correlations have been measured as low as 0.28.
Does interviewing more proxies fix the problem?
No, and this is the key difference from ordinary sampling error. Proxy error is systematic and directional, so additional proxies tighten the confidence interval around a biased estimate. Two proxies inferring from the same visible cues will agree with each other while both being wrong in the same direction.
Which questions are safe to ask a proxy?
Questions about things the proxy could have directly observed: what happened, how often, who was involved, what changed, what was escalated. Questions about internal states - difficulty, frustration, confidence, intent, willingness to pay - should be rewritten into their observable counterparts or routed to the person themselves.
How do I correct for proxy bias if I have no choice?
Bound it rather than believe it. Ask for per-item confidence with a scale question, ask what observation the answer is based on, prefer direction-of-change questions over absolute levels, and if you can reach even a small matched sample of the real people, measure the gap in your own setting and apply it. Modelled corrections of exactly this kind have predicted actual proxy reports at r = 0.95.
Is this the same as using proxies because I cannot reach users?
They are related but different problems. Choosing which stand-in to use when access is blocked is covered in proxy research. This guide covers what happens to the accuracy of any stand-in's answers once you have one - and the answer is that it depends far more on the question than on the person.
Related Resources
- Structured Questions in AI Interviews - the six question types, including the per-item confidence scales that make proxy uncertainty visible
- Key Informant Interviews - the same problem one level up, where a person reports for an organization rather than for another person
- Nobody Made the Decision - what remains unrecoverable after both proxy and informant error are fixed
- Proxy Research - how to rank and use stand-in sources when you cannot reach users at all
- Synthetic Users in Research - the limit case of a proxy with no observation of the person
- Recall Bias - how memory distorts self-report, the error a proxy is not even subject to
- Nonresponse Bias - why the people who need a proxy differ from the people who do not
- Survivorship Bias in Customer Research - the other way a sample quietly stops representing the population
Related Articles
Proxy Research: How to Run User Research When You Can't Talk to Your Users
Blocked from real users by legal, sales, or a locked-down enterprise account? A disciplined framework for using proxies — ranked by evidence quality, with the bias corrections each one requires.
Recall Bias: How Faulty Memory Distorts Research (and How to Prevent It)
Recall bias is the systematic error that arises when respondents remember past events inaccurately or incompletely. Learn why memory is reconstructed not retrieved, how telescoping distorts data, and how to design around it.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Synthetic Users in Research: Validity, Bias, and When AI Personas Are (and Aren't) Trustworthy
A research methodology guide to synthetic users — what they are, the documented bias problems (sycophancy, sign-flipping, shallow insights), the legitimate use cases, and why real AI-moderated interviews are now fast enough that the synthetic-vs-real tradeoff has fundamentally shifted.
5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)
A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.
The 5-Second Test: How to Measure First Impressions and Visual Hierarchy (2026 Guide)
A complete guide to the 5-second test — the lightweight UX research method that measures gut reactions, message clarity, and visual hierarchy. Learn how to design questions, recruit participants, analyze results, and combine 5-second tests with AI interviews.