{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-19T11:33:40.967Z"},"content":[{"type":"documentation","id":"aed47dfd-ae7d-4f51-ab39-3dd913d7111f","slug":"proxy-response-bias-observability","title":"Proxy Response Bias: Why the Question, Not the Person, Decides How Wrong a Proxy Is (2026)","url":"https://www.koji.so/docs/proxy-response-bias-observability","summary":"Proxy response bias is the systematic difference between what people report about themselves and what someone reports on their behalf. It is not random: observability and the interactional nature of an item explained over 60% of the variance in self/proxy differences in national survey data, and the mechanism is that proxies infer rather than retrieve. Because the error is directional, adding proxies reduces variance without reducing bias. The fix is to triage every question by observability and rewrite subjective items into their observable counterparts.","content":"The standard fix for a shaky single source is to add more sources. When the extra sources are **proxies** - people answering on behalf of someone else - that fix quietly stops working, because proxy error is not noise. It has a direction, and the direction is set by the *question*, not by the person.\n\nIn the largest survey study of the effect, researchers regressed the gap between self-reports and proxy reports on independent ratings of how **observable** each item was. Observability and the \"interactional\" nature of the item **accounted for more than 60% of the variance in self/proxy differences** (Todorov and Kirchner, *American Journal of Public Health*, 2000, 90(8), 1248-53). Proxies were not randomly worse. They were accurate about what could be seen and systematically wrong about what could not.\n\nThe practical consequence inverts the usual advice: **adding proxies narrows your confidence interval around a number that is biased, which makes you more certain and no more right.**\n\n## The sign inversion, stated plainly\n\nThis is the second article in a chain, and it reverses the first one's lever.\n\n[Key informant research](/docs/key-informant-interviews-research) says: one person reporting for a whole organization is one measurement with an unchecked calibration, so add informants and measure their agreement. Adding sources reduces variance. That is correct.\n\nThis article says: when the added source is reporting on *another person's* internal state, adding sources **shifts the estimate** as well as tightening it. Variance goes down; bias does not. Two proxies who both cannot see the thing you are asking about will agree with each other beautifully, and their agreement is evidence of nothing except that they are both inferring from the same visible cues.\n\nSame lever, opposite sign. This is why \"we talked to three people at the account\" is not automatically stronger evidence than talking to one - it depends entirely on whether you asked them about things they could observe.\n\n## What \"proxy\" means here, and what it does not\n\nKoji's [proxy research guide](/docs/proxy-research-guide) covers the substitution problem: you cannot get to real users, so you rank the internal roles who can stand in for them - support agents, CS, solution engineers, SMEs - by evidence quality. That is a question about **who** to use.\n\nThis guide is about the layer underneath, and it points at a different defendant. Once you have a proxy - any proxy, high or low on that ladder - the size and direction of their error is governed mostly by **which question you ask them**, not by their job title. A support agent is an excellent instrument for \"what error message did they hit\" and a poor one for \"how frustrated were they\". Moving up the ladder does not fix the second question. Re-sorting your question list does.\n\nThe proxy situations product teams are actually in:\n\n- The **admin or IT contact** answering for 400 seats they provision but do not use\n- The **manager** answering for how their team feels about a workflow\n- The **CSM or AE** answering for what the customer wants, in the customer's absence\n- The **champion** answering for the CFO who overruled them\n- A **parent, teacher or carer** answering for the person the product is actually for\n- [Synthetic users](/docs/synthetic-users-research-methodology) - the limit case, a proxy with no observation of the person at all\n\n## The mechanism: proxies infer, self-reporters retrieve\n\nThe reason proxy error is predictable rather than random was pinned down in a follow-up study (Todorov, *Applied Cognitive Psychology*, 2003, 17, 215-224). Its summary names the source in one line: **\"proxies' higher reliance on inferences.\"**\n\nA self-respondent asked whether they have difficulty learning something searches memory. A proxy asked the same question about someone else cannot search memory, because they were never inside it. So they infer - from what they have seen, from what usually goes with what, from what would be true of a person like this.\n\nThe study tested that directly. Independently collected conditional likelihood judgements - how likely is a person to have difficulty X *given* that they have difficulty Y - predicted the pattern of proxy reports but **not** the pattern of self-reports. A model of self/proxy differences estimated on one year of national survey data and tested against the next year produced a correlation between predicted and actual differences of **0.76**, and between predicted and actual proxy reports of **0.95**.\n\nThat is the single most important fact in this article, and it cuts both ways:\n\n- Proxy error is **systematic**, so it does not average out no matter how many proxies you add.\n- Proxy error is **modellable**, so it can be anticipated, bounded, and in some cases corrected - which is more than you can say for most biases in qualitative research.\n\n## The size and direction of the error\n\nThree findings give you the shape of it.\n\n**Direction depends on who the proxy is talking about.** In the National Health Interview Survey on Disability, proxies **under-reported** conditions for people aged 18 to 64 and **over-reported** them for people aged 65 or older (Todorov and Kirchner, 2000). The authors' conclusion was blunt: use of proxies in representative surveys \"introduces systematic biases, affecting national disability estimates.\"\n\n**Subjective domains are where proxies fall apart.** A systematic review and meta-analysis of nine studies covering **1,980 person-proxy dyads across 10 countries** (Crocker, Smith and Skevington, *Journal of Clinical Epidemiology*, 2015, 68(5), 584-95) found person-proxy correlations ranging from **0.28 for social quality of life to 0.44 for physical quality of life** - and proxies significantly underestimated across the board: social (mean difference 4.7, 95% CI 1.8-7.6), psychological (3.7, CI 0.6-6.8) and physical (3.1, CI 0.6-5.6). The conclusion: \"Proxies tend to be imprecise, underestimating.\"\n\n**Observability explains most of the gap.** This is the Todorov and Kirchner >60%-of-variance result, and it is the one that turns into a working rule.\n\n| Kind of question | Observable to a proxy? | Expected proxy error |\n| --- | --- | --- |\n| Which feature did they open, how often, on what plan | Directly observable | Low - proxies are often *better* than self-report here |\n| What error or blocker did they hit | Observable if it was escalated | Low to moderate, skewed toward escalated cases |\n| What workaround did they build | Observable if visible in shared tooling | Moderate |\n| How difficult did they find it | Inferred from behaviour | High, and directional |\n| How frustrated, anxious or confident did they feel | Not observable | Very high - the 0.28 end of the range |\n| What they would have done instead | Not observable, and not reliably known even to the person | Do not ask a proxy at all |\n\n## The honest counterweight\n\nIt would be too convenient to conclude that proxies are simply bad. Two things stop that.\n\nFirst, Robert Moore's review of the literature (*Journal of Official Statistics*, 1988, 4(2), 155-172) concluded that there was **not substantial evidence that data quality is altered by reporting status** - the differences are real but far more topic-dependent than a blanket \"proxies are worse\" implies. On some factual items proxies do fine, and occasionally better, because they are less motivated to manage the impression.\n\nSecond, the comparison is structurally confounded, and honest researchers say so. As Todorov (2003) puts it, the main problem in studying self/proxy differences is that **respondents are not randomly allocated to self- or proxy-status**. In household surveys, the people who need a proxy differ from the people who do not - they are more likely to be absent, busier, sicker, or less able to respond. Some of what looks like proxy bias is selection.\n\nBoth caveats point the same way: the useful question is never \"should I use a proxy?\" It is \"**for which items is this specific proxy a defensible instrument?**\"\n\n## The observability triage: a two-question test\n\nBefore any question goes to a proxy, run it through two checks:\n\n1. **Could this person have observed the answer?** Not inferred it - observed it. If the answer lives only inside someone's head, a proxy is guessing, and their guess will be systematically shaded by what they can see.\n2. **Is the thing being asked about interactional?** Todorov and Kirchner's second predictor: items involving interaction with others are more visible to a proxy than items that are purely internal. \"Does she ask colleagues for help with this?\" is far more answerable by a proxy than \"does she find this confusing?\"\n\nQuestions that fail both checks should not simply be dropped. Rewrite them into their observable counterpart:\n\n| Do not ask a proxy | Ask instead |\n| --- | --- |\n| \"How frustrated are your users with the import flow?\" | \"In the last month, how many times did someone come to you about the import flow?\" |\n| \"Do your team members find the new dashboard useful?\" | \"Which of these five reports have you personally seen someone open?\" |\n| \"Why did your team stop using it?\" | \"What changed in your team's process in the month it stopped being used?\" |\n| \"Would your users pay for this?\" | Do not ask. Route to the users. |\n\nThe rewrite is not a workaround; it is the correct measurement. You are asking the proxy for what they are an instrument for, and refusing to accept an inference dressed as an observation.\n\n## Bounding the error when you cannot avoid it\n\nSometimes the person is genuinely unreachable and the subjective question genuinely matters. Then the move is to **bound** the answer rather than believe it:\n\n- **Ask the proxy for their confidence, per item.** A `scale` question - \"how confident are you answering this on their behalf?\" - converts an invisible inference into a visible one, and low-confidence items get flagged rather than averaged in.\n- **Ask what they are basing it on.** \"What did you see that makes you say that?\" separates observation from inference inside the transcript itself.\n- **Ask for the direction, not the level.** Proxies are more reliable on change (\"is this getting better or worse?\") than on absolute states, because change is more often marked by visible events.\n- **Collect one matched pair.** If you can reach even five real users alongside twenty proxies, you can measure the gap in your own setting and apply it as a correction - which is exactly the strategy the survey-methodology literature validated when it predicted proxy reports at r = 0.95.\n\n## The modern approach: how Koji handles proxy questions\n\nProxy bias survives in practice because splitting a discussion guide by observability, running two different instruments, and comparing them is more work than most teams can justify. AI moderation changes that arithmetic.\n\n- **Route by observability, not by role.** Build one study where observable items go to the admin, the CSM or the manager, and the subjective items go to the end users - then field both at once. Because Koji interviews are asynchronous and self-serve, running two instruments costs roughly what running one costs.\n- **Measure the gap instead of assuming it.** Ask the same `scale` item of both the proxy and the person - \"how difficult was this, 1-5\" - and the proxy gap becomes a number in your report rather than a caveat in your appendix.\n- **Confidence screens on every proxy item.** Koji's [structured questions](/docs/structured-questions-guide) cover six types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` - so a per-item confidence `scale` and an observation-source `open_ended` follow-up can sit behind every question a proxy answers, automatically, without doubling the guide's length for the humans reading it.\n- **Adaptive probing catches inference in the act.** When a proxy answers a subjective question, the AI can ask \"what did you see that told you that?\" every single time. A human moderator does this when they remember to; an AI interviewer does it on turn after turn, across a hundred interviews, identically.\n- **Reaching the real person gets cheap.** The deepest fix for proxy bias is not to need a proxy. Async AI interviews remove the scheduling cost that pushed teams toward the admin in the first place - which is why proxy use should shrink as a share of your evidence base rather than being optimised forever.\n\nLegacy survey platforms will happily field a proxy questionnaire, but they cannot ask the follow-up that separates *what this person saw* from *what this person assumes* - and that distinction is the entire difference between a usable proxy answer and a confident fabrication.\n\n## Frequently asked questions\n\n### What is proxy response bias?\n\nProxy response bias is the systematic difference between what a person reports about themselves and what someone else reports on their behalf. It is not random error: research on national survey data found that how observable an item is, plus whether it involves interaction with others, explained more than 60% of the variance in self/proxy differences.\n\n### Are proxy respondents always less accurate than the person themselves?\n\nNo. Robert Moore's 1988 review in the *Journal of Official Statistics* found no substantial evidence that data quality is uniformly altered by reporting status. Proxies can be as good or better on observable, factual items, and occasionally less prone to impression management. They degrade sharply on subjective and internal states, where person-proxy correlations have been measured as low as 0.28.\n\n### Does interviewing more proxies fix the problem?\n\nNo, and this is the key difference from ordinary sampling error. Proxy error is systematic and directional, so additional proxies tighten the confidence interval around a biased estimate. Two proxies inferring from the same visible cues will agree with each other while both being wrong in the same direction.\n\n### Which questions are safe to ask a proxy?\n\nQuestions about things the proxy could have directly observed: what happened, how often, who was involved, what changed, what was escalated. Questions about internal states - difficulty, frustration, confidence, intent, willingness to pay - should be rewritten into their observable counterparts or routed to the person themselves.\n\n### How do I correct for proxy bias if I have no choice?\n\nBound it rather than believe it. Ask for per-item confidence with a scale question, ask what observation the answer is based on, prefer direction-of-change questions over absolute levels, and if you can reach even a small matched sample of the real people, measure the gap in your own setting and apply it. Modelled corrections of exactly this kind have predicted actual proxy reports at r = 0.95.\n\n### Is this the same as using proxies because I cannot reach users?\n\nThey are related but different problems. Choosing which stand-in to use when access is blocked is covered in [proxy research](/docs/proxy-research-guide). This guide covers what happens to the accuracy of any stand-in's answers once you have one - and the answer is that it depends far more on the question than on the person.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types, including the per-item confidence scales that make proxy uncertainty visible\n- [Key Informant Interviews](/docs/key-informant-interviews-research) - the same problem one level up, where a person reports for an organization rather than for another person\n- [Nobody Made the Decision](/docs/group-decision-research-reconstruction) - what remains unrecoverable after both proxy and informant error are fixed\n- [Proxy Research](/docs/proxy-research-guide) - how to rank and use stand-in sources when you cannot reach users at all\n- [Synthetic Users in Research](/docs/synthetic-users-research-methodology) - the limit case of a proxy with no observation of the person\n- [Recall Bias](/docs/recall-bias) - how memory distorts self-report, the error a proxy is not even subject to\n- [Nonresponse Bias](/docs/nonresponse-bias) - why the people who need a proxy differ from the people who do not\n- [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) - the other way a sample quietly stops representing the population\n","category":"Research Methods","lastModified":"2026-08-19T03:26:13.365099+00:00","metaTitle":"Proxy Response Bias: The Question Decides the Error (2026)","metaDescription":"Observability explains over 60% of the self/proxy gap. Proxy error is directional, so more proxies tighten the interval around a biased number. Triage your questions instead.","keywords":["proxy response bias","proxy respondent","proxy reporting survey","self vs proxy report","observability bias","answering on behalf of someone","research accuracy"],"aiSummary":"Proxy response bias is the systematic difference between what people report about themselves and what someone reports on their behalf. It is not random: observability and the interactional nature of an item explained over 60% of the variance in self/proxy differences in national survey data, and the mechanism is that proxies infer rather than retrieve. Because the error is directional, adding proxies reduces variance without reducing bias. The fix is to triage every question by observability and rewrite subjective items into their observable counterparts.","aiPrerequisites":["Familiarity with interview or survey question design","Some exposure to B2B or multi-stakeholder research"],"aiLearningOutcomes":["Distinguish proxy error from ordinary sampling error","Predict the direction of proxy bias from the question, not the role","Triage a discussion guide by observability","Rewrite subjective proxy questions into observable ones","Bound proxy error when the real person is unreachable"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}