Research Independence: Why the Team That Built the Feature Should Not Grade It
The five threats to independence from professional ethics codes, applied to product research - and why structural independence is a different problem from cognitive bias.
Answer first: professional ethics codes name five categories of threat to independence - self-interest, self-review, advocacy, familiarity, and intimidation - and product research is exposed to all five, with self-review as the structural default. When the person who wrote the brief also picked the participants, moderated the sessions, analysed the transcripts, and presented the recommendation, they have reviewed their own work at four consecutive checkpoints. No amount of self-awareness fixes that, because it is not a flaw in the person. It is a flaw in the org chart, and the remedy is separation of duties rather than better intentions.
This is the part of research quality that bias training cannot reach.
Five threats, named
The International Code of Ethics for Professional Accountants identifies five categories of circumstance that can compromise objectivity, and requires practitioners to apply safeguards when a threat is not at an acceptable level. The taxonomy transfers to research with almost no adaptation.
| Threat | In auditing | In product research |
|---|---|---|
| Self-interest | A financial interest in the client | The researcher's team is measured on the adoption of the feature being tested |
| Self-review | Evaluating work your own firm produced | The person who designed the flow runs the usability study on it |
| Advocacy | Promoting a client's position | Research commissioned to support a decision already taken |
| Familiarity | A long, close relationship with the client | Interviewing the same friendly power users every quarter |
| Intimidation | Pressure from a dominant client | A senior stakeholder who reacts badly to negative findings |
Most research quality programmes address none of these, because they are not method problems. A perfectly designed study run by a person facing an unmanaged intimidation threat produces a report with the inconvenient section softened, and every method checklist it passes will still say pass.
Self-review is the one product teams have structurally
The self-review threat deserves separate treatment because it is not an occasional risk in product research - it is the normal operating model. The same person or the same small team typically:
- Frames the question, which decides what could possibly be found
- Defines the screener, which decides who is in a position to complain
- Moderates, which decides which answers get probed and which are accepted
- Analyses, which decides what becomes a theme and what stays an anecdote
- Presents, which decides what the decision-maker hears
Each of those is a judgement call about work made at the previous step. By stage five the output has been reviewed exclusively by the party that produced it, five times. In an audit context that arrangement is not merely discouraged, it is prohibited outright.
What the law did about it
After Enron, the Sarbanes-Oxley Act made auditor independence a statutory matter rather than a professional aspiration. Section 201 lists categories of non-audit service an auditor may not provide to an audit client - bookkeeping and services related to the client's accounting records, appraisal and valuation services, and internal audit outsourcing among them.
More useful than the list is the general test behind it. Services are prohibited where they would: (1) result in the auditor auditing its own work, (2) create a mutual or conflicting interest between auditor and client, (3) result in the auditor performing management functions, or (4) place the auditor in the position of advocating for the client.
Hold a typical embedded product research function against those four. It frequently evaluates work its own function specified (1). It shares a success metric with the team it is evaluating (2). It often makes the prioritisation call rather than informing it (3). And it is regularly asked to build the case for a decision rather than test it (4). All four, in one role, as standard practice. The point is not that this is scandalous - it is that a whole profession decided each of these four was individually disqualifying, and research does all four at once without noticing.
Independence is not the same as being unbiased
This distinction is the crux, and it is where most internal debates go wrong.
The corpus already covers the cognitive layer thoroughly. Observer bias covers how a researcher's expectations skew what they see, confirmation bias in user research covers the search for supporting evidence, interviewer bias covers how moderators distort sessions, and avoiding bias in research interviews covers the technique-level remedies. Those are all properties of a mind, and their remedies are awareness, training, and protocol.
Independence is a property of a position. An auditor with flawless self-knowledge and no conscious bias is still not independent if they own shares in the client, and the profession disqualifies them anyway - not because they are presumed dishonest, but because the appearance and structure of the arrangement make the judgement unverifiable. Ethics codes handle this by distinguishing independence of mind from independence in appearance, and requiring both.
The practical consequence for research teams: you cannot train your way out of a self-review threat, and you cannot resolve it by having the researcher declare their assumptions. You resolve it by moving a task to a different person, or by removing the researcher's stake in the outcome.
Segregation of duties for a research study
Segregation of duties is the internal-control answer to the same problem: no single person should control every stage of a transaction. Applied to a study, the useful separations are fewer than teams fear.
| Stage | Must not also be done by | Cheapest safeguard |
|---|---|---|
| Framing the question | The person who owns the feature's success metric | Brief reviewed by someone with no stake in the answer |
| Screening and recruiting | The person who will present the results | Screener criteria pre-registered before recruiting opens |
| Moderating | Anyone who has stated a public prediction about the outcome | A moderator with no position on the question |
| Analysing | The moderator, alone | An independent pass over transcripts, compared afterwards |
| Deciding | The researcher | Research informs; the decision-maker decides and owns it |
The last row is the one research teams resist and it is the most important. A research function that makes product calls has traded independence for influence, and it can no longer credibly evaluate the calls it made. Auditors are barred from performing management functions for exactly this reason.
Where a full separation is unaffordable, the fallback is the pre-registration of judgement calls, so that at least the decisions are made before the results are visible: write the screener, the primary question, the analysis rule, and the threshold for "this is a real finding" into the brief. The pre-launch QA gate is the natural place to check that this happened, and changing a study mid-field covers what to do when a genuinely necessary change arrives after fielding.
Independence is hard even for people whose entire job is independence
It would be easy to read all of this as a counsel of perfection that a real product team can safely ignore. The evidence says otherwise, and it says it in a direction that should be sobering rather than reassuring.
The Public Company Accounting Oversight Board inspects registered audit firms and publishes the share of inspected audits with a Part I.A deficiency - cases where the auditor had not obtained sufficient appropriate evidence to support the opinion it had already signed. In the report published on 31 March 2025, the aggregate rate across all inspected firms was 39% for the 2024 inspection cycle, down from 46% in 2023. Among the Big Four - who audit roughly 80% of the market capitalisation of US-listed public companies - it was 20%, down from 26%. Across the six US Global Network Firms, 26%, down from 34%.
Read those numbers carefully. This is a profession with statutory independence requirements, mandatory rotation, prohibited-service lists, an external inspector, and personal liability - and roughly one in five of the best-resourced audits still lacked sufficient evidence for the opinion that was signed. The improvement trend is real and worth noting. But if that is the floor with all of those controls in place, the honest expectation for an embedded research function with none of them is not that it does better.
The practical inference is not despair. It is that structural safeguards are the thing that moved those numbers, and that self-assessment was never going to.
Where AI moderation actually helps, and where it does not
There is a genuine structural argument here, and it is worth stating precisely rather than overselling.
The moderator has no stake in the answer. An AI interviewer does not have a promotion riding on the feature performing well, has not spent three months arguing for the design, and does not flinch when a participant is scathing. That removes the self-interest and intimidation threats from the moderation stage specifically - the stage where they do the most damage, because a probe not asked leaves no trace in the transcript for a reviewer to catch later.
The instrument is identical for everyone. A human moderator's questions drift across a fieldwork period as their hypothesis firms up. Koji asks the same structured questions of every participant - across the six types: open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - with follow-up probing that adapts to what the participant said rather than to what the researcher hopes to hear. See the structured questions guide.
Independent analysis stops being a budget item. The safeguard in the analysis row of the table above traditionally required a second senior analyst. An automatic independent pass over the same transcripts makes it a default rather than a luxury, and the researcher's own themes can then be compared against it.
What it does not fix. The AI does not choose the question, and framing is the most consequential judgement in the whole chain. It does not choose the screener. And it does not decide what goes in the executive summary. Those three remain human, and they remain exactly where the self-review threat lives. Anyone claiming an AI moderator delivers independence end to end is selling something; it removes the threat from one stage of five, which is worth having and is not the same as solving it. On the limits of trusting automated moderation, see can you trust AI interviewers and synthetic users in research.
Honest objections
The domain expert is the best interviewer we have. Often true. Independence has a real cost in interview quality, and the trade is not free. The usual resolution is to let the expert shape the guide and stay out of the moderation and the analysis, which keeps most of the expertise and removes most of the threat.
Small teams cannot separate five roles across three people. They cannot, and they do not have to. Separate the two that matter most: whoever owns the success metric should not moderate, and whoever presents should not have been the only analyst.
This assumes researchers are dishonest. It assumes the opposite. The entire point of a structural safeguard is that it works on honest people, whose judgement is nonetheless unverifiable from the outside when they occupy both positions. That is why auditors with no conscious bias are still disqualified for owning shares.
Frequently asked questions
What are the five threats to independence?
Self-interest, self-review, advocacy, familiarity, and intimidation. They come from the International Code of Ethics for Professional Accountants, which requires practitioners to apply safeguards whenever a threat is not at an acceptable level. All five have direct analogues in product research, and self-review is the one most teams are exposed to by default.
How is independence different from avoiding bias?
Bias is a property of a mind and is addressed with training, awareness, and protocol. Independence is a property of a position and is addressed by changing who does what. An auditor with no conscious bias is still disqualified for holding a financial interest in the client, because from outside the arrangement the judgement cannot be verified.
Can the person who designed a feature run the usability study on it?
It is a textbook self-review threat, and the safeguard is cheap: let them shape the study, but have someone else moderate and have an independent pass at the analysis. If the team is too small for that, at minimum ensure the person who owns the feature's success metric is not the person deciding which findings make the summary.
Should the research team make the product decision?
No. Performing management functions is one of the four conditions that disqualifies an auditor, for the same reason it compromises research: a function that makes the call cannot credibly evaluate the call. Research informs the decision and the decision-maker owns it.
Does using an AI moderator make a study independent?
It removes the self-interest and intimidation threats from the moderation stage, and it makes independent analysis affordable. It does not choose the research question, define the screener, or write the executive summary, and those remain the stages where self-review does the most damage. It is one safeguard among several, not a substitute for the rest.
What is the minimum viable safeguard for a small team?
Two separations. First, whoever owns the metric the study could embarrass does not moderate. Second, whoever presents the findings is not the only person who analysed them. Add pre-registration of the screener, the primary question, and the finding threshold, and a two-person team has most of the available benefit.
Related Resources
- Structured Questions Guide - the six question types that hold the instrument constant across participants
- Levels of Assurance in Research - what a study is entitled to conclude before independence is even considered
- Corroboration in Research - independence of sources, the companion to independence of people
- Research Peer Review: The Pre-Launch QA Gate - where the separation of duties should be checked
- Observer Bias - the cognitive layer that independence safeguards cannot reach
- The Research Expectation Gap - why even independent, well-evidenced work still disappoints
Related Articles
Confirmation Bias in User Research: How to Recognize and Eliminate It
Confirmation bias quietly corrupts user research by leading teams to hear what they already believe. Learn how it shows up in interviews and analysis, and the practical tactics — and AI moderation — that neutralize it.
Corroboration in Research: Why Three Sources Saying the Same Thing Can Be One Source
Evidence from multiple sources only multiplies confidence when the sources are independent. Four ways research sources secretly share an origin, and a ten-minute test for catching it.
Interviewer Bias: How Moderators Distort Research (and How AI Removes the Variance)
Interviewer bias is the distortion caused by a moderator's wording, reactions, expectations, and characteristics. Learn the types, the evidence, mitigation techniques, and why an AI interviewer eliminates interviewer variance.
Observer Bias in Research: How the Researcher's Expectations Skew What They See
Observer bias is when a researcher's expectations unconsciously shape what they record and how they interpret it. Learn how it works, the evidence behind it, and how to design it out — including with a neutral AI moderator.
Levels of Assurance in Research: How Much Confidence a Study Can Honestly Support
Auditing defines three levels of assurance - reasonable, limited, and none. Research reports use one voice for all three. Here is how to pick and state the level before you field a study.
The Research Expectation Gap: Why Stakeholders Are Disappointed by Studies That Were Done Right
Auditing measured its own credibility gap and found only 16% was sub-standard work. Half was scope. Here is how that decomposition changes what research teams should fix.
Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies
Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.