{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-06T14:11:04.967Z"},"content":[{"type":"documentation","id":"6c8f7dfe-e019-4a10-a10c-9a89dd7102f1","slug":"entrustable-research-activities-supervision","title":"Who Is Cleared to Run It Alone: Entrustment Levels for Democratized Research (2026)","url":"https://www.koji.so/docs/entrustable-research-activities-supervision","summary":"Entrustment levels, borrowed from medical education, replace a single yes-or-no research permission with a per-person, per-activity clearance decision across five graduated supervision levels. The article builds a research activity inventory, sets out what evidence should move someone up, answers the Nielsen Norman Group critique of AI research tools, and separates the activities an AI-native platform genuinely de-risks from the judgement calls it does not.","content":"**The useful question is not whether a colleague has been trained to do research. It is which research activities they can be trusted to run unsupervised, which they can run with a researcher reachable, and which they should not touch yet.** Medical education has a mature framework for exactly this decision, and it transfers to research democratization almost unchanged.\n\nMost democratization programmes fail because they treat access as a training problem. Everyone sits a workshop, everyone gets a seat in the research tool, and six months later the insights library is full of leading questions and confident conclusions drawn from four interviews. The workshop was not the problem. The missing piece was a per-person, per-activity clearance decision, made on evidence, and written down.\n\n## What this article is not\n\nThis is not the question of whether your organisation should democratize research at all, or which activities belong to the research team as a matter of policy. That is covered in the [research democratization playbook](/docs/research-democratization-playbook), which works at the level of the organisation. This article works at the level of one person and one activity: who is cleared to do what, right now, and what evidence moved them up.\n\nThe distinction matters because the two questions have different answers. An organisation can decide that moderated interviews are democratized, and it can still be true that this particular product manager should not run one alone next week.\n\n## Why \"trained\" is the wrong gate\n\nA training record tells you what someone was exposed to. It does not tell you what they can do unobserved, and it does not vary by activity. Yet research competence is radically uneven across activities. The same person who writes a clean screener may be unable to resist pitching their roadmap in a live interview.\n\nNielsen Norman Group has been blunt about where this leads. Kara Pernice notes that in some circles, democratization of user research has come to mean that anyone can do user research, including people who know little about research and who carry no accountability or responsibility for doing it. Her analogy is the one to remember: \"As an analogy, most people can microwave a frozen meal, but only a trained, skilled chef can devise a menu for a 4-star restaurant.\"\n\nThe point is not that non-researchers should be kept out. It is that the tasks are not interchangeable, so a single yes-or-no permission is the wrong shape of decision. Pernice argues that the genuinely complex tasks, such as determining the best research method, identifying the research questions, and planning a contextual inquiry or a quantitative study, are done well only by someone with education, experience and skills.\n\n## The borrowed machinery: entrustable activities and levels of supervision\n\nIn 2005, Olle ten Cate proposed that postgraduate medical training should be organised not around abstract competencies but around the units of work a clinician actually does, and that the assessment question should be whether a trainee can be entrusted with each unit. The idea became known as the entrustable professional activity, or EPA.\n\nTwo features make it worth stealing.\n\nThe first is the distinction between a competency and an activity. A competency describes a person: knowledge, skills, attitudes, things a professional can carry around. An activity describes work. People possess competencies; they do not possess activities. Entrustment is therefore never a certificate about a person in general. It is always a decision about a person doing a specific thing in a specific setting.\n\nThe second is that supervision is graded rather than binary. The widely used scale runs over five levels: the learner may observe only; may act with direct, proactive supervision present in the room; may act with indirect, reactive supervision, where the supervisor is not present but is quickly available; may act unsupervised; and may supervise others. Progress is not a single graduation event. It is a series of small, documented entrustment decisions, each attached to one activity.\n\nTranslate the scale into research and it reads naturally:\n\n- **Level 1, observe.** Sit in on interviews and read transcripts. No contact with participants.\n- **Level 2, act with a researcher present.** Moderate with a researcher in the room or on the call, ready to intervene.\n- **Level 3, act with a researcher reachable.** Run the session alone, with a named researcher available before and after, and a review before anything is published.\n- **Level 4, act alone.** Run and report without review for this activity.\n- **Level 5, supervise.** Clear and coach other people on this activity.\n\n## Build the activity inventory first\n\nAn entrustment ladder is useless without a list of activities fine-grained enough that the levels can differ across them. Research is usually treated as one activity, which is why clearance decisions come out so crude. Split it:\n\n- Writing the research question and deciding whether research is needed at all\n- Choosing the method\n- Writing the discussion guide or the question set\n- Writing the screener and defining the sample\n- Recruiting and scheduling participants\n- Moderating a live session\n- Probing in the moment, including resisting the urge to pitch\n- Coding open text and building a codebook\n- Deciding when there is enough evidence to stop\n- Writing the recommendation\n- Presenting to a decision-making forum\n\nA typical product manager at a company six months into democratization might sit at level 4 on recruiting, level 3 on writing a question set, level 2 on moderating, and level 1 on deciding when to stop. That profile is informative in a way that \"completed the research workshop\" is not. It also tells you precisely what to teach next.\n\n## What evidence should move someone up\n\nThe entrustment literature is clear that the decision rests on observed performance of the activity, not on a course completion. Three rules keep it honest.\n\n**Observe the activity itself.** A person who can describe a good probe in a workshop has demonstrated knowledge of probing, not probing. Watch a real session.\n\n**Use more than one observation, by more than one observer.** A single good interview is weak evidence, and a single observer imports their own habits. This is the same logic that makes [inter-rater reliability](/docs/inter-rater-reliability-qualitative-research) worth measuring in the first place.\n\n**Write down the level and the date.** An undocumented entrustment decision silently becomes permanent. It should be reviewable, and it should be possible to move someone down after a bad study, which is far easier when the original decision was explicit.\n\n## The critique you have to answer\n\nThere is a serious objection to any argument that a tool can make research safer, and it is worth stating at full strength rather than dodging.\n\nIn March 2026, Maria Rosala of Nielsen Norman Group argued that research tools no longer merely host a study: they now plan it, moderate it and analyse it. Her warning follows directly. If those tools lack a solid foundation in research methodology, she writes, the risk is more than inconvenience; it is flawed research presented with confidence at large scale. On who will catch the mistakes, she is specific: \"The results will be flawed research plans, usability-testing tasks that prime participants, interview guides with leading questions, and AI facilitators who don't know how to moderate a usability test well.\"\n\nThat critique lands. It is also, read carefully, an argument for exactly this framework rather than against it. Rosala's objection is that an unsupervised novice plus an opinionated tool produces confident bad work. Entrustment levels are the mechanism that stops a novice from being unsupervised on the activities where the tool cannot help, while letting them move fast on the activities where it can.\n\nThe dishonest version of the AI argument is that the platform raises everyone to level 4 on everything. It does not. The honest version is narrower and more useful: a platform changes the entrustment level for some activities and leaves others untouched.\n\n## What an AI-native platform actually changes\n\nKoji moves the floor on the activities that are mostly a matter of applying a known pattern consistently, and leaves the judgement-heavy activities where they were.\n\nWhere Koji genuinely raises the floor:\n\n- **Question construction.** Koji ships methodology frameworks, and the Mom Test framework encodes its guardrails explicitly rather than leaving them to the moderator's memory. Its stated anti-patterns include never asking whether someone would use a product, because people mispredict their future behaviour, and never asking what they think of an idea, because the answer is a polite lie. Its question patterns are past-tense and concrete: walk me through how you currently handle the task, tell me about the last time the problem happened, what happened after that. A first-time interviewer using that framework in Koji starts from a better guide than most first drafts written unaided.\n- **Consistency of moderation.** A Koji AI interviewer asks the agreed question set the same way every time and probes against the configured points, so the variance that comes from an inexperienced human moderator having a bad afternoon is removed.\n- **Mechanical analysis.** Koji scores every interview for quality on a 1 to 5 scale with a breakdown into relevance, depth and coverage, which gives a reviewer a fast way to find the sessions worth reading closely.\n- **Instrument design.** Koji offers six structured question types, documented in the [structured questions guide](/docs/structured-questions-guide): open_ended, scale, single_choice, multiple_choice, ranking and yes_no. Picking a type is a constrained choice with a defined report output, which is a far easier thing to delegate than open-ended questionnaire design.\n\nWhere it changes nothing, and where your levels must stay low:\n\n- Deciding whether the research question is the right one\n- Deciding whether research is needed at all rather than a decision being made\n- Deciding when the evidence is sufficient, which is why [how many interviews is enough](/docs/how-many-interviews-enough) remains a judgement call\n- Deciding what to do about the answer\n\nThose four are where an unsupervised novice does the damage Rosala describes, and no amount of automated moderation touches them.\n\n## Permissions are not entrustment\n\nOne honest limitation, because it is easy to over-promise here. Koji team roles are owner, admin and member. That is a permissions model, and it is deliberately coarser than a five-level ladder across eleven activities. Nothing in Koji, or in any research platform we are aware of, stores \"this person is cleared at level 3 for moderating.\"\n\nSo the ladder lives in your process, not in a settings page. In practice teams keep it as a short table in the research wiki, review it quarterly, and use Koji roles for the blunt part of the control: who can publish a study, who can only draft one. The useful consequence of the gap is that the review step has to be a human habit. Pairing the ladder with a required pre-publication review by a named researcher at level 3 is what makes it real.\n\n## Common mistakes\n\n**Treating the ladder as a career grade.** Levels attach to activities, not to people. A senior PM can sit at level 4 on five activities and level 1 on another, and that is a normal profile, not a failure.\n\n**Granting level 4 by default to anyone senior.** Seniority is evidence about judgement in their own domain, not about moderating an interview without leading it.\n\n**Never demoting.** If a study goes wrong, the level should move down for that activity. A ladder that only goes up is a training record with extra steps.\n\n**Writing eleven activities and reviewing none of them.** The inventory is cheap and the observations are not. If you cannot afford to observe, start with three activities that matter most.\n\n**Confusing the clearance question with the policy question.** Keep [the organisational decision](/docs/research-democratization-playbook) separate from the per-person one, and make sure your [research maturity](/docs/user-research-maturity-model) assessment does not quietly assume that access equals capability.\n\n## Frequently asked questions\n\n### What is an entrustable professional activity?\n\nAn entrustable professional activity, or EPA, is a unit of real professional work that a learner can be entrusted to perform once they have shown they are ready. The concept was introduced by Olle ten Cate in 2005 as an alternative to assessing abstract competencies, and it is now widely used in medical education. The key move is that the assessment attaches to the activity rather than to the person: you do not certify a researcher in general, you decide that this person can run this activity at this level of supervision.\n\n### How is this different from a research training programme?\n\nA training programme records exposure; an entrustment ladder records demonstrated performance of a specific activity, observed by someone qualified to judge it. The practical difference is granularity and reversibility. Training is usually one event with one outcome, while entrustment is a set of per-activity decisions that are dated, reviewable and can be moved down as well as up.\n\n### What are the five levels of supervision?\n\nThe commonly used scale runs: observe only; act with direct supervision present in the room; act with indirect supervision where a supervisor is not present but quickly reachable; act unsupervised; and supervise others. In a research context these map onto sitting in on sessions, moderating with a researcher on the call, running sessions alone with a review before publication, running and reporting without review, and clearing other people.\n\n### Does an AI interviewer mean everyone can be cleared for moderating?\n\nNo, and claiming otherwise is the mistake Nielsen Norman Group warns about. An AI interviewer removes a specific risk, which is inconsistent human moderation, and Koji frameworks encode anti-patterns such as never asking whether someone would use a product. It does not decide whether the research question is worth asking, whether enough evidence has been collected, or what to do with the answer. Keep levels low on those activities regardless of tooling.\n\n### Which research activities should stay with trained researchers?\n\nThe judgement-heavy ones: choosing the method, deciding whether research is needed at all, deciding when the evidence is sufficient, and translating findings into a recommendation. Nielsen Norman Group singles out method selection, identifying research questions and planning contextual or quantitative studies as tasks done well only with education, experience and skills. Activities with a constrained, checkable output, such as picking a structured question type or recruiting to a defined screener, are the natural first candidates for delegation.\n\n### How do I document entrustment levels in Koji?\n\nKoji does not store entrustment levels; its team roles are owner, admin and member, which is a permissions model rather than a competence ladder. Keep the ladder as a short table outside the tool, and use Koji roles for the coarse control of who can publish a study versus draft one. Pair it with a required pre-publication review by a named researcher for anyone below level 4, and use the Koji per-interview quality scores to make that review fast.\n\n## Related Resources\n\n- [Research Democratization: The 2026 Playbook](/docs/research-democratization-playbook) - the organisation-level question of what gets democratized, which this article deliberately leaves aside\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and why a constrained instrument is the easiest thing to delegate first\n- [User Research Maturity Model](/docs/user-research-maturity-model) - where a clearance ladder fits in a wider maturity picture\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) - why one observer is not enough evidence to move someone up\n- [How Many Interviews Is Enough?](/docs/how-many-interviews-enough) - the judgement call that stays with trained researchers\n- [Research Workspace Roles and Permissions](/docs/research-workspace-roles-permissions) - what the Koji roles actually control\n- [Stakeholder Buy-In for User Research](/docs/stakeholder-buy-in-user-research) - making the case for a staged rollout rather than open access\n\n## Sources\n\n- ten Cate O. Entrustability of professional activities and competency-based training. *Medical Education*. 2005;39(12):1176-1177. doi:10.1111/j.1365-2929.2005.02341.x\n- ten Cate O. A primer on entrustable professional activities. *Korean Journal of Medical Education*. 2018;30(1):1-10. doi:10.3946/kjme.2018.76\n- ten Cate O, Chen HC, Hoff RG, Peters H, Bok H, van der Schaaf M. Curriculum development for the workplace using Entrustable Professional Activities (EPAs): AMEE Guide No. 99. *Medical Teacher*. 2015;37(11):983-1002. doi:10.3109/0142159X.2015.1060308\n- Chen HC, van den Broek WES, ten Cate O. The Case for Use of Entrustable Professional Activities in Undergraduate Medical Education. *Academic Medicine*. 2015;90(4):431-436. doi:10.1097/ACM.0000000000000586\n- Pernice K. Democratize User Research in 5 Steps. Nielsen Norman Group, 12 June 2022.\n- Rosala M. The Methodological Problems Hiding in Your Research Tools. Nielsen Norman Group, 13 March 2026.\n","category":"Research Methods","lastModified":"2026-10-06T07:54:52.037019+00:00","metaTitle":"Entrustment Levels for Democratized Research (2026)","metaDescription":"Decide which research activities each person can run unsupervised using graduated entrustment levels borrowed from medical education.","keywords":["research democratization governance","who can run user interviews","entrustable professional activities","supervision levels research","research enablement","democratized research guardrails"],"aiSummary":"Entrustment levels, borrowed from medical education, replace a single yes-or-no research permission with a per-person, per-activity clearance decision across five graduated supervision levels. The article builds a research activity inventory, sets out what evidence should move someone up, answers the Nielsen Norman Group critique of AI research tools, and separates the activities an AI-native platform genuinely de-risks from the judgement calls it does not.","aiPrerequisites":["Familiarity with basic user research methods","An existing or planned research democratization effort"],"aiLearningOutcomes":["Split research into activities fine-grained enough to clear separately","Assign and document graduated supervision levels per person per activity","Identify which activities an AI-native platform de-risks and which it does not","Avoid granting unsupervised clearance on judgement-heavy activities"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}