Back to docs
Research Methods

Medical Device Usability Testing: IEC 62366-1 and FDA Human Factors Requirements (2026)

Regulators do not accept "we tested it and users liked it." Here is how usability engineering actually works under IEC 62366-1 and the FDA's May 2026 human factors guidance — critical tasks, formative studies, summative validation, and the parts of the file AI-moderated research can legitimately fill.

Medical device usability testing is a regulated engineering process, not a design nicety. Under IEC 62366-1 and the FDA's human factors guidance, you must identify the tasks where a use error could hurt someone, design the interface to prevent those errors, and then prove with representative users that you succeeded. The FDA issued its final guidance, Content of Human Factors Information in Medical Device Marketing Submissions, on 29 May 2026, replacing the December 2022 draft — and it changed what many manufacturers have to submit.

The short version for teams under time pressure: summative validation still requires observed, simulated use with representative users performing critical tasks on the final design. No AI platform replaces that. But summative validation is the last 10% of the usability engineering file. The other 90% — understanding your user groups, discovering known use problems, testing whether people can understand your instructions for use, gathering formative feedback, and running post-market use surveillance — is interview and comprehension work, and that is exactly where a platform like Koji collapses weeks into days.

What the regulations actually require

Two documents drive the work, and they overlap heavily without being identical.

IEC 62366-1FDA human factors guidance
StatusHarmonised standard; FDA-recognised consensus standardAgency guidance (recommendations, but reviewers apply them)
Core artifactUsability engineering fileHFE/UE report in the marketing submission
Sample size for validation"Representative users" — no number givenMinimum 15 participants per distinct user group
Risk linkageVia ISO 14971:2019 risk managementVia use-related risk analysis (URRA)
EmphasisProcess conformity and traceabilityEvidence that critical-task risk is acceptable

The unifying concept is use error: not "user error." The standard deliberately avoids blaming the person. A use error is an act or omission that produces a different result than the manufacturer intended — and the assumption is that the interface invited it. Your job is to find the acts and omissions that could cause harm, then engineer them out.

That chain runs: use specification → user profiles → known use problems → task analysis → use-related risk analysis → critical tasks → formative evaluation → summative validation → HFE/UE report.

What changed in the May 2026 FDA guidance

The final guidance keeps the risk-based framework but makes it meaningfully less mechanical. Five changes matter for research planning:

  1. Three human factors submission categories. Category 1 applies to certain modified devices and asks for a conclusion plus a high-level summary of the HF evaluation. Category 2 applies when you can document why there are no critical tasks — or, for a modification, no new or impacted critical tasks. Category 3 requires a full HFE/UE report including validation testing data.
  2. A new Decision Point D. The flowchart now asks whether validation data actually needs to be submitted, weighing the user interface's history of use, its complexity, and the adequacy of existing risk controls. Having a critical task no longer automatically forces a new validation study.
  3. Justification in place of testing. For modifications and well-understood interfaces with a safe use history, a rigorous justification built on a comprehensive URRA can substitute for new validation data.
  4. Leveraging existing data. The FDA explicitly encourages referencing prior submissions and data from similar devices rather than duplicating studies.
  5. Far more worked examples. The example section roughly tripled versus the draft, now covering pediatric users, augmented-reality interfaces, and devices with known use-related problems.

The practical consequence is that the quality of your use-related risk analysis and your evidence about real-world use now carries more weight than the number of studies you ran. Justification-based pathways only survive review when the underlying understanding of users, environments and known problems is deep and documented. That raises the value of upstream research considerably.

Context for why reviewers are strict: there were 111 Class I recall events and early alerts in 2025, and 2024 saw 1,059 US device recall events — a four-year high, with Class I recalls at their highest level in fifteen years. Design-related causes lead the list. Every use error found in a formative study is one that does not become a field action.

Step 1 — Write the use specification with users, not about them

The use specification defines intended medical indication, patient population, intended user profiles, use environment and operating principle. Most teams write it from internal assumptions, then discover in summative testing that a whole user group was missing.

Interview each candidate user group before you write it. The questions that matter are unglamorous: Who actually performs this step in your clinic at 3am? What else is happening in the room? What do you do when the alarm sounds and you are two rooms away? Which steps do experienced staff skip? What did the previous device get wrong?

This is where AI-moderated interviews earn their keep. Clinicians are hard to schedule and geographically scattered; median physician response rates in research sit around 18%, with a range roughly 10–60%. An asynchronous study that a nurse can complete by voice at the end of a shift, in her own language, without a moderator's calendar, converts far better than a booked video call. Koji's AI interviewer probes automatically when an answer is thin — "you said you usually skip the priming step; walk me through the last time you did that" — so you get the incident detail that a static survey never surfaces.

Step 2 — Hunt for known use problems

IEC 62366-1 expects you to consider known use problems with your device and with similar devices. Teams typically satisfy this with a MAUDE search and a literature scan, then stop. That is a thin file.

A stronger approach adds primary evidence:

  • Interview users of the predicate or competitor device about workarounds and near misses.
  • Interview your own service and complaints staff — they hold the richest failure narratives in the company.
  • Interview trainers about which steps consistently need re-teaching.

Structure these so they aggregate. In Koji, a multiple_choice question that lists candidate problem steps gives you frequency, a scale question captures perceived severity, and an open_ended follow-up captures the story behind each selection. You end up with a ranked, quotable, traceable input to the URRA rather than a folder of notes.

Step 3 — Task analysis and the use-related risk analysis

Decompose use into tasks and subtasks, then for each ask what could go wrong perceptually, cognitively, and in action. Which failures could lead to harm? Those are your critical tasks, and they define the scope of validation and the depth of the submission.

Two disciplines separate a strong URRA from a weak one. First, derive tasks from observed practice, not from the draft IFU — the IFU describes intended use, and use errors live in the gap between intended and actual. Second, keep the traceability explicit: every critical task should trace to a hazardous situation in the ISO 14971 file and to a risk control, and every risk control should trace to the evaluation that tested it.

Step 4 — Formative evaluation: cheap, early, repeated

Formative studies are exploratory. They have no pass/fail criteria, they can use prototypes, and their entire purpose is to find and fix problems before validation. The classic finding is that about five representative users surface roughly 85% of usability problems and ten reach about 95% — which is why running three small formative rounds beats running one big one.

Formative work splits cleanly into two kinds:

  • Interaction studies — someone must handle the device or prototype while an observer watches. Do these in person or over supervised video. This is not something to automate.
  • Comprehension and expectation studies — does the label make sense, does the alarm mean what people think it means, does the IFU step read as one action or two, would a user expect the device to do X after Y? These are pure comprehension work, and running them asynchronously with dozens of clinicians instead of six is a strict upgrade.

Koji handles the second category natively. Show the artefact, ask a single_choice recall question with one correct answer and plausible distractors, capture self-rated clarity on a scale, then let the AI probe every wrong answer with an open_ended follow-up to learn why it was misread. Set the pass criterion before you field it — for example, 90% correct identification of the dose-confirmation step — and record the before-and-after result across redrafts. That produces exactly the kind of documented, criterion-based evidence a reviewer can follow.

Step 5 — Summative validation: what cannot be automated

Be clear with yourself and with your quality team about this boundary.

ActivityCan AI-moderated research do it?
Use specification and user-profile interviewsYes — voice or text, async, any language
Known use problems and near-miss discoveryYes
IFU, labeling and training comprehension testingYes, with pre-set pass criteria
Formative feedback on concepts, alarms, wordingYes
Formative hands-on interaction studiesNo — requires observation of use
Summative human factors validationNo — requires observed simulated use with the final design
Post-market use surveillance and complaint follow-upYes

Summative validation means representative users from each distinct user group performing critical tasks in a realistic simulated-use environment, with observation of what they do, followed by a knowledge-task debrief and root-cause analysis of every use error and difficulty. The FDA expects a minimum of 15 participants per user group, more for higher-risk products, and a root cause for each observed problem — not a satisfaction score.

Where AI-moderated research does contribute to summative work is the debrief. Structured post-task interviewing at scale, with consistent probing and automatic transcription, removes a real source of variability between moderators. But the observation itself stays human.

Step 6 — Post-market use surveillance

Both frameworks expect production and post-production information to feed back into the usability engineering file. Complaint data tells you that something went wrong; it rarely tells you why. A short quarterly AI interview study with active users — "walk me through the last time the device did something you did not expect" — turns a complaint trend into an identified use error with a candidate root cause. That evidence is what supports your next submission's justification-based pathway under the 2026 framework.

Common mistakes that cost submissions

  • Treating validation as a usability test. Reviewers want risk evidence, not SUS scores. A benchmark study is not a summative validation.
  • Missing a user group. Home caregivers, cleaning and reprocessing staff, and remote monitoring staff are routinely forgotten, and each needs its own 15.
  • Testing the IFU you wish you had. Validate the final labeling, not a cleaned-up draft.
  • No pass criteria in advance. A comprehension result without a pre-stated threshold reads as post-hoc.
  • Untraceable findings. If a use error in the report cannot be traced to a task, a hazard and a risk control, the file does not hold together.
  • Confusing formative and summative. Formative studies have no acceptance criteria; using formative data as validation evidence is a predictable rejection.

How Koji fits a regulated programme

Koji is an AI-native research platform, not a regulatory submission tool. It runs AI-moderated interviews by voice or text, probes follow-ups automatically, supports six structured question types — open_ended, scale, single_choice, multiple_choice, ranking and yes_no — and produces analysis with quote-level traceability back to transcripts. In a human factors programme that means: faster access to scattered clinicians, comprehension evidence with pre-set criteria, multilingual studies for multi-market submissions, and exports and transcripts you can attach to the usability engineering file as method and data lineage.

What it does not do is watch someone use your device. Plan the observed studies properly, and use AI research to make everything around them faster and better evidenced.

Frequently asked questions

Can AI-moderated interviews replace summative human factors validation? No. Summative validation requires representative users performing critical tasks with the final design in a realistic simulated-use environment, observed so that use errors and difficulties can be recorded and root-caused. AI-moderated research covers the work around it: use specification interviews, known use problem discovery, IFU and labeling comprehension testing, formative comprehension studies and post-market use surveillance.

How many participants does FDA expect in a summative usability study? A minimum of 15 participants per distinct user group, with more expected for higher-risk devices. Distinct user groups are counted separately, so a device used by nurses, home caregivers and reprocessing staff needs 15 of each. IEC 62366-1 itself specifies representative users without naming a number.

What changed in the FDA human factors guidance issued in May 2026? The final guidance, published 29 May 2026, replaced the December 2022 draft. It confirms three human factors submission categories, adds Decision Point D — which weighs user interface history of use, complexity and existing risk controls before requiring validation data to be submitted — permits robust justification in place of new testing for well-understood interfaces, encourages leveraging data from prior submissions and similar devices, reorders the HFE/UE report sections, and roughly triples the worked examples.

What is the difference between formative and summative usability evaluation? Formative evaluation happens during development, uses prototypes, has no pass or fail criteria, and exists to find and fix problems. Summative evaluation is the final validation of the finished design against pre-defined acceptance criteria with representative users performing critical tasks. Presenting formative data as validation evidence is a predictable reason for a deficiency letter.

How do we test whether users understand our instructions for use? Run a comprehension test rather than a satisfaction survey. Define the key points the IFU must convey, set a pass threshold in advance, present the real artefact, test recall with single_choice questions that have plausible distractors, capture self-rated clarity on a scale question, and probe every wrong answer with an open_ended follow-up. Redraft and retest until the threshold is met, and record the before-and-after result.

Does a device modification always require a new validation study? Not since the 2026 final guidance. If a modification introduces no new or impacted critical tasks it can fall into Category 2 with documented justification, and Decision Point D allows well-understood interfaces with a safe use history and adequate risk controls to rely on justification and leveraged data. That places more weight on the quality of the use-related risk analysis and on real-world evidence about how the device is actually used.

Related resources

Related Articles

Accessibility Research: How to Include Users with Disabilities in Your Studies

A practical guide to designing and conducting accessible user research — how to recruit participants with disabilities, adapt your methods, and use async AI interviews to remove barriers to participation.

AI-Powered Patient and Provider Research for Healthcare

How healthcare organizations use Koji to conduct patient experience research, provider feedback studies, and clinical workflow analysis at scale — while maintaining HIPAA-aware research practices.

Content Testing: How to Test Microcopy, Labels, and UX Writing With Real Users (2026)

Six methods for testing whether your words actually work — cloze tests, highlighter tests, comprehension checks, term-choice tests, expectation tests, and label first-click — plus how to run them conversationally at scale instead of one participant at a time.

Formative vs. Summative Research: When to Use Each Method (And Why It Matters)

Formative research shapes a product while it's still being built. Summative research evaluates how it performs after it ships. Confusing the two is the most common reason research budgets get wasted on the wrong question at the wrong time.

HIPAA-Compliant AI User Research: A Practical Playbook for Healthcare and HealthTech

Run AI-moderated customer research in healthcare contexts without putting PHI at risk. Patterns for HIPAA alignment, anonymous-mode interviews, BYOK, sub-processor handling, and what Enterprise teams need from a vendor.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Usability Metrics: Task Success Rate, Time on Task, and Error Rate Explained

The complete guide to the core usability metrics — task success rate, time on task, and error rate — including industry benchmarks, formulas, sample sizes, and how to capture them automatically with AI-moderated research.

How to Conduct Usability Testing: The Complete Guide

A comprehensive guide to usability testing for UX researchers and product managers. Covers types of testing, participant numbers, step-by-step facilitation, and the most common mistakes to avoid.