Content Testing: How to Test Microcopy, Labels, and UX Writing With Real Users (2026)
Six methods for testing whether your words actually work — cloze tests, highlighter tests, comprehension checks, term-choice tests, expectation tests, and label first-click — plus how to run them conversationally at scale instead of one participant at a time.
Content Testing: How to Test Microcopy, Labels, and UX Writing With Real Users (2026)
The bottom line: Content testing measures whether people understand and act correctly on your words — button labels, error messages, form hints, empty states, pricing copy, onboarding instructions. It is not proofreading, and it is not A/B testing: A/B tells you which variant converted, content testing tells you why one variant confused people, which is the part you can generalise to the next hundred strings. Six methods cover almost every case, and with an AI-moderated platform like Koji you can run them as a single conversational study across a hundred participants instead of scheduling twelve sessions.
Why words deserve their own test
Copy is the interface. A user does not experience your information architecture; they experience the eleven words on the screen that tell them what will happen if they click. Yet copy is usually the least-tested layer of a product: it changes late, it changes often, and it is written by people who have spent months internalising the exact vocabulary the user has never heard.
That last point is the core problem. Content testing exists because the author cannot un-know the meaning. You cannot self-assess comprehension of your own terminology, and neither can your colleagues — insiders share the company-specific assumptions that make unclear copy feel obvious. Testing with representative users is not a nice-to-have here; it is the only way to obtain the signal at all.
Three symptoms suggest you have a content problem rather than a design problem:
- Support tickets quote your own UI text back at you as the source of confusion.
- Users complete a task but describe what they did incorrectly afterwards.
- Two internal teams argue about a label and both cite "clarity" as the reason.
The six methods
1. Cloze test — does the copy hold together?
The classic comprehension measure. Take a passage, delete every nth word, and ask participants to fill in the blanks from context. For long-form prose, n = 6 is conventional; for microcopy you need a much smaller interval, and you want roughly 25 to 50 blanks in total to get a stable score.
Scoring: count exact and near-synonym matches. A commonly used rule of thumb is that copy is comprehensible when participants restore around 60% or more of the missing words. Below that, the text is not predictable enough — usually because sentence structure is convoluted or terminology is unfamiliar.
Use it for: legal-adjacent explanations, policy text, onboarding paragraphs, security or billing copy where misunderstanding is costly.
Do not use it for: three-word button labels. There is nothing to restore.
2. Highlighter test — where does confidence break?
Show the copy and ask participants to mark anything confusing, and separately anything reassuring. Digitally this is a two-pass exercise: first the friction, then the trust.
The output is a heat map of specific strings rather than an overall score, which makes it the most directly actionable method for a writer. In a conversational study you replicate it by asking participants to quote back the exact phrase that felt unclear, then probing why. Koji's follow-up probing does the "why" automatically, which is where the rewrite instruction actually comes from.
Use it for: pricing pages, plan comparison tables, consent and permission dialogs, anything where hesitation kills conversion.
3. Comprehension check — did the meaning survive?
Show the copy, remove it, then ask what it meant and what would happen next. The rule is to ask for a restatement in the participant's own words, never a yes/no "did you understand?" — everyone says yes.
Strong variants:
- "In your own words, what does this setting do?"
- "If you click this, what happens to your existing data?"
- "Who is this plan for?"
Use it for: destructive actions, permissions, data-sharing explanations, plan selection.
4. Term-choice test — which word wins?
Present two to five candidate terms for the same concept and ask which one participants would expect to lead where. This is nomenclature research and it is the highest-leverage content test in most products, because the term propagates across navigation, docs, support macros, and sales collateral.
Run it as a single_choice question for the winner plus a ranking question when you need the full preference order — then probe the choice: "what did you expect Workspace to contain that Project would not?" The reasoning matters more than the vote count, because it tells you what mental model the term evoked.
5. Expectation test — what do they think the button does?
Show the control in isolation, before the click. Ask what the participant expects to happen, how confident they are, and what would make them hesitate. A scale question captures confidence; the probed open-ended answer captures the reason.
This catches the most expensive class of copy failure: labels that are perfectly clear and describe the wrong thing.
6. Label first-click — can they find it at all?
Present the navigation or menu labels and ask where they would go to accomplish a specific task. Findability failures are frequently vocabulary failures in disguise; if 40% pick the wrong label, the problem is rarely the layout.
Pair it with the term-choice test: first learn what people call the thing, then verify that the label you chose is where they look.
Readability scores are not content testing
Flesch-Kincaid, grade-level scores, and their cousins measure sentence and word length. They do not know whether reconciliation, entitlement, or seat means anything to your audience, and they cannot detect a perfectly readable sentence that describes the wrong behaviour. Use readability tools as a drafting aid and a floor, never as evidence. Comprehension is a property of the reader, not of the text — which is why it has to be measured with readers.
Running content tests at conversational scale
The traditional constraint on content testing is throughput. Each method above is cheap to run once and painful to run twenty times: you schedule sessions, read copy aloud, take notes, and hand-code the answers. In practice teams test the redesign and skip the two hundred strings that ship every quarter.
An AI-moderated study removes the scheduling layer. A few specifics that matter for content work:
- Use text mode for reading tasks. Copy has to be seen. Text interviews let participants read the string and respond to it; voice is better suited to the expectation and reasoning parts of the study. Koji supports both from a single link, so a mixed design works.
- Mix structured and open in one pass. Term choice as
single_choice, preference order asranking, confidence asscale, comprehension restatement asopen_endedwith probing, and a quickyes_noon whether the participant would take the action. The six structured question types mean one study returns both the vote counts and the reasoning — no second round. - Let the AI ask the "why". The value in content testing sits entirely in the follow-up: not "which label did you pick" but "what did you expect that label to contain?" Koji probes each answer up to three times automatically, which is exactly the interviewer behaviour a form cannot replicate and a busy team rarely sustains by hand.
- Recruit outsiders. Never test copy on colleagues, and be careful with power users, who have already learned your vocabulary. Screen for the segment whose comprehension you actually care about — often new or prospective users.
- Watch for order effects. Showing variant A before variant B primes the comparison. Randomise where you can and keep the sequence identical across participants where you cannot, so the bias is at least constant. See question order bias.
How many participants?
More than qualitative usability work, less than a survey. Comprehension is a proportion, and proportions need a bit of sample: 20 to 30 participants gives a usable read on a term-choice test between two clearly different candidates; 50 or more if the options are close or you need to compare segments. Cloze tests are more forgiving because each participant contributes dozens of data points. For the reasoning layer, the usual qualitative saturation logic applies — see how many user interviews you need.
A ready-made study structure
A single 8-to-12 minute conversation can cover an entire feature's copy:
- Warm-up (
open_ended) — What do you use a tool like this for today? - Expectation (
open_ended+scale) — Here is the button. What happens if you click it? How confident are you, 1 to 7? - Comprehension (
open_ended, probed) — Here is the explanation text. In your own words, what does it mean for your data? - Highlight (
open_ended, probed) — Which exact phrase, if any, made you pause? Why that one? - Term choice (
single_choice+ probe) — Which of these would you click to find your saved work? - Preference order (
ranking) — Rank these four headings by how clearly they describe the page. - Action (
yes_no, probed) — Based only on this copy, would you turn the setting on?
Koji aggregates the structured items into distributions automatically and themes the open-ended answers with supporting quotes, so the writer receives "62% expected Archive to delete the file, and here are the eleven quotes explaining why" rather than a folder of recordings.
What to do with the results
- Rewrite against the failure, not the score. A 45% cloze score tells you the passage failed; the highlighter quotes tell you which clause did it.
- Fix the term everywhere at once. A nomenclature change that lands in the UI but not in docs, emails, and support macros creates a new comprehension problem.
- Keep a decision log. Record the tested alternatives and the winning rationale. It ends the recurring internal argument and it is the fastest onboarding artefact a new writer can get.
- Re-test after the rewrite. Content changes are cheap enough that a second wave is realistic, and comprehension gains are the easiest research win to demonstrate to stakeholders.
Related Resources
- Structured Questions in AI Interviews — the six question types used throughout this guide
- Cognitive Interviews — the sibling method for testing whether your questions are understood
- Open-Ended vs Closed-Ended Questions — when to count and when to probe
- Choice and Ranking Questions in AI Interviews — running term-choice and preference-order tests
- Question Order Bias — avoiding priming when comparing copy variants
- AI Usability Testing — where content testing fits inside a broader usability study
- How Many User Interviews Do You Need? — sample size for the qualitative layer
Related Articles
AI Usability Testing: How AI Moderates and Analyzes Usability Studies in 2026
A practical guide to AI usability testing in 2026 — what AI can moderate and analyze, where it fits alongside click-based testing, and how to capture the "why" behind every usability result.
Choice and Ranking Questions in AI Interviews: Capture Preference Data at Scale
Learn how to use single choice, multiple choice, ranking, and yes/no questions in Koji AI interviews — with automatic report charts that show preference distributions across all your participants.
Cognitive Interviews: How to Test Your Survey Questions Before You Launch
A practical guide to cognitive interviewing — the pretesting technique that reveals whether your survey questions and interview guides are understood as intended. Covers think-aloud, verbal probing, sample sizing, and AI-powered approaches.
How Many User Interviews Do You Need? The Sample Size Guide for Qualitative Research
Discover the right number of user interviews for your research. Learn about data saturation, theoretical saturation, and practical frameworks for knowing when you've collected enough qualitative data.
Open-Ended vs. Closed-Ended Questions: Examples and When to Use Each
Open-ended questions reveal the "why" in respondents'' own words; closed-ended questions deliver clean, countable data. Learn the difference, see examples of both, and discover why the best research pairs them — and how AI captures both at once.
Question Order Bias: How Survey & Interview Sequencing Skews Your Data (2026)
Why the sequence of your questions changes the answers — the classic Pew and Schwarz findings, the four main order effects, a practical sequencing checklist, and how AI moderation neutralizes the risk.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.