Clickwrap vs Browsewrap: How to Research Whether Users Actually Agreed to Your Terms
Courts decide whether your terms are enforceable by asking what a reasonably prudent Internet user would have seen and understood. That is an empirical question. Here is how to answer it with evidence instead of opinion.
Whether your terms of service bind a user is decided by a two-part test: the interface must give reasonably conspicuous notice of the terms, and the user must take an action that unambiguously manifests assent to them. Both prongs turn on what a hypothetical "reasonably prudent Internet user" would have seen and understood, which is an empirical claim about real people. Almost nobody tests it. Teams ship a sign-up screen, a court later reconstructs it from screenshots, and a design decision made in an afternoon determines whether an arbitration clause, a limitation of liability, or a subscription term survives. Running a short comprehension study on your own sign-up flow converts that guess into a record. With an AI interview platform like Koji you can put the actual screen in front of 60 to 150 real users, ask what they noticed and what they believed they were agreeing to, and have the analysis back the same day.
The test courts actually apply
In Berman v. Freedom Financial Network, LLC, 30 F.4th 849 (9th Cir. 2022), the Ninth Circuit set out the rule plainly. Unless the operator can show the consumer had actual knowledge of the agreement, an enforceable contract exists on an inquiry notice theory only if:
- the website provides reasonably conspicuous notice of the terms to which the consumer will be bound; and
- the consumer takes some action, such as clicking a button or checking a box, that unambiguously manifests assent to those terms.
The court drew directly on the Second Circuit's formulation in Specht v. Netscape Communications Corp., 306 F.3d 17 (2d Cir. 2002), which held that reasonably conspicuous notice of contract terms and unambiguous manifestation of assent are essential if electronic bargaining is to have integrity and credibility. The other two anchors are Nguyen v. Barnes & Noble Inc., 763 F.3d 1171 (9th Cir. 2014), which refused to enforce an arbitration provision the consumer did not unambiguously assent to, and Meyer v. Uber Technologies, Inc., 868 F.3d 66 (2d Cir. 2017), which upheld an agreement where the hyperlinks were both blue and underlined and the screen warned that creating an account meant agreeing to the terms.
The three interface patterns
| Pattern | How assent is claimed | Typical enforceability |
|---|---|---|
| Clickwrap | User checks a box or clicks a button labelled to indicate agreement, next to visible terms | Strongest. The action itself carries the legal meaning |
| Sign-in wrap | User clicks a general button such as Continue or Sign Up, with a nearby line saying that doing so means agreeing | Contested. Falls in a gray zone where enforceability depends on conspicuous textual notice tied to the specific act |
| Browsewrap | Terms sit behind a link somewhere on the page; continued use is said to imply agreement | Weakest. Usually fails without explicit textual notice that continued use manifests intent to be bound |
Most modern sign-up flows are sign-in wrap. That is precisely the pattern courts scrutinise hardest, and precisely the one whose outcome depends on pixel-level design choices that no one on the team ever validated with users.
Why this is a research problem, not a legal one
Read the Berman opinion closely and it is a usability critique written by judges. The notice was printed in a tiny gray font considerably smaller than the surrounding text, so small the court called it barely legible to the naked eye. The larger surrounding text, the court said, naturally directs the user's attention everywhere else. The hyperlinks were underlined but appeared in the same gray as the rest of the sentence rather than in blue, the color typically used to signify a hyperlink. Consumers, the court held, cannot be required to hover over plain-looking text or aimlessly click words in an effort to ferret out hyperlinks.
Then comes the sentence that should reorganise how product teams think about this. Because online providers have complete control over the design of their websites, the onus must be on website owners to put users on notice of the terms to which they wish to bind consumers. Design control is what creates the burden. You chose the font size, the contrast, the button copy and the visual hierarchy, so you own the question of whether they worked.
The second prong is even more clearly empirical. Merely clicking a button, the court said, viewed in the abstract, does not signify agreement to anything. A click counts only if the user is explicitly advised that the act of clicking will constitute assent. The webpages in Berman said "I understand and agree to the Terms & Conditions" but never told the user which action would constitute that agreement. The court offered the fix in one line: language such as "By clicking the Continue button, you agree to the Terms & Conditions."
That is a testable hypothesis. Does your copy tell users what act binds them, and do they come away believing they were bound? A comprehension study answers it in days. Judicial reconstruction answers it years later, expensively, and once.
What to measure
The mistake teams make is testing whether users can find the terms. That is not the standard. The standard is what a reasonably prudent user would have noticed and understood while behaving normally, which means your study has to let people behave normally first and ask questions second.
Structure it in four layers, moving from unprompted to prompted. Koji supports all six structured question types in the same session, so a single study can carry the whole ladder without splitting it across tools.
| Layer | What you are establishing | Question type to use |
|---|---|---|
| Unaided notice | Did anything about terms register at all, without being led | open_ended, with AI follow-up probing |
| Aided recognition | Shown the screen again, can they point to the notice | yes_no then open_ended |
| Assent comprehension | Do they know which action bound them | single_choice across the plausible actions |
| Belief about content | What do they think they agreed to | multiple_choice plus ranking on perceived importance |
| Confidence | How sure are they, which separates knowledge from guessing | scale |
A few design rules that make the difference between a study a lawyer can use and one they cannot:
- Do not lead. Asking "did you see the terms of service link?" manufactures the yes you were hoping for. Start with a completely open recall question and let the AI probe what the participant raises on their own. This is the same discipline described in open-ended vs closed-ended questions and the reason leading phrasing shows up in survey response bias.
- Use a control screen. Run a second cell with a corrected design, blue and underlined links, explicit button-tied language, adequate contrast. The comparison is what makes the finding causal rather than descriptive. Without a control, every result is vulnerable to the objection that the question itself produced the answer.
- Recruit to your actual user base. A study on general consumers tells you little about a flow that only enterprise buyers see. See how to find and recruit research participants and research panel management.
- Let people act before you ask. Put the screen in front of them as a task, not as an exhibit. Moderated usability testing and unmoderated usability testing both work; Koji's AI moderator gives you the probing depth of the former at the scale of the latter.
Reading the results honestly
The single most common failure is treating a comfortable number as a pass. If 40 percent of participants cannot say which action bound them, that is not a communications problem to be solved with a help article. It is direct evidence on the second prong of the legal test, and it will be just as true when someone else runs the study.
Watch for three specific patterns:
The recognition illusion. Aided recognition is almost always far higher than unaided notice. People shown a screen and asked whether they saw a line of text will say yes. That gap is the whole reason the unaided layer has to come first, and it is why a study that only asks the aided question overstates conspicuousness badly.
Confident wrongness. A participant who says with high confidence that they agreed to nothing binding is a worse result than one who says they are not sure. Pairing a scale confidence question with the comprehension question separates these. The same logic underpins warranty comprehension research and testing a price display without making a deceptive claim.
Segment collapse. Mobile users see a materially different screen from desktop users, and small screens hide text that large ones reveal. Report the two separately. Averaging them produces a number that describes nobody.
If your result is that the current design performs no worse than the corrected one, resist the urge to call that a win. A non-significant difference in an underpowered study is not evidence of equivalence, a trap covered in equivalence testing and what no difference really means and in statistical power and minimum detectable effect.
How Koji fits
Traditional survey tools make this study harder than it needs to be. A static questionnaire in SurveyMonkey or Typeform cannot ask why someone missed the notice, cannot chase an ambiguous answer, and cannot tell the difference between a participant who understood the agreement and one who pattern-matched to the phrase terms of service. Qualtrics can field the quantitative layer but leaves you to schedule and moderate the qualitative half separately.
Koji runs the whole ladder in one conversational interview, by voice or by text, with no moderator scheduling. The AI asks the open recall question, hears "I think there was something at the bottom", and follows up on its own to find out whether the participant knew what it was for. Structured questions carry the countable part, so you get a distribution for the comprehension question and a percentage for the assent question alongside the verbatim explanations. Analysis is automatic, so the report exists the moment fielding closes rather than a fortnight later.
That speed matters here more than in most research, because assent design changes ship constantly. A growth team testing a shorter sign-up flow is, whether it knows it or not, running an experiment on contract formation. Being able to attach a comprehension check to that experiment for the cost of an afternoon is the difference between knowing and hoping.
Common mistakes
- Testing the terms instead of the screen. Nobody reads the terms. The legal question is about the notice and the act, not the document.
- Running it only after a dispute. Research commissioned during litigation carries an obvious credibility problem. Research run in the ordinary course of product development does not, which is the central argument of survey evidence in court.
- Ignoring the flow that most users take. Test the path with the highest volume, including any express checkout, social sign-in or invite-acceptance path that bypasses the main screen.
- Forgetting accessibility. A notice that fails contrast requirements fails conspicuousness for a large group of users by definition. See accessibility compliance research.
- Not writing down the method. An undocumented study is not evidence. Record the population, the recruitment, the exact wording and the screens shown.
Frequently asked questions
Is clickwrap always enforceable and browsewrap never?
No. Clickwrap is the strongest pattern because the user's action carries the legal meaning directly, and browsewrap is the weakest because nothing the user does signals agreement. But enforceability turns on the specific design, not the label. A clickwrap box buried in illegible gray text can still fail the conspicuousness prong, and courts have enforced browsewrap-style terms where explicit textual notice told users that continued use would bind them.
How many participants do I need?
For a two-cell comparison of a current screen against a corrected one, 60 to 150 per cell is a reasonable planning range, driven by how large a difference you need to detect rather than by any fixed rule. Work it out in advance using statistical power and minimum detectable effect rather than picking a round number, because an underpowered study that finds no difference is easily misread as a clean bill of health.
Does research like this actually get used in court?
Survey evidence is routinely admitted on questions of consumer perception, and Federal Rule of Evidence 703 long ago settled the sampling and hearsay objections by redirecting attention to the validity of the techniques used. Whether a specific study is admitted depends on its methodology. The practical value is usually earlier than trial anyway: a study that shows a notice is not working lets you fix it before anyone is disputing it.
Should the study be run by the legal team or the product team?
Product teams should run it as part of normal design work, with legal reviewing the instrument. Research produced in the ordinary course of business is more credible and far cheaper than research commissioned once a dispute exists. Keep the outputs in a shared research repository so the record survives staff turnover.
What about consent flows for privacy rather than contract terms?
Same method, different legal standard. Privacy consent has its own requirements around symmetry and clarity, covered in dark patterns testing, and research-participant consent is a separate topic again, handled in research consent form templates and research ethics and informed consent.
Can I run this on a competitor's sign-up flow?
Yes, and it is a useful benchmark, since the reasonably prudent user standard is informed by prevailing design conventions. Test two or three comparable flows alongside your own so you can say where yours sits rather than only whether it passed.
Related Resources
- Structured Questions Guide - the six question types and when to use each
- Warranty Comprehension Research - the same comprehension method applied to warranty documents
- Reference Prices and Drip Pricing - testing a price display without making a deceptive claim
- Dark Patterns Testing - measuring whether an interface impairs autonomous choice
- Survey Evidence in Court - what makes a study hold up under cross-examination
- Moderated Usability Testing - running the task portion of a notice study
Related Articles
Moderated Usability Testing: How to Run Sessions That Surface Real Problems (2026 Guide)
A practical 2026 guide to moderated usability testing: how to write tasks, run think-aloud sessions, measure task success and SEQ, choose sample size, and scale moderation with AI on Koji.
Open-Ended vs. Closed-Ended Questions: Examples and When to Use Each
Open-ended questions reveal the "why" in respondents'' own words; closed-ended questions deliver clean, countable data. Learn the difference, see examples of both, and discover why the best research pairs them — and how AI captures both at once.
Reference Prices and Drip Pricing: How to Test a Price Display Without Making a Deceptive Claim
Willingness-to-pay research tells you what people will pay. It says nothing about whether your was-now price or your checkout fees are lawful. Here is the price-display law that changed in 2025 and the comprehension study that produces evidence for it.
Research Ethics and Informed Consent: A Practical Guide for UX Teams
A practical guide to ethical UX research — covering the Belmont Report's three principles, GDPR informed consent requirements, how to handle AI tools responsibly, and how to build ethical maturity in your research practice.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Warranty Comprehension Research: Testing Whether Buyers Understand What Your Warranty Covers
The Magnuson-Moss Warranty Act requires warranty terms in simple and readily understood language, but never defines the test. Here is how to measure what buyers actually understand about coverage, remedy, and process.