{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-10T16:16:33.497Z"},"content":[{"type":"documentation","id":"fe3b0ede-7200-4956-94ce-1a0e60a56657","slug":"dark-patterns-user-testing","title":"Dark Patterns: How to Test a Flow for Deceptive Design Before a Regulator Does","url":"https://www.koji.so/docs/dark-patterns-user-testing","summary":"Dark pattern rules judge effect rather than intent, which makes deceptive design a measurement problem. This guide gives an operational test: a variant that lifts conversion while lowering accurate user belief is a dark pattern by regulators own logic. It covers the FTC 2022 staff report four families, DSA Article 25 (including the widely misreported paragraph 2 carve-out), California symmetry in choice under 11 CCR 7004, control-cell study design, and magnitude evidence from Luguri and Strahilevitz (11.3 to 37.2 percent) and Mathur et al (1,818 instances across 11,000 sites).","content":"**A dark pattern is defined by its effect on the user, not by anyone's intention.** The California Privacy Protection Agency put it plainly in Enforcement Advisory 2024-02, issued 4 September 2024: what matters is effect rather than intent, and consent obtained through a dark pattern is not valid consent at all. That single doctrinal fact reorganises the whole problem. You cannot audit your way to safety by asking the team whether they meant to manipulate anyone, because their answer is legally irrelevant. The only way to know whether an interface deceives or materially impairs a user's ability to decide freely is to put it in front of users and measure what they came away believing. Conversion analytics cannot do that. A conversational research platform like Koji can, and can do it fast enough to sit inside a normal design cycle.\n\n## Why your A/B test cannot tell you\n\nThe FTC's September 2022 staff report *Bringing Dark Patterns to Light* contains an example that should be pinned above every growth team's desk. According to the FTC's complaint, Credit Karma ran an A/B test comparing telling consumers they were \"pre-approved\" for a credit card against telling them they had \"Excellent\" odds of approval. The pre-approved claim was, the FTC alleged, false. The A/B test showed it yielded a greater click rate. The company deployed it.\n\nNothing in that process was unusual. It is textbook conversion rate optimisation, executed competently. **Conversion optimisation and dark pattern design are the same activity, distinguished only by whether you also measure what the user believed.** An A/B test scores a variant on whether people clicked. A dark pattern is precisely a variant that increases clicks by degrading understanding. The instrument is structurally blind to the harm.\n\nThis gives you an operational definition you can actually run against, and it is the most useful thing in this guide:\n\n| | Comprehension intact | Comprehension degraded |\n| --- | --- | --- |\n| **Conversion up** | Genuine improvement. Ship it | **Dark pattern signature.** The lift is being bought with misunderstanding |\n| **Conversion down** | Honest friction. May still be correct, for example an age check or a confirmation step | Simply bad design. Fix or revert |\n\nThe top-right cell is the one that matters. A variant that lifts conversion while lowering the share of users who can correctly state what they signed up for, what they will be charged, or what data they shared is a dark pattern by the regulators' own logic, whatever the team intended. Running a comprehension measurement alongside your conversion metric turns an unanswerable question about intent into a two-number comparison.\n\n## What the rules actually say\n\nThree regimes matter for most product teams, and they are not identical.\n\n**The FTC (United States).** The 2022 staff report groups common dark patterns into four families: design elements that induce false beliefs; that hide or delay disclosure of material information; that lead to unauthorized charges; and that obscure or subvert privacy choices. The report cites the FTC's action against LendingClub, where the company allegedly promised no hidden fees while placing the fee disclosure behind tooltip buttons consumers were unlikely to click and burying it in an un-bolded itemisation sandwiched between more prominent bolded paragraphs. Note the mechanism: nothing was concealed. Everything was disclosed. It was disclosed in a way that did not work.\n\n**The EU Digital Services Act.** Article 25 of Regulation (EU) 2022/2065, applicable from 17 February 2024, provides that providers of online platforms shall not design, organise or operate their online interfaces in a way that deceives or manipulates the recipients of their service, or in a way that otherwise materially distorts or impairs their ability to make free and informed decisions. Article 25(3) names three practices the Commission may issue guidance on: giving more prominence to certain choices; repeatedly requesting a choice the user has already made, especially through pop-ups that interfere with the experience; and making termination harder than subscription.\n\nOne detail is widely reported backwards and worth getting right. Article 25(2) says the prohibition **does not apply** to practices already covered by the Unfair Commercial Practices Directive or the GDPR. Article 25 is a gap-filler, not an extra layer stacked on top of existing law. Practically, that means a privacy dark pattern is a GDPR problem and a misleading commercial claim is a UCPD problem, and Article 25 catches the manipulation that falls between them.\n\n**California.** The CCPA regulations at 11 CCR section 7004 require that methods for submitting requests and obtaining consent use easy-to-understand language and provide symmetry in choice. Symmetry means the path to the more privacy-protective option must not be longer, harder or more time-consuming than the path to the less protective one. The regulations give a concrete illustration: offering only \"Yes\" and \"Ask me later\" is asymmetric, because the symmetrical pairing is \"Yes\" and \"No\".\n\nSymmetry is unusual among these rules in being measurable without any users at all. Count the clicks, the screens and the seconds on each path. If they differ, you have a finding before you recruit anybody. Do that audit first; it is free.\n\n## Designing the study\n\nMap each of the four FTC families to something you can put a number on. The point is to measure belief and choice accuracy, not satisfaction.\n\n| Dark pattern family | What to measure | Question types |\n| --- | --- | --- |\n| Induces false beliefs | Share of users who hold a specific incorrect belief after the flow, versus a control | single_choice for the belief, scale for confidence |\n| Hides or delays material information | Unaided recall of the fee, term or limitation before it is pointed out | open_ended with AI follow-up, then yes_no on awareness |\n| Leads to unauthorized charges | Whether users can state what they will be charged, when, and how to stop it | open_ended plus multiple_choice on expected charges |\n| Obscures privacy choices | Whether the choice users made matches the choice they intended | single_choice on intent versus logged actual setting, ranking on which settings matter most |\n\nFive rules separate a study that will survive scrutiny from one that will not:\n\n- **Always run a control cell.** A finding that 30 percent of users held a false belief means nothing on its own; people misremember, guess, and satisfice. What matters is the gap between your variant and a neutral version. This is exactly why control groups are the backbone of deception research, as covered in [survey evidence in court](/docs/survey-evidence-court-daubert).\n- **Measure before you reveal.** Ask what the participant expects to be charged before you show them the terms. Once you have shown them, the measurement is gone for that participant.\n- **Do not ask people whether they felt manipulated.** They will underreport. The FTC report notes that many consumers do not realise they are being manipulated at all, and that some who do realise it stay quiet out of embarrassment. Measure accuracy of belief, which people cannot flatter themselves about, not perception of fairness.\n- **Watch for acquiescence.** Agree or disagree formats carry a documented inflation effect of roughly 10 percent, so a flat \"this page was clear, agree or disagree\" will overstate clarity. Use neutral, specific questions instead. See [survey response bias](/docs/survey-response-bias) and [open-ended vs closed-ended questions](/docs/open-ended-vs-closed-ended-questions).\n- **Test on the device people use.** The FTC report notes that some techniques work better on smaller screens, because the scrolling required makes it unlikely people will see hidden information, and that this may fall hardest on lower-income users who rely on a phone as their only internet access. A desktop-only study can miss the entire effect.\n\n## What the evidence says about magnitude\n\nTwo findings are worth carrying into any internal argument about whether this is a real risk.\n\nLuguri and Strahilevitz, writing in the *Journal of Legal Analysis* in 2021, ran large experiments exposing consumers to dark patterns while signing up for a dubious data protection plan. Acceptance ran at 11.3 percent in the control condition, 25.4 percent under mild dark patterns and 37.2 percent under aggressive ones. Mild manipulation more than doubled sign-ups; aggressive manipulation more than tripled them. The FTC report also records that the effects grew significantly when subjects were exposed to more than one pattern at a time, which matters because real flows stack them.\n\nMathur and colleagues, in work presented at CSCW in 2019, analysed roughly 53,000 product pages across about 11,000 shopping websites and identified 1,818 dark pattern instances spanning 15 types in 7 broader categories, along with 183 sites engaged in outright deceptive practices. They also found 22 third-party entities selling dark patterns as a turnkey product. That last number is the uncomfortable one: a meaningful share of this is bought from a vendor rather than designed in-house, which means it can enter your product through a procurement decision that never reached design review.\n\n## How Koji fits\n\nThe measurement this requires is awkward for traditional tools. You need a participant to go through a realistic flow, then answer an unaided recall question, then be probed on what they meant, then answer countable structured questions, all without being tipped off about what you are testing. A static form in SurveyMonkey, Typeform or Qualtrics can do the countable part but cannot probe, so ambiguous answers stay ambiguous and you cannot tell a confused participant from a careless one.\n\nKoji runs it as a single conversational interview, by voice or text, with no moderator to schedule. The AI asks the open question first, hears \"I think it was about twenty dollars a month, maybe\", and follows up automatically to find out whether that was a memory or a guess. All six structured question types sit in the same session, so the belief-accuracy percentage and the confidence distribution come back alongside the verbatim explanations that tell you which design element caused the error. Analysis is automatic, so a comprehension check can run in parallel with the A/B test rather than months behind it.\n\nThat cadence is the real unlock. The reason teams ship dark patterns is rarely malice; it is that the conversion number arrives in a day and the comprehension number, under traditional research, arrives in six weeks or never. Equalise the latency and the two numbers get weighed together, which is the only durable fix.\n\n## Common mistakes\n\n- **Auditing intent instead of effect.** Workshops about whether the team meant well produce no evidence a regulator would credit.\n- **Relying on a heuristic checklist alone.** Pattern lists are a useful screen, but whether a specific implementation impairs decision-making is an empirical question about your users. Use the list to generate hypotheses, then test them.\n- **Testing only the variant that won.** The comparison you need is against a neutral control, not against the previous champion, which may itself be manipulative.\n- **Confusing friction with honesty.** Some friction is legitimate and required. See [age assurance research](/docs/age-assurance-user-research) for a case where friction is the legal obligation, and [the friction log method](/docs/friction-log) for separating the two.\n- **Stopping at the sign-up flow.** Article 25(3)(c) specifically names making termination harder than subscription. Test the cancellation path with the same rigour, and read [onboarding drop-off diagnosis](/docs/onboarding-drop-off-research-guide) for the entry side.\n\n## Frequently asked questions\n\n### What legally counts as a dark pattern?\n\nDefinitions vary by regime but converge on effect. The DSA prohibits interface design that deceives or manipulates users or otherwise materially distorts or impairs their ability to make free and informed decisions. California requires symmetry in choice and treats consent obtained through a dark pattern as invalid. The FTC groups them into inducing false beliefs, hiding material information, causing unauthorized charges, and subverting privacy choices. None of them require proof that anyone intended harm.\n\n### Can a pattern be a dark pattern if every fact was disclosed?\n\nYes, and this is the most common misunderstanding. In the LendingClub matter the fees were disclosed; they were placed behind tooltips users were unlikely to click and inside an un-bolded paragraph between bolded ones. Disclosure that predictably fails to reach the user is the mechanism, not a defence against it.\n\n### How is this different from ordinary usability testing?\n\nUsability testing asks whether people can complete a task. Dark pattern testing asks whether the people who completed it understood what they did. A flow can score extremely well on task success and completion time precisely because it is manipulative. Run both, and read [usability metrics](/docs/usability-metrics-guide) for why success rate alone is insufficient.\n\n### Do we need a control group?\n\nFor any claim that a design caused a false belief, yes. Without a control you cannot separate the effect of the design from background misunderstanding, guessing, or the wording of your own question. A control also neutralises the objection that a question was leading, because both cells answered the identical question.\n\n### How large a comprehension drop should worry us?\n\nTreat any statistically reliable drop in accurate belief that accompanies a conversion lift as a finding worth escalating, rather than setting a tolerance threshold. Regulators do not apply a percentage cut-off. Also avoid the opposite error of reading a non-significant difference as proof of no harm, which is covered in [equivalence testing](/docs/equivalence-testing-no-difference).\n\n### Does this apply to B2B products?\n\nThe consumer protection statutes are aimed at consumers, but the DSA applies to online platforms broadly, and enterprise buyers are covered by contract doctrines with their own standards for notice and assent, discussed in [clickwrap vs browsewrap](/docs/clickwrap-vs-browsewrap-assent-research). The measurement method is identical regardless.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types and when each one fits\n- [Clickwrap vs Browsewrap](/docs/clickwrap-vs-browsewrap-assent-research) - testing whether users understood what they agreed to\n- [Age Assurance Research](/docs/age-assurance-user-research) - where friction is the legal requirement\n- [Survey Evidence in Court](/docs/survey-evidence-court-daubert) - control-group design that withstands challenge\n- [Reference Prices and Drip Pricing](/docs/price-display-comprehension-research) - price display comprehension\n- [Usability Metrics](/docs/usability-metrics-guide) - why task success alone hides manipulation","category":"Research Methods","lastModified":"2026-08-10T03:21:10.624704+00:00","metaTitle":"Dark Patterns Testing: Measure Deceptive Design Before a Regulator Does","metaDescription":"Dark pattern law judges effect, not intent. Learn the conversion-versus-comprehension test that identifies manipulation, plus what the FTC report, DSA Article 25 and California symmetry rules actually require.","keywords":["dark patterns testing","deceptive design research","choice architecture testing","symmetry in choice","dsa article 25","dark pattern user research","consent banner testing"],"aiSummary":"Dark pattern rules judge effect rather than intent, which makes deceptive design a measurement problem. This guide gives an operational test: a variant that lifts conversion while lowering accurate user belief is a dark pattern by regulators own logic. It covers the FTC 2022 staff report four families, DSA Article 25 (including the widely misreported paragraph 2 carve-out), California symmetry in choice under 11 CCR 7004, control-cell study design, and magnitude evidence from Luguri and Strahilevitz (11.3 to 37.2 percent) and Mathur et al (1,818 instances across 11,000 sites).","aiPrerequisites":["Familiarity with A/B testing and conversion metrics","Access to the live flow you want to evaluate"],"aiLearningOutcomes":["Apply the conversion-versus-comprehension matrix to classify a design change","Distinguish DSA Article 25 from UCPD and GDPR coverage","Audit symmetry in choice without recruiting participants","Design a control cell that neutralises leading-question objections","Avoid measuring perceived fairness instead of belief accuracy"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min"}],"pagination":{"total":1,"returned":1,"offset":0}}