Age Assurance and Age Gates: Testing Accuracy, Friction, and Abandonment
Age assurance is the rare feature where friction is the legal deliverable. That inverts the usual research question, and it creates a blind spot analytics physically cannot see: the adult who was wrongly blocked and left.
Age assurance is the one place in your product where friction is the requirement, not the defect. Every other flow gets optimised toward fewer steps; an age check exists precisely to stop some people getting through. That inversion breaks the usual instincts, and it creates a measurement problem most teams get wrong: your analytics can tell you who completed a check, but they cannot tell you who was wrongly turned away. An adult misjudged as a child and a child correctly blocked look identical in the logs. Only research reaches the first group. This guide covers what to measure, using the strongest public evidence base that now exists, and how to run it quickly with conversational research.
The scale this has reached
Ofcom published its statutory report on the use of age assurance on 15 July 2026, prepared under section 157 of the Online Safety Act 2023, covering the first six months of the protection of children duties. The numbers reset the baseline for anyone still treating age checks as a niche concern.
Between July and December 2025, over 69 million age checks were completed across a sample of just 32 services operating in the UK, a 23-fold increase on the previous six months. Ofcom notes the true UK total is likely materially higher. Facial age estimation dominated, making up 68 percent of all age check completions across analysed pornography services between July and November 2025; a parallel survey of users found 58 percent had used facial age estimation and 28 percent photo ID matching.
Ofcom's framework defines highly effective age assurance against four criteria: it must be technically accurate, robust, reliable and fair. The guidance names seven methods capable of meeting that bar and explicitly rules out self-declaration of age as incapable of it. Two implementation steps recur throughout the report as things services skipped: liveness detection, and a challenge age approach, the online analogue of the Challenge 25 policy used in retail, where anyone estimated below a buffer age is routed to a second, stronger check.
Cost is not the obstacle people assume. Ofcom found per-check costs generally sit toward the lower end of a GBP 0.05 to GBP 0.30 range.
| Method | Share of completions | What it collects | The question to test |
|---|---|---|---|
| Facial age estimation | 68 percent | A selfie, processed to an age band | Accuracy in the buffer band; failure under poor light and older cameras |
| Photo ID matching | Second most used in survey data (28 percent) | A government document plus a selfie | Willingness to submit ID at all, and what users believe happens to it |
| Digital identity services | 2 percent | A reusable verified credential | Why adoption stays low despite reuse across services; flow length |
| Email-based estimation | Popular among completions | An email address, inferred from its history | Perceived intrusiveness versus accuracy for light users |
| Mobile network operator check | Popular among completions | Carrier-held age status | Coverage gaps for pay-as-you-go and shared accounts |
| Self-declaration | Not permitted | A tick box | Not capable of being highly effective under Ofcom guidance |
The finding every product team should read twice
Buried in Ofcom's analysis is a single sentence that justifies this entire article. Digital identity services accounted for only 2 percent of completed age checks. One pornography service provider removed its digital identity option after a week, because in its experience the user flow was long and it saw evidence of users abandoning the process.
That company discovered an abandonment problem by shipping it to production and watching a week of real traffic drain away. A comprehension and friction study on the same flow would have surfaced the same finding before launch, at a fraction of the cost, and would additionally have explained why people abandoned, which the traffic data never does.
The wider pattern in the report is the same at ecosystem scale. Many porn sites that introduced age checks experienced sharp declines in traffic while some sites without age assurance gained. Ofcom found that almost half, 47 percent, of the pornography services children visited had no age checks at all, and that in May 2026 around 33 percent of first-page Google results for general porn queries pointed to sites without protections. Only about 5 percent of children reported using a VPN. The dominant evasion route is not circumvention technology; it is walking to a service that did not implement a check. Friction that is heavier than a competitor's does not just cost you conversions, it moves the user somewhere with no protection at all.
What to measure
Map your study directly onto Ofcom's four criteria, and add a fifth that the evidence shows drives behaviour: acceptability. Koji supports all six structured question types in one session, so a single interview can carry the countable measures and the explanations together.
| Criterion | What to measure | Question types |
|---|---|---|
| Technically accurate | Estimated age versus true age, especially near the threshold and buffer band | scale for confidence, yes_no on outcome correctness |
| Robust | Completion under real conditions: poor light, old phones, shared devices, low bandwidth | open_ended with AI follow-up on where it broke |
| Reliable | Repeatability. Same person, same method, second attempt, same result | yes_no plus open_ended on retry experience |
| Fair | Whether accuracy and completion differ by skin tone, age band, disability or device tier | single_choice on demographics, compared across segments |
| Acceptable | Willingness to submit the data at all, and what people believe happens to it | ranking across methods, multiple_choice on perceived data use |
Five design rules matter more here than in ordinary usability work:
- Recruit the people who failed, not just the people who finished. This is the whole ballgame. Completion analytics are a census of successes, the classic shape described in survivorship bias in customer research. Your recruitment has to reach abandoners deliberately, through in-product intercepts or a panel screened on having encountered an age check, or the study will simply confirm that people who succeeded succeeded.
- Test the buffer band, not the average user. A model that is accurate on a 40-year-old and a 12-year-old tells you nothing. All the risk sits within a few years of the threshold, which is exactly why a challenge age approach exists. Oversample that band deliberately.
- Test fairness as a first-class outcome. Ofcom looked specifically for demographic testing when assessing whether services showed regard for fairness. Report accuracy by segment, and treat a gap as a finding rather than noise. See accessibility compliance research for the adjacent obligations.
- Ask what people think happens to their ID. Stakeholders told Ofcom that high-profile breaches have raised privacy concerns, with smaller platforms seeing higher drop-off due to lower user trust. One cited incident involved a third-party service provider where roughly 70,000 users may have had government ID photos exposed. Perceived data handling is a live driver of abandonment, so measure belief about it, not just satisfaction.
- Do not ask minors to self-report age in your study. Research involving children carries its own consent architecture. See user research with children and teens and research ethics and informed consent.
Reading the results
Three interpretation traps recur.
Completion rate is not effectiveness. A check that everyone passes quickly may simply be weak. Ofcom's central criticism of some large services was that proprietary age inference models, which analyse behaviour to guess whether a user is a child, may have failed to identify large numbers of children on the platform. High throughput and low protection are perfectly compatible.
A false rejection and a correct block are the same event in your data. Only the user knows which happened. This is why the qualitative layer is not optional: you need the adult who was refused to tell you they were refused. Build the study so that outcome and ground truth are both captured.
Do not read a small measured difference as no difference. If your new method appears no worse than the old one on a modest sample, that is usually an underpowered result rather than evidence of equivalence. Work out the detectable difference in advance using statistical power and minimum detectable effect, and if the claim you want to make is genuinely one of parity, use the method in equivalence testing rather than a non-significant p-value.
Segment your reporting by device and method. Averaging a facial estimation flow on a new phone with the same flow on a five-year-old handset in poor light produces a number describing nobody, the same failure mode covered in sampling bias.
How Koji fits
The awkward part of this research is that you need people to describe a failure they experienced, often days ago, on a topic they may find embarrassing or intrusive. A static questionnaire in Typeform or SurveyMonkey gets you a completion checkbox and a satisfaction score, neither of which explains why someone gave up at the camera step. Scheduling moderated sessions gets you the depth but not the sample size, and not on the timeline a compliance deadline runs to.
Koji runs these as AI-moderated conversational interviews, by voice or text, with no moderator to book. The AI asks what happened, hears "it kept telling me to move into better light and I gave up", and probes on its own to find out whether the participant retried, whether they went elsewhere, and what they assumed would happen to the photo. Structured questions carry the countable layer, so accuracy rates, method preference rankings and demographic breakdowns arrive as distributions alongside the verbatim accounts. Reports generate automatically, which matters when the alternative is discovering your abandonment rate the way that provider in the Ofcom report did, a week into production.
Common mistakes
- Optimising the check for speed. Speed is not the goal, correct classification is. A frictionless check that self-declaration would have passed is not highly effective under any regulator's definition.
- Treating age assurance as identity verification. They are different products with different data footprints. Estimation methods that return only an age band collect far less than ID matching, which matters for both abandonment and breach exposure.
- Testing only your own flow. Because users migrate to whichever service asks less, a competitive benchmark is part of the risk picture, not a nice-to-have.
- Skipping the retry path. Reliability means the same person gets the same answer twice. Most studies never test a second attempt.
- Assuming the vendor tested fairness. Ask for segment-level accuracy evidence, and validate it on your own population. Vendor benchmarks rarely match your device mix or user base.
Frequently asked questions
What counts as highly effective age assurance?
Under Ofcom's guidance, a method must be technically accurate, robust, reliable and fair. Ofcom lists seven methods capable of meeting the standard, including facial age estimation, photo ID matching, digital identity services, open banking, mobile network operator checks and email-based estimation. Self-declaration of age is expressly identified as not capable of being highly effective.
What is a challenge age?
It is the online version of the Challenge 25 policy familiar from retail. Rather than checking against the legal threshold directly, the system sets a higher buffer age; anyone estimated below that buffer is routed to a second, stronger check. It reduces false positives from estimation error near the threshold, and Ofcom flagged its absence as a shortcoming at several services.
How much does an age check cost?
Ofcom found per-check costs across analysed providers generally fall toward the lower end of a GBP 0.05 to GBP 0.30 range, with limited variation between methods. Cost is rarely the binding constraint; abandonment and accuracy usually are.
Which method should we choose?
Facial age estimation currently dominates by volume, accounting for 68 percent of completions across the pornography services Ofcom analysed, largely because it is fast and collects comparatively little. Digital identity services accounted for only 2 percent despite wide availability, with flow length cited as a reason. The right answer depends on your user base, which is exactly what a comparative study using ranking questions across methods will tell you.
Do age checks actually work?
Ofcom's assessment is mixed and worth stating honestly. It found highly effective age assurance is helping prevent children accessing pornography, and that circumvention appears low, with roughly 5 percent of children reporting VPN use. But it also found almost half of the pornography services children visited had no checks, and that search results readily surface unprotected sites. Checks work where they are deployed; the gap is coverage.
How does this differ from testing a consent banner?
Age assurance has a defined regulatory standard with named acceptable methods and a formal effectiveness bar. Consent design is judged instead on whether it deceives or impairs free choice, covered in dark patterns testing. The research methods overlap heavily, but the pass condition does not: here you are proving a check works, not proving a choice was free.
Related Resources
- Structured Questions Guide - the six question types and when to use each
- Survivorship Bias in Customer Research - why completion analytics cannot see false rejections
- Dark Patterns Testing - measuring whether an interface impairs free choice
- User Research With Children and Teens - consent architecture for minors
- Accessibility Compliance Research - fairness and access obligations
- Clickwrap vs Browsewrap - testing notice and assent in the same sign-up flow
Related Articles
Accessibility Compliance Research: What WCAG, the ADA, and the European Accessibility Act Require You to Test
WCAG conformance is an audit standard, not proof your product works for disabled users. Here is what the EAA, ADA Title II, and Section 504 actually demand in 2026 — and how to run the user research that closes the gap.
Clickwrap vs Browsewrap: How to Research Whether Users Actually Agreed to Your Terms
Courts decide whether your terms are enforceable by asking what a reasonably prudent Internet user would have seen and understood. That is an empirical question. Here is how to answer it with evidence instead of opinion.
Dark Patterns: How to Test a Flow for Deceptive Design Before a Regulator Does
Dark pattern rules judge effect, not intent. That makes deceptive design a measurement problem, and the measurement is the one thing conversion testing never captures: what the user actually believed.
Research Ethics and Informed Consent: A Practical Guide for UX Teams
A practical guide to ethical UX research — covering the Belmont Report's three principles, GDPR informed consent requirements, how to handle AI tools responsibly, and how to build ethical maturity in your research practice.
Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)
Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Universe: How to Define Who Counts Before You Collect a Single Answer (2026)
The universe is the population whose opinion is actually relevant to your claim. Get it wrong and no sample size, weighting or analysis can rescue the study. A protocol, four documented failures, and how to enforce it at the door.
Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story
Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.
User Research With Children and Teens: COPPA, Parental Consent, and Assent
Researching under-13s triggers COPPA verifiable parental consent - including a separate consent before any child data trains AI. Here is the compliance path, the parent-mediated pattern most teams should use instead, and how to design sessions that actually work with young participants.