Back to docs
Research Methods

Contestability and Redress: How to Design and Research the Appeal Flow for AI Decisions (2026)

EU AI Act Article 86 became applicable on 2 August 2026. Learn what GDPR Article 22 and the CJEU already require of an appeal flow, why appeal rates are a misleading metric, and how to research whether users can actually contest a decision.

If your product makes automated decisions about people - approving, ranking, scoring, flagging, pricing, or declining - then the appeal path is now a regulated part of the product, not a support process you can improvise. EU AI Act Article 86 became applicable on 2 August 2026. It gives affected people a right to an explanation of individual decision-making, on demand, from the deployer.

Most teams have a lot of research on the decision and none at all on the appeal. That is exactly backwards in terms of risk: the decision affects everyone a little, and the appeal affects a small number of people enormously. This guide covers what the law now requires, why the obvious metric is misleading, and how to research a redress flow that works.

What the law actually requires

Three instruments stack, and they say different things.

GDPR Article 22 gives a data subject the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects or similarly significantly affects them. Where such processing is permitted - by contractual necessity, by Union or Member State law, or by explicit consent - Article 22(3) requires safeguards, and it names three specific ones: the right to obtain human intervention, to express his or her point of view, and to contest the decision.

Those three are not one right written three ways. A flow that lets a user submit a complaint into a queue satisfies none of them properly. "Human intervention" means a person with the authority and the information to change the outcome. "Express a point of view" means the user can introduce facts the system did not have. "Contest" means the outcome is genuinely reversible.

The CJEU has twice narrowed the room to manoeuvre. In SCHUFA Holding (C-634/21, 7 December 2023) the Court held that a credit reference agency producing a probability score is itself engaged in automated individual decision-making under Article 22, where lenders draw strongly on that score - so the Article 22 obligations fall on the scoring party, not only on the party that formally issues the decision. If you supply scores that other people act on, this is about you.

In Dun and Bradstreet Austria (C-203/22, 27 February 2025) the Court addressed what "meaningful information about the logic involved" actually means. A data subject may require the controller to explain, in a concise, transparent, intelligible and easily accessible form, the procedure and principles actually applied to their personal data to reach the specific result. The Court also held that where a controller claims trade secret protection, it must disclose the protected information to the competent supervisory authority or court for assessment - trade secrecy is a reason to route the disclosure, not a reason to refuse the explanation.

EU AI Act Article 86 adds an explanation right specific to high-risk systems, applicable since 2 August 2026:

"Any affected person subject to a decision which is taken by the deployer on the basis of the output from a high-risk AI system listed in Annex III, with the exception of systems listed under point 2 thereof, and which produces legal effects or similarly significantly affects that person in a way that they consider to have an adverse impact on their health, safety or fundamental rights shall have the right to obtain from the deployer clear and meaningful explanations of the role of the AI system in the decision-making procedure and the main elements of the decision taken."

Two design consequences are easy to miss. The trigger is the affected person considering the impact adverse - it is subjective, not a threshold you get to define. And the right is reactive and on demand, which means you must be able to reconstruct, months later, what the system contributed to a specific decision. That is a logging and retention requirement disguised as a transparency right, and it should be scoped alongside the rest of your EU AI Act obligations.

Why this is worth more than compliance effort

Two state-scale failures show what happens when automated decisions have no working appeal path.

In the Netherlands, the childcare benefits scandal - the toeslagenaffaire - saw roughly 26,000 families wrongly accused of benefit fraud between 2005 and 2019, with a risk model that used nationality as an input in breach of non-discrimination law. Families faced demands to repay sums they did not owe with, in practice, no effective route to challenge them. The government resigned in January 2021. By 2024 more than 33,000 people had been formally acknowledged as victims.

In Australia, the Robodebt scheme used automated income averaging to raise welfare debts. In May 2020 the government conceded that 470,000 debts raised under the scheme were insufficient under the law; the Federal Court recorded that roughly 751 million Australian dollars had been recovered from about 381,000 people. A Royal Commission reported in 2023. The original class action settled in 2021 with 112 million dollars of compensation, and a further settlement of 548.5 million dollars was agreed in July 2025 and approved by the Federal Court on 23 June 2026. Including roughly 1.763 billion dollars of cancelled debts, more than 2.4 billion dollars has been returned.

Neither failure was primarily a model-accuracy failure. Both were contestability failures. The systems were wrong at a rate that a functioning appeal process would have surfaced within months, and in both cases the appeal process was the thing that did not work.

Contestability is a design property, not a support queue

The most useful academic framing comes from Kars Alfrink, Ianus Keller, Gerd Kortuem and Neelke Doorn, whose paper "Contestable AI by Design: Towards a Framework" appeared in Minds and Machines (33(4), 613-639). Their premise:

"As the use of AI systems continues to increase, so do concerns over their lack of fairness, legitimacy and accountability. Such harmful automated decision-making can be guarded against by ensuring AI systems are contestable by design: responsive to human intervention throughout the system lifecycle."

The phrase doing the work is throughout the system lifecycle. Contestability is not a button added at the end. If a challenge cannot reach the people who set the thresholds, retrain the model, or change the policy, then the appeal flow is theatre - it processes individual complaints without ever changing what generates them.

The appeal funnel, and why it collapses

Here is the arithmetic that should govern how you resource this. For a wrongly-decided user to obtain redress, four things must all happen:

  1. Discovery - they realise a decision was made and that it can be challenged.
  2. Comprehension - they understand the stated reason well enough to argue against it.
  3. Effort - they are willing and able to pay the cost of appealing.
  4. Capacity - the reviewer has the authority and information to reverse it.

These multiply. If each gate passes a generous 50%, only 6% of wrongly-decided users obtain redress (0.5 to the fourth power). The flow can look reasonable at every individual step and still fail 94% of the people it exists for.

Run it the other way to see the standard required: to get half of wrongly-decided users through, each gate must pass around 84% (0.84 to the fourth power is roughly 0.5). That is a demanding bar, and it is the number that tells you an appeal flow needs real design investment rather than a form and an inbox.

GateTypical failureWhat to measure
DiscoveryThe decision arrives with no mention that it can be challengedUnprompted awareness that an appeal exists
ComprehensionThe reason is a generic code or a policy citationCan the user restate why, and name what would change it
EffortLogin walls, document uploads, 20-day windows, phone-only routesTime to submit, drop-off per step, abandonment reasons
CapacityThe reviewer sees the same score and no new evidenceOverturn rate, and whether reviewers can see the user submission

Gate four is the quiet killer. If your human reviewer is shown the model output and nothing else, they will mostly confirm it - which is automation bias operating inside your own compliance control. A rubber-stamp review satisfies the letter of "human intervention" and none of its purpose.

Appeal rate is not a quality metric

A low appeal rate is celebrated in most organisations. It should not be, for exactly the reason a low thumbs-down rate should not be.

The number of appeals you receive is the product of the rate at which you are wrong, the rate at which users notice, and the rate at which they think appealing is worth it. It falls when your decisions improve - and it falls just as reliably when the appeal path gets harder to find, when the deadline shortens, or when the affected population learns that appealing does not work. The same decomposition applies as in in-product feedback signals, and the trap is identical.

Two metrics are far more diagnostic:

  • Overturn rate among appeals. A high overturn rate means your decisions are wrong often; it also means the appeal path works. A near-zero overturn rate is not a clean bill of health - it usually means the review is not real, or only the most hopeless cases make it through the funnel.
  • Estimated redress coverage. Of the people you decided against wrongly, what share obtained relief? You cannot measure this directly, which is precisely why you have to research it: audit a sample of negative decisions independently and see how many of the erroneous ones ever produced an appeal.

What to research, and how

The population you need is hard to reach by design. People who were declined have left, are angry, or never knew there was a decision to question. Standard recruitment misses them almost perfectly.

Interview people who did not appeal. This is the single highest-value study on this page and almost nobody runs it. Take users who received an adverse decision and never contested it, and find out which of the four gates stopped them. Non-appealers are the denominator that makes every other number interpretable.

Test comprehension, not satisfaction. Show a real decision notice and ask the user to explain, in their own words, why the decision went that way and what they would have to change. If they cannot, the notice fails the Dun and Bradstreet standard of concise, intelligible information about the procedure and principles actually applied - regardless of how clear your legal team finds it.

Time the flow end to end. Measure with a stopwatch, not an estimate. Count required documents, screens, and days.

Research the reviewer, not just the appellant. Ask reviewers what they see, what authority they have, and how often they overturn. This is where you find out whether gate four is real.

Study the loop back into the system. Ask whether anything an appeal reveals ever reaches the team that owns the model or the threshold. If nothing does, you have a complaints desk, not contestability - and the pattern belongs in your AI incident postmortem process.

Running this study with Koji

The practical obstacle to redress research is reach. Declined and lapsed users will not book a 45-minute call with the company that just turned them down, and moderated sessions on a sensitive topic are slow and expensive. Traditional survey tools can reach these people but collect only the shallow layer, and a form cannot ask the follow-up question that matters.

Koji AI-moderated interviews change the economics. The interview runs on the participant schedule, in text or voice, with no human moderator on the other side - which for a sensitive topic like an unfair decision often produces more candour than a call would. The AI probes each answer automatically, so "I did not bother appealing" becomes a real explanation rather than a dead end. You can also tune the AI interviewer tone and give it your company context, so it handles an angry participant appropriately.

A redress study using Koji six structured question types:

QuestionTypeWhat it establishes
"Tell me what happened when you got that decision."open_endedThe narrative, with automatic follow-up probing
"Did you know you could challenge it?"yes_noGate 1, discovery, cleanly measured
"How clear was the reason you were given?"scaleGate 2, comparable across notice versions
"What stopped you from appealing?"single_choiceGate 3, the reason non-appealers never give you
"Which of these did you try?"multiple_choiceActual recovery behaviour
"Rank these changes by how much they would have helped."rankingPrioritised fixes, from the affected population

Because the metrics sit in structured fields, they aggregate automatically and stay comparable release over release, while the open-ended answers carry the reasoning. Thematic analysis runs as responses arrive, so you get a live report instead of a fortnight of transcript coding. Voice interviews work well here: people explain an injustice more fully out loud than they type it.

One caution specific to this topic. Research about a decision is not part of the decision, and it must not become a condition of appealing. Keep the study separate from the redress mechanism itself, be explicit that participation does not affect the outcome, and handle the data under your normal privacy and security and GDPR commitments.

The short version

  • GDPR Article 22(3) requires human intervention, the ability to express a point of view, and genuine contestability - three distinct things.
  • SCHUFA puts those duties on whoever produces the score, not only on whoever issues the decision.
  • Dun and Bradstreet sets the explanation standard, and makes trade secrecy a routing question rather than an exemption.
  • EU AI Act Article 86 has applied since 2 August 2026 and is triggered by the affected person view of the impact.
  • Four multiplicative gates mean a flow that is 50% effective at each step reaches only 6% of the people it is for.
  • Appeal rate is not a quality metric. Overturn rate and redress coverage are.
  • The most valuable interviews are with the people who never appealed at all.

Try it yourself. Koji gives you 10 free credits at signup - enough to run a real redress study with the population your current research is missing. No seat licences and no sales call required.

Frequently asked questions

What does EU AI Act Article 86 require, and when does it apply?

It has applied since 2 August 2026. Any affected person subject to a decision taken by a deployer on the basis of output from a high-risk Annex III system (excluding point 2), which produces legal effects or similarly significantly affects them in a way they consider to have an adverse impact on their health, safety or fundamental rights, has the right to obtain clear and meaningful explanations of the role of the AI system in the decision procedure and the main elements of the decision. Two design consequences: the trigger is the affected person own view of adverse impact, not a threshold you define, and the right is reactive and on demand, so you must be able to reconstruct months later what the system contributed to a specific decision.

How is Article 86 different from GDPR Article 22?

Article 22 concerns decisions based solely on automated processing and, where such processing is permitted, requires safeguards under 22(3): the right to obtain human intervention, to express a point of view, and to contest the decision. Article 86 is narrower in scope - high-risk Annex III systems - but does not require the decision to be solely automated, and it grants an explanation right specifically. Article 86(3) applies only to the extent the right is not otherwise provided under Union law, so the two are designed to complement rather than duplicate each other.

Does the SCHUFA ruling apply to us if we only supply a score?

Very likely yes. In SCHUFA Holding (C-634/21, 7 December 2023) the CJEU held that a credit reference agency producing a probability score is itself engaged in automated individual decision-making under Article 22 where the recipient draws strongly on that score. The obligations therefore fall on the scoring party, not only on the party that formally issues the decision. If you supply scores, rankings or risk flags that customers act on, you should assume the Article 22 safeguards attach to you.

Can we refuse to explain a decision because the model is a trade secret?

No. In Dun and Bradstreet Austria (C-203/22, 27 February 2025) the CJEU held that a data subject may require the controller to explain, in a concise, transparent, intelligible and easily accessible form, the procedure and principles actually applied to their personal data to reach the specific result. Where the controller claims trade secret protection, it must disclose the protected information to the competent supervisory authority or court for assessment. Trade secrecy determines who sees the underlying detail, not whether an explanation is owed.

Why is a low appeal rate a bad thing to celebrate?

Because the count of appeals is the product of how often you are wrong, how often users notice, and how often they believe appealing is worth the effort. It falls when decisions improve, and it falls just as reliably when the appeal path becomes harder to find, the deadline shortens, or the affected population learns that appealing does not work. Use overturn rate among appeals and estimated redress coverage instead - the share of wrongly-decided people who actually obtained relief.

What is the appeal funnel and why does it collapse?

Four things must all happen for a wrongly-decided user to obtain redress: they discover the decision can be challenged, they comprehend the stated reason well enough to argue against it, they are willing and able to pay the effort cost, and the reviewer has the authority and information to reverse it. These multiply. At a generous 50% per gate only 6% of wrongly-decided users obtain redress; reaching half of them requires roughly 84% at every gate. Reviewer capacity is the quiet killer, because a reviewer shown only the model output will mostly confirm it - automation bias operating inside your own compliance control.

Who should I interview for a redress study?

The people who never appealed. They are the denominator that makes every other number interpretable, and standard recruitment misses them almost perfectly because they have left, are angry, or never realised a decision was made. Koji AI-moderated interviews reach them because the interview runs on their schedule with no human on the other side, which on a sensitive topic often produces more candour than a call. Keep the study strictly separate from the redress mechanism and state plainly that participating does not affect the outcome of any appeal.

Related Resources

Related Articles

AI Explainability Testing: How to Find Out Whether Your Explanations Actually Help Users (2026)

Explanations that users rate highly often fail to improve their decisions — and in one 3,800-person experiment, the more transparent model made people worse at catching its mistakes. This guide covers the four outcome measures that separate a useful explanation from a satisfying one, and how to test yours.

AI Governance for Customer Research: ISO 42001, the NIST AI RMF, and What Procurement Actually Asks

Security review is asking whether your AI research platform is ISO 42001 certified and NIST AI RMF aligned. Here is what each framework covers, what a certificate does and does not buy you under the EU AI Act, and how to answer.

AI Model Cards and User Disclosure: Documenting Intended Use, Limitations, and What You Tell People (2026)

A practical guide to model cards, system cards, and user-facing AI disclosure — what belongs in each section, what the EU AI Act's Article 50 has required since 2 August 2026, and how to source the Limitations section from real user research instead of guesswork.

AI Over-Reliance and Automation Bias: How to Research Whether Users Trust Your AI Too Much (2026)

Users who accept every AI suggestion are a product risk, not a success metric. How to measure over-reliance and automation bias, why self-report fails, and the study designs that produce honest reliance data.

The EU AI Act and User Research: What AI-Moderated Interviews Actually Require (2026)

AI-moderated customer interviews sit in the EU AI Act's limited-risk transparency tier, not the high-risk tier. Here is exactly what Article 50 requires from 2 August 2026, the two things that escalate a study to high-risk, and a compliance checklist you can run this week.

GDPR-Compliant AI User Research: A Practical Guide

How to run AI-moderated customer interviews under GDPR. Lawful basis, consent flows, data minimization, retention, sub-processors, and how Koji handles each requirement.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.