Back to blog
Research

When AI Describes Your Product Wrong (2026): Why More Mentions Can Make Your Positioning Worse

Visibility and description accuracy are independent variables. If the model has settled on the wrong description of your product, every extra mention distributes it further. Here is how to find out what it is saying.

Koji

Koji Team

Research · · 11 min read

The short answer

An AI answer is not a retrieved page. It is a paraphrase assembled from many sources, most of which you did not write. That means visibility and accuracy are two independent variables, and almost every team is optimising only the first. If the model has settled on a wrong description of your product, then every additional mention distributes that wrong description more widely. For a brand being described incorrectly, more AI visibility is a larger liability, not a smaller one.

This is the second of three articles on AI answers and brand measurement. The first covered how to measure AI search visibility honestly. The third covers why AI answers carry no timestamp.

What an AI answer actually is

The mental model most marketers carry over from search is wrong in one specific way, and the error is consequential.

A search result is a pointer. Google decides which of your pages to show; the words the user reads are your words. If the description is wrong, you can edit the page.

An AI answer is a synthesis. The model produces new prose about you, drawing on training weights, retrieved documents, and whatever third-party pages the retrieval layer happened to surface. Nobody wrote the sentence the buyer reads. It did not exist before the question was asked, and it will be phrased differently for the next person.

This has an uncomfortable implication that follows directly and that most AI-visibility strategy ignores: your description in AI answers is a lagging index of what other people have written about you, not of what your product does. Your own site is one input among many, and it is the input the system is least likely to treat as a neutral account, because it is obviously promotional. Review sites, forum threads, comparison roundups, news coverage and competitors' comparison pages all get a vote, and in aggregate they outvote you.

The evidence on how often the description is wrong

The best public evidence is not from a marketing vendor. In October 2025 the European Broadcasting Union and the BBC published News Integrity in AI Assistants, a study involving 22 public service media organisations across 18 countries and 14 languages, in which professional journalists evaluated AI responses against 3,113 questions.

The headline findings:

  • 45% of all AI responses contained at least one significant issue.
  • Counting lower-severity problems as well, 81% of responses had an issue of some form.
  • Sourcing was the single biggest cause of significant issues, at 31% of all responses. That category covers "information in the response not supported by the cited source, providing no sources at all, or making incorrect or unverifiable sourcing claims."
  • Accuracy problems affected 20% of responses, and the assistants were closely grouped, all between 18% and 22%.
  • Sourcing performance varied enormously by assistant. Gemini had significant sourcing issues in 72% of responses, three times ChatGPT's 24%, with Perplexity and Copilot both at 15%.

One line in the EBU report ports to brands so exactly it is worth quoting in full: "Of particular concern for publishers are sourcing errors that misrepresent them, such as when a response misattributes an incorrect claim to them."

Substitute "brands" for "publishers" and that is the risk in a sentence. The failure mode is not only that the model says something false in general. It is that the model attaches a claim to you that you never made and cannot find the origin of.

Forrester's The State Of Business Buying, 2026 corroborates this from the buyer's side, and it is notable that an analyst firm rather than a critic is saying it. AI search tools, Forrester writes, "also known as answer engines, offer speed and efficiency, but they often deliver incomplete or unreliable information, creating mistrust. Buyers compensate for this by seeking validation from trusted sources, emphasizing the value of human contact in the buying process."

To be clear about what this evidence is and is not: the EBU study evaluated news questions, not product questions. Nobody has published an equivalent audit of brand descriptions at that rigour. The mechanism, though, is not news-specific. Synthesis from mixed-quality third-party sources with imperfect attribution is how these systems answer every question, including "is Acme any good for a team of 50?"

Why volume and accuracy move independently

This is the crux, and it is where most AI-visibility programmes have a logical hole.

What you optimiseWhat it changesWhat it does not change
More third-party mentions and citationsHow often you appear in answersWhich description the model has settled on
Better structured data on your siteHow parseable your own pages areHow much weight your pages get against third-party accounts
More comparison-page placementsWhich consideration set you appear inWhether the framing in those pages is yours or a competitor's
Higher review volumePerceived credibility signalsWhether reviews are paraphrased into the claim you want

Nothing in the left-hand column controls the description. The two variables are genuinely orthogonal, which is why a brand can run a successful AEO programme, watch its mention rate climb every month, and end the year with worse-qualified pipeline than it started with. The mentions went up. The sentence the buyer read stayed wrong, and more buyers read it.

This is not the same question as whether third-party reviews are honest. That question, including what the FTC's Consumer Review Rule now makes illegal, is covered in our analysis of what software review sites can and cannot tell you. The problem here survives perfect honesty upstream: entirely accurate reviews, written in good faith, get compressed into a one-paragraph description that no reviewer wrote and nobody checked.

The four ways a description goes wrong

Failure modeWhat it looks likeWhat it costs
Category misassignmentYou sell a research platform; the model files you under "survey tools"You are compared on the wrong axis, against the wrong rivals, and lose on criteria that do not apply
Attribute errorWrong pricing, a free tier that no longer exists, "enterprise only" when you self-serveBuyers disqualify you before contact, and you never learn it happened
Misattributed claimA limitation from a competitor's comparison page, or a criticism from one review, stated as fact about youYou spend the first call defending something you never said
Framing inheritanceThe model adopts a reviewer's or competitor's framing as neutral descriptionYour positioning is written by the party with the most published content, not the best product

The first and fourth are the expensive ones, and they are the ones that never show up in a dashboard. A tracker records that you were mentioned. It scores that as a win. It does not notice that you were mentioned as a cheaper alternative to a category you deliberately exited two years ago.

Category misassignment is particularly costly because it is self-reinforcing. Once the model files you alongside a set of competitors, buyers ask follow-up questions in that frame, comparison content gets written in that frame, and the corpus that trains the next answer contains more of that frame.

Route around it: you cannot edit the model, but you can find the belief

There is no correction desk. No major assistant offers brands a mechanism to file a factual dispute about how they are described and have it fixed. Publishing a correction on your own site enters your correction into the same contested corpus as everything else, competing with third-party accounts, on a timeline nobody controls.

So the tractable question is not "how do I fix the model?" It is "what do buyers currently believe, how wrong is it, and which wrong belief is costing the most?" That question is fully answerable, and the answer tells you where to aim your corpus effort.

The instrument is an unprompted-description study, and its design has one rule that decides whether the data is worth anything: ask what they believe before you tell them anything. Any positioning statement you show first contaminates every answer after it.

  • Open with open_ended: "In your own words, what does this product do, and who is it for?" Nothing before it. No logo, no tagline, no category label.
  • Follow with single_choice on category placement, offering your intended category alongside the three you keep getting confused with. This quantifies misassignment.
  • Use multiple_choice for specific attribute beliefs: pricing model, deployment, company size served, whether there is a free tier. This catches attribute errors cheaply.
  • Use ranking to order the vendors they consider comparable. This surfaces the consideration set the machine built, which is often not the one your competitive deck assumes.
  • Use a scale on confidence in their own description, which separates a firmly held wrong belief from an idle guess. These need different responses.
  • Use yes_no on whether they checked any claim against your own site. A high "no" rate means the third-party description is the only one operating.

Run the same instrument against buyers who chose you and buyers who did not. The gap between the two unprompted descriptions is your positioning problem, stated in customers' words and quantified.

Our brand perception survey guide covers the unprompted-description technique in depth, and positioning research covers validating category framing with real buyers. For the consideration-set half, competitive intelligence interviews covers what customers know about your rivals, and the competitive research guide covers the wider market picture.

One warning specific to 2026. It is tempting to study this with synthetic respondents, since the subject is AI. Do not. When the thing you are measuring is a distortion introduced by a language model, asking a language model to simulate the buyer measures the distortion against itself. Our guide to synthetic users sets out where AI personas are and are not defensible, and this is squarely on the "not" side. The same caution applies to respondent quality generally, covered in survey fraud and respondent quality.

What to do in the next quarter

  1. Capture the actual description. Run your top 20 buyer questions across four assistants and save the prose, not just whether you were mentioned. This takes an afternoon.
  2. Score each answer on the four failure modes above. Most teams find at least one systematic category or attribute error immediately.
  3. Field the unprompted-description study with 20 to 30 buyers, split between won and lost.
  4. Compare the two. Where the buyers' wrong belief matches the model's wrong description, you have a causal story worth acting on and a specific claim to go correct in the corpus.
  5. Add description accuracy to your reporting next to mention rate. Reporting volume alone is what created this blind spot.

Where Koji fits

The study above is easy to design and historically painful to run, because unprompted description requires a moderator who asks the question cold and does not react to the answer. Human moderators leak. They nod at the right answer, follow up harder on the wrong one, and by respondent twelve they are subtly coaching.

Koji's AI interviewer does not. It asks all 30 buyers the same opening question in the same neutral way, probes for specifics without signalling which answer is wanted, and never reveals your positioning because it was not asked to. That is the difference between measuring what buyers believe and measuring what buyers will agree to.

Because Koji supports all six structured question types inside a conversational interview, the unprompted open_ended description, the single_choice category placement, the ranking of competitors and the scale on confidence come from the same respondent in one sitting, which is what lets you say "the buyers who put us in the wrong category also rank us against the wrong three vendors." Thematic analysis runs automatically across every transcript, so the recurring misconceptions surface as themes with quotes attached rather than as a pile of recordings. See the structured questions guide for how the types combine and aggregate.

You cannot make the model say the right thing. You can find out exactly what it is saying wrong, to whom, and what that costs, and then go change the sources it reads.

Run an unprompted-description study with Koji and see your positioning in your buyers' words.

Frequently Asked Questions

Why does AI get my brand wrong?

Because an AI answer is synthesised, not retrieved. The model composes new prose about you from training data, retrieved pages and third-party accounts, weighting sources it treats as independent above your own marketing site. If the third-party corpus is thin, contradictory or out of date, the answer can be fluent, confident and wrong at the same time. Nobody wrote the sentence the buyer reads, so there is no page to go and edit.

Can I ask OpenAI or Google to correct information about my company?

There is no established correction channel for brands at any major assistant comparable to a press correction or a search removal request. The practical levers are indirect: improve the accuracy and clarity of your own pages, and work on the third-party sources the models draw from. Both operate on a timeline you do not control, which is why measuring what buyers currently believe matters more than waiting for a fix.

How often do AI assistants make significant errors?

In the EBU and BBC study News Integrity in AI Assistants, published in October 2025 and covering 3,113 questions across 22 public service media organisations, 45% of AI responses contained at least one significant issue and 81% had an issue of some form. Sourcing was the largest single category at 31%, and accuracy problems affected 20%. The study covered news questions rather than product questions, so treat it as evidence of mechanism rather than a brand-specific error rate.

Does more AI visibility always help my brand?

No. Visibility and description accuracy are independent variables. If the model has settled on a wrong category or a wrong attribute for your product, increasing the number of answers you appear in increases the number of buyers who receive the wrong description. Track what the answers say about you, not only how often they mention you.

What is category misassignment and why does it matter?

Category misassignment is when an AI answer files your product under the wrong category, so buyers compare you against the wrong competitors on criteria that do not apply to what you actually sell. It is expensive because it is self-reinforcing: buyers ask follow-up questions in the wrong frame, more content gets written in that frame, and future answers inherit it. It rarely appears in visibility dashboards, because a mention in the wrong category still counts as a mention.

Should I use synthetic respondents to study how AI describes my brand?

No. When the phenomenon you are measuring is a distortion introduced by a language model, using a language model to simulate your buyers measures that distortion against itself and will tend to confirm it. This is one of the clearest cases for real respondents. The point of the study is to find the gap between what the machine says and what humans actually took away, and only humans can supply the second half.

Run your first AI-moderated study in 10 minutes

10 free credits on signup. No credit card required.

GDPR compliantEU or US data residencyNo AI training on your data
Koji

Koji Team

Research

Share this article

Keep reading