The short answer
You can measure whether an AI assistant mentions your brand. You cannot measure how often buyers ask about you, because nobody knows the denominator. Every "AI visibility score" on the market is a sample drawn from a prompt list that you or your vendor invented, run on a schedule, from a clean session, in one language, with no conversation history. That is a useful instrument. It is not a market measurement, and the difference matters more than any number it produces.
This is the first of three articles on what AI answers do to brand measurement. This one is about what you can legitimately count. The second, why more AI mentions can make your positioning worse, is about what those counts hide. The third, why AI answers have no clock, is about the timestamp that does not exist.
Why this is not a ranking problem
Search visibility had a stable shape for twenty years. There was a finite list of keywords, each with an estimated monthly volume, and ten blue links per keyword. You could compute a share. The universe was public, countable, and the same for everyone.
An AI answer breaks all three properties.
The prompt is free text. It is not drawn from a keyword list, it is conditioned on whatever the buyer said earlier in the conversation, and no vendor publishes prompt volumes the way keyword tools publish search volumes. Two buyers with the same intent can produce different brand sets by phrasing the question differently, and neither phrasing is more "correct" than the other.
The scale is not in doubt. OpenAI announced 900 million weekly active users on 27 February 2026, up from the 800 million it reported in October 2025, alongside 50 million paying subscribers. Forrester's report The State Of Business Buying, 2026, published 21 January 2026, states plainly that "genAI searches are the starting point for B2B buyers." The behaviour is real and it is at the front of the funnel. What is missing is the measuring apparatus, and pretending otherwise is how teams end up reporting a number they cannot defend.
The click evidence is unusually good, because Pew Research Center measured behaviour rather than opinion. Pew tracked 900 U.S. adults on its KnowledgePanel Digital through 68,879 unique Google searches during March 2025, of which 12,593 produced an AI summary. When an AI summary was present, users clicked a traditional search result on 8% of visits, against 15% without one. They clicked a source link inside the AI summary itself on just 1% of visits. They ended their browsing session on 26% of pages with an AI summary, against 16% without. Fifty-eight percent of respondents ran at least one search that month that produced an AI summary.
Read the 1% carefully, because it is the load-bearing number in this article. The answer that most influenced the buyer is the answer that produced no click. Your most effective appearance in AI search is, by construction, invisible in your analytics.
What an AI visibility tracker actually measures
Every tool in this category works the same way: it runs a list of prompts on a schedule and records whether you appear. The list is the product. And the list is a hypothesis about what buyers ask, written by people who are not the buyers.
| What the tracker does | What the buyer does |
|---|---|
| Runs a fixed prompt list you or the vendor wrote | Types whatever is in their head, in their words |
| Starts from a clean session with no history | Arrives mid-conversation, often several turns in |
| Uses one phrasing per intent | Rephrases when the first answer is unsatisfying |
| Runs from one locale, one language, logged out | Is logged in, in their market, with account memory |
| Samples on a fixed cadence (daily, weekly) | Asks once, at the moment of need |
| Records the answer text | Reads it, believes part of it, and acts |
None of these gaps make the tool useless. They make it a weather station rather than a census, and the distinction is exactly the one we drew for share of search on the digital shelf. A weather station is genuinely informative about direction and change over time. It cannot tell you the size of the sky.
The practical consequence: treat your prompt list as a research instrument that needs validating, not as a fact. If you have never checked your prompt list against language real buyers use, your visibility score is measuring your own vocabulary. That check is an interview question, not a scraping problem, and we will come back to it.
The four numbers worth tracking, and the one that lies
| Metric | What it can prove | How it misleads |
|---|---|---|
| Mention rate (share of your prompt set where you appear) | Direction over time on a fixed instrument | Reads as market share; it is share of a list you wrote |
| Citation share (whose URLs are linked as sources) | Which third-party pages the model leans on | Cited sources are often not what shaped the answer text |
| Description accuracy (what the model says you do) | Whether the answer is commercially safe | Usually untracked entirely, and it is the one that costs money |
| AI referral traffic | That someone clicked through | Pew: the source link is clicked on 1% of visits, so this counts your least influential answers |
The fourth is the trap. AI referral traffic is the only metric in the list that lands in your existing analytics, so it is the one that gets reported to the board. It is also the one with the most severe selection bias in the set. A buyer who reads a complete, satisfying answer about you and never clicks appears nowhere. A buyer who reads a thin, unsatisfying answer and clicks through to check appears as a success. Optimising for AI referral traffic optimises for answers that fail to satisfy.
Attribution surveys make this worse rather than better, because buyers routinely cannot separate "I read it in an AI answer" from "I read it somewhere." If you plan to ask, ask about the belief and its source, not about the channel.
The third metric, description accuracy, is where the actual commercial risk sits, and almost nobody tracks it. That is the subject of the next article in this series.
The gap no tracker closes: what the buyer did next
A tracker can tell you that on 12 August, asked a particular question, Perplexity named you third of five. It cannot tell you any of the four things that determine whether that mattered:
- What the buyer actually asked, in their own words, including the two turns before the one that named you.
- Whether they believed the answer, or opened three tabs to verify it.
- What in the answer made them shortlist you, or drop you.
- What they thought your product was, based only on what the machine said.
Forrester found the same behaviour from the buyer side, and its language is worth quoting because it comes from an analyst firm rather than a vendor: AI search tools, "also known as answer engines, offer speed and efficiency, but they often deliver incomplete or unreliable information, creating mistrust. Buyers compensate for this by seeking validation from trusted sources, emphasizing the value of human contact in the buying process."
That sentence should reorganise your measurement plan. If buyers are compensating for unreliable answers by seeking validation elsewhere, then the AI answer is not the end of the journey and your mention rate is not the outcome variable. The outcome variable is what belief the buyer carried into the validation step. Forrester's data on how much validation now happens is striking on its own: the typical buying decision involves 13 internal stakeholders and nine external influencers, procurement professionals are decision-makers in 53% of business buying cycles, and more than 60% of buyers now use a trial, rising to 78% for purchases of $10 million or more.
Barbara Winters, vice president and principal analyst at Forrester, put the implication for vendors directly: providers must "ensure that their claims can be validated through trusted external voices."
How to measure the part that matters
You cannot interview the model. You can interview the person who wrote the prompt, read the answer, and decided what to do with it. Those people are reachable, and the four variables above live only in their heads.
A tight study looks like this, and it is small: 15 to 25 recent buyers and near-misses, which is a week of fieldwork rather than a quarter.
- Ask them to reconstruct the first question they asked an AI assistant about your category, in their own words. This validates or destroys your prompt list, and it is the single highest-value question in the study. Use an
open_endedquestion so they produce their phrasing, not yours. - Ask what they came away believing your product does, before you tell them anything. Again
open_ended, asked first, because any description you offer contaminates it. - Ask which of a list of sources they consulted after the AI answer, using
multiple_choice, then arankingquestion ordering those sources by influence. This measures the validation step Forrester identified. - Ask, on a
scale, how much they trusted the AI answer at the time. - Ask a
yes_noon whether they verified any specific claim the assistant made about you. A high "no" rate means the description is being taken at face value, which raises the stakes on accuracy enormously. - Ask a
single_choiceon which vendor the assistant appeared to recommend most strongly, then follow up on why they think so.
That instrument produces something no tracker can: the distribution of beliefs in the actual buyer population, with the prompt language attached. Koji supports all six of these question types natively in a single study, which matters here because you need the open-ended reconstruction and the quantified trust and ranking data from the same respondent, in the same session, or you cannot connect them. The structured questions guide covers how the six types combine in one interview and how each is aggregated in reporting.
For the standing version of this measurement, treat it as an addition to brand tracking rather than a new discipline. Our brand tracking study guide covers the cadence and wave design, and brand perception surveys cover the unprompted-description question in detail. If you are testing whether your category framing survives contact with the machine, positioning research is the closer fit.
What to do in the next quarter
- Keep the tracker. Report it as a directional index on a fixed instrument, with the prompt list published inside the report so nobody mistakes it for market share.
- Validate the prompt list against 20 real buyer reconstructions. Expect to be wrong about a third of it; that finding alone usually justifies the study.
- Add description accuracy as a tracked metric, with a named owner. It is the one with money attached.
- Stop reporting AI referral traffic as a visibility KPI. Report it as what it is: a count of your least satisfying answers.
- Put one question about AI-sourced beliefs into your existing win/loss interviews. The win/loss analysis guide covers where it fits without lengthening the call.
Where Koji fits
Koji runs AI-moderated voice and text interviews with your real buyers, so the four variables a tracker cannot see get collected at the scale of a survey and the depth of an interview. You write the brief; the AI interviewer runs every conversation, probes the reconstruction ("what did you type first?"), and holds the same structure across all 25 respondents, which is precisely where human moderators drift. Thematic analysis and a shareable report are generated automatically, so a study fielded on Monday is a readout on Thursday rather than a transcript pile in three weeks.
Two Koji specifics matter for this particular study. Interviews mix all six structured question types with open conversation, so the unprompted description and the trust scale come from the same respondent. And there is no moderator to lead the witness on the one question where leading is fatal: what the buyer thinks you do. Because the interviewer is not a person with a stake in the answer, the unprompted description stays unprompted.
If you want to know what AI is telling your market about you, the tracker is step one and it is cheap. Step two is asking the humans, and that is the step almost nobody has taken yet.
Start a study with Koji and get your first AI-sourced-belief readout inside a week.
Related reading
- Share of Search and the Digital Shelf (2026)
- When AI Describes Your Product Wrong (2026)
- AI Answers Have No Clock (2026)
- Agentic Commerce Research (2026)
- G2 vs Capterra vs TrustRadius (2026)
Frequently Asked Questions
What is AI search visibility?
AI search visibility is how often, and how favourably, an AI assistant such as ChatGPT, Gemini, Copilot or Perplexity mentions your brand when someone asks a question in your category. Unlike search rankings, there is no public list of prompts with volumes attached, so visibility is always measured against a sample of prompts someone chose rather than against the full universe of what buyers ask.
How is AI visibility different from SEO?
SEO measures position against a finite, public keyword list with estimated volumes, and the same results page is broadly shared across users. AI answers are free-text, conditioned on conversation history and account context, and are synthesised per conversation, so two buyers with identical intent can receive different brand sets. There is no shared page and no fixed number of positions, which means there is no true denominator to compute a share against.
Is AI referral traffic a good measure of AI visibility?
No, and it is systematically biased. Pew Research Center found that users clicked a source link inside a Google AI summary on just 1% of visits to pages containing one. The answers that fully satisfy a buyer generate no click at all, so referral traffic disproportionately counts your thin, unsatisfying answers. It is worth tracking as a floor, never as the headline visibility metric.
How many people actually use AI assistants for buying research?
OpenAI reported 900 million weekly active users for ChatGPT in February 2026, up from 800 million in October 2025. On the B2B side, Forrester's The State Of Business Buying, 2026 found that genAI searches are now the starting point for business buyers, though Forrester also notes buyers distrust answer engines enough to seek validation from other sources afterwards.
Can I just ask buyers where they heard about us?
Not reliably. Buyers routinely cannot distinguish an AI answer from a search result or an article they half-remember, so channel-attribution questions produce noisy data. Ask instead about the belief and its content: what they came away thinking your product does, what they asked first in their own words, and whether they verified any specific claim. Those answers are recoverable and far more actionable.
How large does an AI-visibility interview study need to be?
Fifteen to twenty-five recent buyers and near-misses is usually sufficient to validate a prompt list and surface the dominant misconceptions, because you are looking for recurring beliefs rather than population estimates. If you want to quantify how widespread a specific misconception is, scale to 100 or more and combine open-ended reconstruction with scale and single-choice items in the same interview.