Back to blog
Research

Agentic Commerce Research (2026): When an AI Agent Shops for Your Customer, What Is Left to Measure?

AI-sourced traffic to retail is growing fast, but the first generation of in-chat checkout has already been pulled. Here is what the 2026 evidence actually shows, why agent-mediated buying destroys every impression-based shelf metric, and the one research move that still works.

Koji

Koji Team

Research · · 12 min read

The two preceding articles in this series assume something that is quietly becoming optional: that a human being looks at a page.

Share of search measures the results grid a shopper sees. The Buy Box measures which offer a shopper is shown. Both metrics require an impression delivered to a pair of eyes. When an AI agent does the searching, the comparing, and sometimes the buying, there is no grid, no position, and no impression - and every shelf metric you own becomes undefined rather than merely inaccurate.

That is the measurement problem. It is worth being precise about how big it is right now, because the honest answer is more interesting than the hype.

The short answer

Agentic commerce is real as a discovery channel and, as of 2026, largely unproven as a checkout channel. The first serious attempt at in-chat purchase has already been rolled back. Meanwhile AI-sourced traffic to retail is growing at rates no other channel has matched, and it converts well once it lands on your site.

The strategic implication is not "prepare for agents to buy everything." It is narrower and more urgent: an increasing share of your consideration-stage influence is happening inside a system that sends you no impression data, no query log, and no ranking report. You find out you lost only when the sale does not happen.

What the 2026 evidence actually shows

Take the growth first. Adobe Analytics reported traffic to US retail sites from generative AI sources up roughly 693% year over year across the 2025 holiday season, with AI referrals converting around 31% better than other traffic sources - a reversal from the prior year, when AI traffic lagged on conversion. Whatever else is true, people are arriving from these tools and arriving with intent.

Now take the checkout side, which is where the reporting gets much less flattering.

OpenAI launched Instant Checkout on 29 September 2025, letting people buy from US Etsy sellers inside ChatGPT via the Agentic Commerce Protocol, an open standard co-developed with Stripe, with "over a million" Shopify merchants promised to follow. By March 2026, CNBC reported that OpenAI was ending Instant Checkout in favour of dedicated retailer apps inside the chatbot. The specifics are the useful part:

  • Forrester principal analyst Emily Pfeiffer put the number of Shopify merchants actually live on Instant Checkout at roughly 30, against the million promised.
  • Walmart made about 200,000 products available in ChatGPT, and found that conversion rates were three times lower for products bought directly in ChatGPT than for those that rerouted shoppers to Walmart's own site for checkout.
  • A Semrush survey of more than 1,000 US consumers found just 22% had bought a product inside an AI tool, while about half had made a purchase after using AI during research.
  • Walmart's EVP of AI acceleration, Daniel Danker, told Morgan Stanley's TMT conference on 4 March 2026 that Instant Checkout was "a very temporary moment in time."

On why it struggled, Pfeiffer identified a data problem rather than a demand problem: "Crawling and scraping is inadequate to get the full breadth of product data that you need to do a good job of commerce." Items surfaced in chat could be out of stock, or carry wrong delivery timing or shipping costs.

Put those together and the shape is clear. The agent is currently a research assistant that hands off, not a buyer that completes. Half of consumers buy after using AI in research; 22% buy inside it. Etsy described ChatGPT as a valuable discovery channel with low Instant Checkout purchase volume. That is a consideration-stage phenomenon, and consideration is exactly the stage your existing measurement is worst at.

Why this breaks shelf measurement specifically

A physical shelf and a results page share a property that makes them measurable: they are a shared, inspectable universe. Everyone standing in that aisle sees the same twelve facings. Everyone who types that keyword sees approximately the same grid, which is why a scrape is a defensible sample of it.

An agent's answer has none of those properties.

PropertyRetail shelfResults pageAgent answer
Shared across shoppersYesApproximatelyNo
Has a fixed number of positionsYesYesNo
Inspectable by a competitorYesYesNot reliably
Produces an impression you can countYesYesNo
Reproducible on demandYesRoughlyNo - varies with phrasing and context

That last row is the one that ends the metric. Share of search requires a denominator: the set of results shown for a term. An agent does not return a set of results for a term. It returns a synthesised answer to a conversation, conditioned on everything said before it, and rephrasing the same request produces a different set of brands.

So the adversary here is not a rival who outranks you. It is that nobody owns the universe of what was considered. In the retail measurement lane we have worked through several versions of this: in category management a competitor advises the retailer on your shelf; in personalized pricing the single posted price stops existing. Agentic commerce is the strongest form: the consideration set is generated per conversation and retained by nobody.

There is a second-order consequence that most 2026 commentary misses. If the agent's product knowledge comes substantially from crawling - which Pfeiffer's criticism implies - then your discoverability now depends on the machine-readability of your product data, not on your merchandising. Wrong stock status, missing attributes, and unstructured specs are no longer a conversion-rate problem on your PDP. They are an eligibility problem for being in the answer at all.

What you can still research, and it is more than you would think

You cannot interview an agent. You can interview the person who wrote the prompt, read the answer, and decided whether to accept it - and that person is the actual decision-maker in every agentic transaction that matters commercially.

This is where agentic commerce turns out to be a qualitative research problem rather than an analytics problem. The variables that determine outcomes are all in the human's head:

1. What they actually asked. The prompt is the new query, and it is far richer and far less predictable than a keyword. People describe constraints, occasions, budgets, and prior bad experiences. None of this appears in any keyword tool.

2. Whether they verified. Did they click through, open a second tab, check reviews, or simply accept the recommendation? This single behaviour splits your addressable population into two groups that need completely different strategies.

3. What made them override it. When someone rejects the agent's recommendation, the reason is a direct measurement of the brand equity, trust, or prior experience that survived an algorithmic recommendation against you. That is arguably the most valuable brand metric available in 2026 and almost nobody is collecting it.

4. What they believed they were getting. Agents summarise. Summaries drop caveats. The gap between what the agent implied and what the product delivers is a post-purchase risk that lands on your reviews and your returns, not on the model's.

Using the six structured question types, a first agentic-commerce study looks like this:

What you need to learnTypeExample
The prompt in their own wordsopen_endedType or say roughly what you asked the assistant
Whether verification happenedyes_noDid you check the recommendation anywhere else before buying?
What would make them overriderankingRank what would stop you accepting a recommendation
Which brands the agent surfacedmultiple_choiceWhich of these did it mention?
Where the purchase completedsingle_choiceDid you buy in the chat, on the retailer site, or in store?
Trust in the recommendationscaleHow much did you trust it, 1 to 7?

The multiple_choice item is doing something no other instrument can: reconstructing the consideration set from the only copy that still exists, which is the shopper's memory. It is imperfect recall, and it should be reported as such - but an imperfect reconstruction of a set that is otherwise unrecorded beats a precise measurement of a page nobody looked at. The general principle is covered in triangulation in research, and the honesty requirement in research assurance levels.

One caution worth stating, because it will bite teams this year. Do not run this study by asking an AI to simulate shoppers. Synthetic respondents cannot tell you what a real person typed into a real assistant or why they overrode it, and the failure modes of AI-generated responses in research are well documented - see survey fraud and respondent quality. The subject being AI-mediated makes real respondents more necessary, not less.

What to do in the next quarter

Instrument the handoff. Tag and separate AI-sourced sessions in analytics. Given the conversion differential Adobe reports, blending them into "referral" is discarding your highest-intent traffic segment.

Fix the product data before the merchandising. Structured, accurate, complete attributes and live stock status are now a discoverability input. This is unglamorous and it is the highest-leverage work available.

Run a small, recurring study of agent-assisted buyers. Twenty to thirty interviews a quarter with people who used an assistant in the category, focused on the four variables above. Because the platforms are changing every few months - Instant Checkout existed for roughly six - a one-off study will be stale before it is socialised. This needs to be a repeating instrument.

Do not rebuild your strategy around in-chat checkout yet. The 2026 evidence says the buying still mostly happens on your site. Optimise for being chosen in the answer and arriving well, not for transacting inside someone else's window.

Where Koji fits

Every recommendation above depends on talking to a moving target: shoppers whose tools change quarterly, in a channel with no logs. That is a recurring qualitative programme, which under legacy tooling means a standing agency retainer or a research hire, and in practice means it does not happen.

Koji is built for exactly this cadence. You describe what you need to learn; the AI consultant writes the guide; AI-moderated voice interviews run with real buyers in parallel, at any hour, asking every participant the same questions with the same follow-up logic; thematic analysis and a shareable report arrive in one click. Rerunning the study next quarter is a duplicate, not a new project - which is what makes a moving target trackable at all.

Against the alternatives: UserTesting and dscout give you recordings you still have to watch and code; Qualtrics and SurveyMonkey give you counts but cannot chase down why someone overrode a recommendation; Dovetail organises insights you have already paid someone else to gather. Koji runs the interviews, analyses them, and returns the report - 10x faster than the traditional cycle, with no research expertise required and no moderator steering the answer.

The agent will not tell you why it did not recommend you. The customer will. Start a free study on Koji and go ask them.

Frequently Asked Questions

What is agentic commerce?

Agentic commerce is the use of AI agents to carry out shopping tasks on a person's behalf - researching options, comparing them, and in some implementations completing the purchase. In 2026 the dominant real-world pattern is agent-assisted discovery followed by checkout on the retailer's own site, rather than fully autonomous buying.

Did ChatGPT's Instant Checkout succeed?

No. Launched on 29 September 2025 with Etsy and a promise of over a million Shopify merchants, roughly 30 Shopify merchants were actually live by early 2026 according to Forrester, and OpenAI moved to ending it in favour of retailer apps in the chatbot by March 2026. Walmart reported conversion three times lower for in-chat purchase than for redirecting shoppers to its own site.

Does share of search still work if shoppers use AI assistants?

Not for the AI-mediated portion of demand. Share of search needs a fixed set of results for a keyword; an agent returns a synthesised answer that varies with phrasing and conversational context and produces no countable impression. It remains valid for traditional retailer search, which is still where most volume is.

How do I find out whether an AI assistant recommends my brand?

Prompt testing gives you a rough, unstable read and is worth doing as a monitor, but it has the same weakness as scraping a personalised results page: your prompt is not their prompt. The more reliable route is asking recent category buyers who used an assistant which brands were mentioned, then treating that reconstruction as the imperfect but only surviving record of the consideration set.

Should I use AI-generated synthetic respondents to study this faster?

No. Synthetic respondents cannot report what a real person typed into a real assistant, whether they verified the recommendation, or why they overrode it - which are the four variables that determine the outcome. Use real participants and screen for response quality.

What is the single highest-leverage thing to fix first?

Product data quality and machine-readability. Assistants that build product knowledge by crawling need accurate, complete, structured attributes and live availability, and Forrester's assessment of Instant Checkout was that inadequate product data - wrong stock, wrong delivery timing, wrong shipping cost - was a core reason the experience underperformed. That work also improves your traditional listings, so it pays regardless of how fast agentic buying grows.

Run your first AI-moderated study in 10 minutes

10 free credits on signup. No credit card required.

GDPR compliantEU or US data residencyNo AI training on your data
Koji

Koji Team

Research

Share this article

Keep reading