Every analysis of shopper behaviour runs on a dataset with an edge. The edge is not drawn by the shopper, the category, or the question you are asking. It is drawn by whoever owns the sensor - the retailer whose loyalty card recorded the transaction, the marketplace whose account tracked the order, the panel whose members agreed to be observed.
This article is the third in a series on the shopping trip as a unit of analysis, and it argues that the first two rest on an assumption that is usually false. Market basket analysis assumes the basket you can see is the basket. Shopper mission research assumes you can observe the trip. Neither is true, because no dataset in existence contains a whole shopping trip.
The short answer
Cross-shopping is the normal state, not an edge case. The average American visits 5.4 separate grocery banners in a single month, according to FMI's U.S. Grocery Shopper Trends 2026, released in May 2026 and developed with The Hartman Group from a nationally representative survey of 2,023 respondents fielded 4-18 February 2026. Younger shoppers spread further still: Gen Z visits 6.7 banners a month and millennials 6.1.
If a household shops five or six banners, then any single retailer's data contains roughly one fifth of that household's category behaviour, and no field in that dataset indicates what is missing. The absence is invisible, which is materially worse than the absence being large.
What each sensor can and cannot see
| Data source | Universe it defines | What it cannot see, by construction |
|---|---|---|
| Retailer loyalty card | Transactions at one banner by identified members | Every trip to a competitor; unidentified trips; the reason for any trip |
| Point-of-sale / scanner | All transactions at participating stores | Who the shopper was; what they intended; what they did not find |
| Consumer panel | Purchases members remember to record | Under-recorded small trips; non-member households; the same bias that made them join |
| Card spend data | Spend at merchant level on tracked cards | Basket contents; cash trips; which household member; why |
| Retail media network | Ad exposure and conversion inside one retailer | Conversions elsewhere; the counterfactual purchase |
| Your own DTC store | Your customers, your site | The 80% of their category spend that happens elsewhere |
Read the right-hand column as a single sentence and the capstone argument appears: each source is blind to a different thing, and none of them is blind to nothing. The habit of triangulating across several of them helps only if their blind spots differ, which brings us to the trap.
Adding more sources does not automatically fix this
The instinct when told a dataset is partial is to add another. That works only where the second source could have disagreed with the first. Several of the sources above share an origin - retail media measurement is downstream of the same retailer's loyalty data; a syndicated read and a retailer read of the same banner are two views of one transaction stream. Stacking them raises apparent confidence without adding independent evidence. Our guide to triangulation covers the general form of this, and retail media measurement covers the specific case where the network reporting the result also sold the media.
The corollary is uncomfortable but clean: the fact that all your data agrees is not evidence that it is right, if all your data comes through the same sensor.
The loyalty data illusion
Loyalty programme data feels like it solves cross-shopping, because it identifies the household. It does not, and the numbers are stark.
Cardlytics analysed the entire US grocery category in its card-spend data - representing over $200bn in annual card spend - across eight quarters from Q1 2023 to Q4 2024. On average, 58% of a merchant's customers are not loyal to it. Loyal customers hold an 84% share of wallet, while non-loyal customers hold 18%, close to a 3x gap. And more than a third of a merchant's top 20% most frequent transactors are still not loyal customers by that definition.
That last figure is the one to keep. The shoppers a retailer considers its best - the ones at the top of the frequency ranking, the ones whose behaviour drives the category plan - include a large group giving most of their money to somebody else. The loyalty card sees their frequency and reads it as devotion. It cannot see the four other banners where the rest of the basket went.
Cardlytics also found non-loyal customers are 42% to 185% more likely to lapse or churn. So the segment your data systematically mischaracterises is also the segment most likely to leave, which is a fairly direct description of how retention programmes end up surprised.
Why this is a research problem and not a data problem
The reflex is to solve this by acquiring more data: buy a panel, license card spend, join a data collaboration. These help, and they are the right move for sizing. But each one is another sensor with another boundary, and the shopper's actual decision process still spans all of them.
Consider what a shopper does that no combination of transaction datasets will ever contain:
- They chose which store to go to, for reasons that happened before any sensor switched on
- They did not buy something they intended to buy, and no dataset records an intention
- They bought at store A because of something they saw at store B
- They split one shopping mission across two retailers deliberately
- They would have bought more if a condition had been different
Every item on that list is a counterfactual or an intention. Neither is a recordable event. This is the same structural point made in survey universe definition: the boundary of your frame determines what conclusions are even available to you, and a boundary you did not choose is one you probably have not noticed.
The only instrument that spans retailers is the shopper. They were present at all 5.4 banners. They are the sole entity in the system with a view of the whole trip.
How to run cross-shop research
- Define the universe as the household's category behaviour, not your transactions. Write down the denominator you actually care about before you look at any data you happen to own.
- Recruit on category behaviour, not on your customer list. Recruiting from your own CRM guarantees a sample selected on the outcome you are studying - see survivorship bias in customer research.
- Reconstruct a period, not a purchase. Ask about every place they bought the category in the last two weeks, in order, and what each trip was for.
- Ask for the split and the reason for the split. Share of wallet is a number you can ask for directly and reasonably accurately at the household level, and the reason behind it exists nowhere else.
- Size it afterwards. Once you know the shapes, use panel or card data to estimate how common each is. Qualitative first to find the structure, quantitative second to weight it - not the reverse.
Where Koji fits
Cross-shop research has historically meant a custom study through a large agency, because it needs real conversations with a population defined by category behaviour rather than by any one retailer's list. That is a six-figure project with a three-month timeline, which is why most teams settle for the loyalty data they already have and quietly accept the blind spot.
Koji collapses that. You define the audience by category behaviour, run AI-moderated voice interviews in parallel, and get trip-by-trip reconstructions across every banner the household uses in hours rather than months - with no moderator drifting into the retailers they personally find interesting.
Koji's six structured question types capture the wallet split and the narrative in one session: open_ended for the two-week reconstruction across all banners, multiple_choice for every retailer used in the period, scale for estimated share of spend at each, single_choice for the primary store and what makes it primary, ranking to order the reasons for splitting the trip, and yes_no for whether a specific trip was planned as a split or became one. One-click reports turn that into a picture of the household's whole category behaviour - the view no retailer, no panel, and no card-spend file contains on its own.
Ten years of transaction data will tell you what happened inside your boundary with great precision. It will never tell you where the boundary is. For that you have to ask the only participant who was on both sides of it.
Frequently Asked Questions
What is cross-shop analysis?
Cross-shop analysis examines how households distribute their category spending across multiple competing retailers or brands rather than concentrating it in one. It answers who else your customers buy from, how much of their spend you actually capture, and why they split it. It requires data that spans retailers, which most transaction datasets do not.
How many grocery stores does the average shopper use?
FMI's U.S. Grocery Shopper Trends 2026, released in May 2026, found Americans visit 5.4 separate grocery banners in an average month. Gen Z shoppers visit 6.7 and millennials 6.1. The findings come from a nationally representative survey of 2,023 respondents fielded 4-18 February 2026, conducted with The Hartman Group.
Does loyalty card data solve cross-shopping?
No. Loyalty data identifies the household but only within one banner. Cardlytics found that on average 58% of a merchant's customers are not loyal to it, and that more than a third of a merchant's top 20% most frequent transactors are non-loyal. A loyalty card reads frequency as devotion because it cannot see the other stores.
What is share of wallet in retail?
Share of wallet is the proportion of a household's total category spending that goes to one retailer or brand. Cardlytics grocery data shows loyal customers hold an 84% share of wallet compared with 18% for non-loyal customers, roughly a 3x difference across an analysis covering over $200bn in annual card spend from Q1 2023 to Q4 2024.
Why does adding more data sources not fix a blind spot?
A second source improves confidence only if it could have disagreed with the first. Sources that share an origin - such as a retail media network's reporting and the same retailer's loyalty data - produce agreement that reflects a common upstream feed rather than independent confirmation. Agreement across sensors that share a boundary is not evidence.
How do you research cross-shopping behaviour?
Define the universe as household category behaviour rather than your own transactions, recruit on category behaviour instead of from your customer list, and reconstruct a period rather than a single purchase. Ask which retailers were used in the last two weeks, what each trip was for, and the reason spend was split. Then use panel or card data to size the patterns you found.
Related reading
- Market Basket Analysis: What Co-Purchase Data Proves and What It Cannot
- Shopper Mission Research: Why the Same Customer Is a Different Shopper on Every Trip
- Retail Media Measurement: Why Network-Reported ROAS Is Not Evidence
- NielsenIQ vs Circana vs Numerator: Retail Measurement Data Compared
- Why Did Market Share Drop? How to Diagnose Share Loss Your Dashboard Cannot Explain