Once a multi-unit brand accepts that it cannot see its customers through its operators, the obvious fix is to mandate measurement: require every location to collect customer feedback, standardise the instrument, and put the resulting score on the operator's scorecard so it actually gets done. Response rates go up. Coverage goes up. Compliance goes up.
And the data stops being evidence.
This is the counter-intuitive part of multi-location customer feedback, and it is worth stating plainly: attaching a satisfaction score to the consequences of the person who collects it does not make the score more reliable, it makes it less reliable, and it does so while every visible quality indicator improves. You end up with a more precisely measured version of a number that no longer means what its name says.
The mechanism: you handed sample selection to the person being graded
Nobody has to lie for this to happen. A location manager whose renewal, bonus or league position depends on a satisfaction score controls three levers, all of them legitimate-looking:
- Who gets asked. The manager decides which customers get handed the card, hear the reminder, or get walked to the tablet. Asking the delighted customer and forgetting the annoyed one is not fraud. It is a completely natural human response to being measured.
- When they get asked. Immediately after the recovery, not during the wait. At the end of a good visit, not a bad one. Timing alone moves a score several points without anyone touching a response.
- What the customer is told the score means. "Anything less than a 10 counts as a fail for me" is a sentence that gets said in thousands of service businesses every day. It is not a request to lie. It reframes the scale.
The output is a genuine set of responses from real customers who really felt that way. The number is not fabricated. The sample is. And a biased sample of honest answers is harder to detect than a fabricated number, because every integrity check you run on the responses comes back clean.
This is not marking your own homework
It is worth separating this from a problem it superficially resembles. When a retail media network reports the return on ad spend for campaigns it also sold, the concern is that the seller is grading its own performance with its own metric definition. That is a self-reported metric problem, and it is covered separately in retail media measurement.
What happens on a location scorecard is different and, in a sense, worse. The metric definition is fine. The responses are real. The arithmetic is correct. What has been quietly delegated is the sample, and sampling decisions leave no trace in the data. There is no field in the export that records the customers who were never invited.
The most controlled version of this experiment has already been run
The best available evidence on what happens to satisfaction measurement when it is standardised, audited and tied to money does not come from franchising. It comes from US hospitals, and it is public.
HCAHPS is the federal patient experience survey. The instrument is mandated, the survey modes are mandated, scores are adjusted for survey mode and patient mix, an approved-vendor regime governs administration, and the results feed hospital payment. If measurement under high stakes could be made to work by standardisation and audit alone, it would work here.
The July 2026 public report, covering discharges from October 2024 through September 2025 across 3,949 hospitals, shows what it actually produces:
- The national response rate is 23%. More than three quarters of discharged patients never answer, in a survey that is legally mandated, professionally administered and payment-linked. State-level response rates range from 16% to 31%.
- On the Discharge Information measure, the entire national distribution spans 14 points. A hospital at the 5th percentile, described in the published table as "near worst", scores 79. A hospital at the 95th percentile, "near best", scores 93.
- Half of all 3,949 hospitals sit inside a five-point band. The 25th percentile is 84 and the 75th percentile is 89. Roughly 1,974 institutions are separated by less than the width of a rounding convention.
Two honest caveats, because the point is not that all such measures collapse. Compression is measure-specific: the overall Hospital Rating measure, the share of patients giving a 9 or 10, spans 30 points from the 5th to the 95th percentile, and that measure discriminates perfectly well. And HCAHPS is a healthcare survey, not a franchise scorecard, so this is an analogy about structure rather than a finding about your brand.
But the structure is the same structure. Independent units, a common instrument, mandated collection, published rankings, and money attached. And the lesson transfers cleanly: when you rank units on a measure whose distribution is compressed, you are ranking noise, and the units at the top learned to be at the top.
If you are building league tables from location scores, read internal benchmarks and percentile norms before you publish one, and customer experience benchmarking before you compare against an industry figure.
The regulator drew the line exactly where the incentive is, and your scorecard is on the far side
The Federal Trade Commission's Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465, took effect on 21 October 2024. Two of its provisions describe the scorecard problem almost exactly.
Section 465.4 makes it an unfair or deceptive practice for a business to provide "compensation or other incentives in exchange for, or conditioned expressly or by implication on, the writing or creation of consumer reviews expressing a particular sentiment, whether positive or negative".
Section 465.7(b) reaches the selection side. It prohibits a business from materially misrepresenting that displayed reviews "represent most or all the reviews submitted" when reviews are being suppressed "based upon their ratings or their negative sentiment", and it sets out the test that makes the difference: a review is not treated as suppressed if the withholding follows "criteria for withholding reviews that are applied equally to all reviews submitted without regard to sentiment".
Applied equally, without regard to sentiment. That is the standard, and it is the right one.
Now the part that should worry any brand running a location scorecard: Part 465 governs public consumer reviews. It does not govern your internal CSAT programme. The behaviour the FTC considered serious enough to regulate, conditioning an incentive on sentiment and selecting who is heard by predicted sentiment, is not unlawful when it happens inside your own survey tool against your own dashboard. Nobody is coming to enforce a standard on your scorecard. Which means the only thing standing between your programme and the same distortion is how you designed it.
Brands that also publish or quote this feedback in marketing have a second exposure here, covered in using research quotes in marketing.
How to separate measurement from evaluation
The fix is structural, not motivational. You cannot train your way out of this, and you should not try, because the operators are responding rationally to the system you built.
- Draw the sample centrally, from the transaction log. A random or census draw from the brand's record of what was actually sold. Not a list the location supplies, and not a card the location hands out.
- Never let the unit know who was drawn. If the location can identify the selected customers before they respond, you have reintroduced every lever above.
- Keep the incentive on the response rate, not on the score. Rewarding a location for the number of customers reached is safe, because the location cannot control what those customers say. Rewarding it for the score is not.
- Use the same instrument, asked the same way, every time. Human administration varies by definition. That variance lands in the location comparison and looks like performance.
- Report the distribution, not just the mean. A location's mean says nothing when the network's interquartile range is five points. Show the spread and let people see how little of it is real.
- Ask why, not only how much. A compressed score cannot tell you what to change. A transcript can. And if the reason a location scores badly is a method-of-operation detail rather than a staffing one, you can fix it once across the network. Which findings you can actually act on is a separate and surprisingly restrictive question, covered in brand standards and customer research.
Points one through four are all versions of the same instruction: take the sampling decision away from the person whose number depends on it. See also research participant incentives on what incentive structures do to who responds, and panel conditioning on what repeated asking does to the same people over time.
Where Koji fits
The problem above is a sampling problem wearing a survey costume, so the fix has to happen in how the study is fielded, not in how the responses are analysed.
- Koji draws the sample and runs the interview. The invitation goes from the brand to the customer. The operator never receives the list, never hands out a device, and never learns who was contacted before they answer.
- AI-moderated voice interviews remove the person from the room. There is no local manager standing nearby, no "anything less than a 10 counts as a fail", and no moderator whose tone varies by location or by time of day. Every respondent is asked the same way, which is what makes a location comparison mean anything.
- Six structured question types give you the comparable number and the reason in one session: open_ended, scale, single_choice, multiple_choice, ranking and yes_no. You get the scale score for the dashboard and the open-ended explanation for the operations team from the same conversation.
- Automatic thematic analysis surfaces what actually differs between your top and bottom locations, rather than confirming that a five-point gap exists.
- One-click reports put the distribution and the themes in front of the operations director in hours, not at the end of a quarterly cycle.
Traditional CX suites like Qualtrics and Medallia will happily build you the league table, because the league table is the product they sell to head office. Typeform and SurveyMonkey will collect responses from whoever is handed the link, which is precisely the failure mode described above. UserTesting and dscout recruit from panels, so they cannot reach the customer who visited unit 412 at all. Koji is AI-native and built the other way round: brand-controlled sampling, unbiased AI moderation, and 10x faster insight than a moderated multi-location study, with no research expertise required.
Frequently asked questions
Why are all our location CSAT scores 9s and 10s?
Almost always because the people collecting the responses are graded on them. Location staff control who is asked, when they are asked, and how the scale is framed, and none of those levers require anyone to falsify a response. The result is a real set of answers from an unrepresentative set of customers, which is much harder to detect than a fabricated number.
Is it illegal to ask only happy customers for feedback?
For public reviews, it can be. The FTC's Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465, prohibits incentives conditioned on sentiment (465.4) and misrepresenting that displayed reviews represent most or all submitted reviews when negative ones are suppressed (465.7(b)). For an internal satisfaction survey shown only on your own dashboard, Part 465 generally does not apply, which is exactly why internal programmes drift further than public ones. This is general information, not legal advice.
Should we tie franchisee bonuses to customer satisfaction scores?
Tie them to response rate or coverage instead. A location can control how many customers are reached and cannot control what those customers say, so incentivising reach improves the data while incentivising the score corrupts it. If the score must carry consequences, the sample has to be drawn centrally and kept invisible to the location.
What response rate should a multi-location feedback programme expect?
Lower than most brands assume. HCAHPS, a federally mandated, professionally administered, payment-linked patient survey, reports a 23% national response rate across 3,949 hospitals, with state figures ranging from 16% to 31%. A voluntary commercial programme with no mandate behind it should plan for that range or below, and should design for non-response rather than hoping to eliminate it.
How do we know if our location rankings are real?
Look at the spread before you look at the order. If the gap between your 25th and 75th percentile locations is small relative to the measurement error of the instrument, the middle of your ranking is noise and reshuffles every period regardless of what anyone did. Publishing the distribution alongside the ranking makes this visible to everyone reading it.
Can we fix a gamed programme without replacing it?
Only by changing who draws the sample. Retraining, auditing and adding attention checks all operate on the responses, and the responses were never the problem. Until the location stops choosing which customers are invited, every other control leaves the underlying selection untouched.
Measure your locations without asking them to measure themselves
Koji runs brand-controlled AI interviews with customers your operators never select, and returns both the comparable score and the explanation behind it. From question to insight in hours, across every location at once.
Start a study with Koji and see what your customers say when nobody is standing next to them.