Back to docs
Research Methods

The Differencing Attack: Why Suppressing the Small Segment Publishes It (2026)

Hiding a small cell and publishing the totals discloses the cell exactly, by subtraction. The same arithmetic works on dashboards - and rounding does not save you.

TL;DR: The standard fix for a too-small segment - hide the cell, publish the totals - usually publishes the cell. If a row shows a total of 13 and two of its three cells are visible at 7 and 4, the hidden cell is 2, exactly, by subtraction. Statistical agencies have known this since the 1970s and call the fix complementary suppression: you must also blank cells that were never sensitive, in every row and column the sensitive cell touches. The same arithmetic works on dashboards, where two filtered views one respondent apart disclose that respondent's answer - and rounding to one decimal place does not save you. In a worked case below, rounded means still pin a single respondent's 0-10 score to exactly 6.

The reflex, and why it inverts

You run a study, you build the segment table, and one cell reads n = 2. You know better than to publish it, so you blank it out and ship the rest.

You have just published it.

Here is the table, with the small cell hidden:

Plan tierEMEAAMERAPACTotal
Starter1812939
Growth22151148
Enterprise74(suppressed)13
Total473122100

The suppressed cell is 13 - 7 - 4 = 2. It is also 22 - 9 - 11 = 2 down the column. Two independent routes to the same answer, both requiring a subtraction that a reader does in their head.

The US Federal Committee on Statistical Methodology states the principle in Statistical Policy Working Paper 22: "In a row or column with a suppressed sensitive cell, at least one additional cell must be suppressed, or the value in the sensitive cell could be calculated exactly by subtraction from the marginal total." The additional cells the paper requires - "certain other non-sensitive cells must also be suppressed" - are what it calls complementary suppressions.

Note the direction of the failure. The protective act - blanking the cell - is what creates the disclosure, because it advertises where the small group is while leaving the arithmetic that recovers it intact. Had you published the raw 2, a reader would have had to notice it. Suppressing it and publishing the total tells them exactly which subtraction to perform.

This is the structural inversion at the heart of statistical disclosure control: the more visibly you protect a cell, the more precisely you locate it.

Complementary suppression is harder than it looks

The obvious remedy is to blank a second cell in each affected row and column. The FCSM working paper is careful to say that this is not sufficient either. Of a worked example with "at least two suppressed cells in each row and column," it observes: "This table appears to offer protection to the sensitive cells, however, a closer review shows disclosure of sensitive data still occurs."

The reason is that a table with margins is a system of linear equations. Each row total is an equation, each column total is an equation, and each suppressed cell is an unknown. If the system has a unique solution - or bounds the unknown tightly enough - the suppression pattern has failed no matter how many cells you blanked. The paper's own appendix lists the two audits an agency runs on any proposed pattern: whether "Implicitly Published Unions of Suppressed Cells Are Sensitive," and whether "Row, Column and/or Layer Equations Can Be Solved for Suppressed Cells."

The working paper's conclusion is one that any research team should take seriously before rolling their own: "While it is possible to select cells for complementary suppression manually, in all but the simplest of cases, it is difficult to guarantee that the result provides adequate protection."

For a research report, this means the honest options are narrower than they look:

  • Collapse the category. Merge APAC into a Rest-of-World bucket so no small cell exists. This is generalisation, and it is nearly always the right answer.
  • Drop the margin. If you publish the interior cells without row and column totals, the subtraction has nothing to work with. Readers dislike this, and they are right to, but it is at least sound.
  • Publish nothing at that granularity. If the cut is too fine for the sample, the cut is too fine for the report.

What is not an option is blanking one cell and shipping the totals, which is what nearly everyone does.

The dashboard version, which is worse

Tables at least make the arithmetic visible. Interactive dashboards hide it, and generate the equations automatically.

Consider a satisfaction tracker on a 0-10 scale. A colleague runs two views:

  • Filter: Enterprise - n = 41, mean = 4.20
  • Filter: Enterprise, excluding EMEA - n = 40, mean = 4.15

Both views comfortably exceed any minimum base size. Neither is a small cell. But they differ by exactly one respondent, so:

41 x 4.20 - 40 x 4.15 = 172.2 - 166.0 = 6.2

The single EMEA Enterprise respondent scored 6.2. That is not an inference about a group; it is one person's answer, recovered from two aggregates that both passed the base-size check.

The instinctive objection is that the dashboard rounds, so the recovery is approximate. Work it through. If both means are rounded to one decimal place, the true sums lie in intervals, and the difference lies in the interval:

41 x 4.195 - 40 x 4.155 = 5.795 to 41 x 4.205 - 40 x 4.145 = 6.605

The recovered value is somewhere in (5.795, 6.605), a window of width 0.81. On a 0-10 integer scale there is exactly one integer in that window: 6. Rounding narrowed eleven possible answers to one. It did not protect anybody.

This generalises unpleasantly. Any dashboard that lets a user (a) apply arbitrary filters and (b) see counts and means will let a determined user recover individual values, and the more precisely it reports, the fewer queries they need. The base-size rule that guards each view cannot see the difference between views, because the difference is not a view.

Where research reports leak by differencing

Four patterns account for most of it.

Wave-over-wave trackers. You publish quarterly. Q3 has 84 respondents in a segment, Q4 has 85. If the tracker reports both means, the new respondent's score is recoverable by the arithmetic above. Trackers are especially exposed because the same segments recur by design and the sample changes by small increments.

The "excluding" cut. Any report that shows both a total and a subset - "all customers," "all customers except churned" - hands the reader the complement for free. If the complement is small, it is disclosed.

Longitudinal panels with attrition. A panel that loses one member between reports is a differencing attack that ran itself.

Overlapping segment definitions. "Enterprise" and "ACV above 100k" overlap in all but two accounts. Publishing both means the symmetric difference is a two-person group with a computable average.

The common thread: none of these involves publishing a small cell. Every individual number passed review. The disclosure lives in the relationship between numbers, which is precisely what a per-artefact review cannot see - a theme that reaches its logical conclusion in the privacy budget.

What to do instead

Fix the segment grid once, and freeze it. The single most effective control is to define a standing set of report segments - coarse enough that every cell clears your k threshold with room to spare - and report only on that grid, every time. Frozen grids are immune to the "excluding" cut and to overlapping-definition differencing, because there is only one definition.

Report bands, not point estimates, for anything small. A mean reported as "between 4 and 5" cannot be differenced usefully. This costs less than it seems, because a mean computed on 40 people has a confidence interval wider than the band anyway; you are removing precision that was never real. Publishing the interval is more honest as well as safer.

Never publish both a set and its near-complement. If the report shows "all," it should not also show "all except a small group."

Ban ad-hoc filter combinations in shared dashboards. Give people the frozen grid. If someone needs a bespoke cut, it should be a request that a human reviews - which is also the point at which you notice that the cut has an n of 3.

Add a minimum-difference rule to your review. Alongside "no cell below k," add no two published figures may be based on respondent sets differing by fewer than k people. That second rule is the one that catches differencing, and almost nobody has it written down.

How Koji reduces the surface

Koji is an AI-native research platform: the AI interviewer runs voice and text conversations, probes with its own follow-up questions, and produces the analysed report automatically. Three consequences matter for differencing risk.

Reports are generated artefacts, not live query surfaces. A Koji report is produced for a study with a defined respondent set and a defined set of cuts derived from the study's structured questions. That is a fundamentally smaller attack surface than a self-serve BI dashboard where any user can compose arbitrary filters, because the set of published aggregates is finite, enumerable and reviewable before it ships. Most differencing risk in practice comes from the unbounded-query pattern, and a generated report simply does not have one.

Segments come from declared questions, so the grid is stable by construction. Koji's six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - carry stable IDs from the interview plan through analysis into report aggregation. A single_choice question with five declared options produces the same five segments in every study that uses it. That is a frozen grid arriving as a side effect of good question design, rather than as a policy somebody has to enforce.

Depth does not require a finer grid. The reason teams slice into two-person cells is that they are hunting for the story, and on a survey platform the only way to find a story is to keep cutting. In an AI interview the story is already in the transcript: the follow-up probing - up to three follow-ups per question, driven by what the respondent actually said - produces the why at the individual level, so the quantitative grid can stay coarse. You quote the Enterprise APAC customer's reasoning without publishing an n = 2 cell containing their score. The qualitative and quantitative layers carry different loads, and only one of them needs to be fine-grained.

The responsibility is still yours. Koji does not know which of your segments are sensitive or who might be reading. What it can do is keep the published aggregate set small, stable and enumerable - which is the precondition for auditing it at all.

Frequently asked questions

What is a differencing attack on a research report?

It is the recovery of a hidden or individual value by subtracting two published aggregates. The simplest form is a suppressed table cell recovered from a row total; the more common form in practice is two dashboard views that differ by one respondent, where the difference of the two totals is that respondent's answer.

If I suppress a small cell, is the report safe?

Usually not. Statistical Policy Working Paper 22 is explicit: "the value in the sensitive cell could be calculated exactly by subtraction from the marginal total" unless additional, non-sensitive cells are also suppressed. Suppressing one cell while publishing row and column totals is the single most common disclosure-control mistake in reporting.

Does complementary suppression fix it?

It is necessary but hard to get right. The same working paper shows an example with at least two suppressed cells in every row and column and notes that "a closer review shows disclosure of sensitive data still occurs," because the table is a solvable system of equations. Its own conclusion is that manual selection of complementary cells cannot be guaranteed adequate "in all but the simplest of cases." Collapsing the category is the more reliable fix.

Does rounding protect against differencing?

Only partially, and often not at all. In the worked example above, two means rounded to one decimal place still bound one respondent's 0-10 score to the interval (5.795, 6.605), which contains exactly one integer. Rounding turns an exact recovery into a narrow one; whether that matters depends on how many values the answer could take, and for a rating scale the answer is usually "not enough to help."

How do I stop dashboards from leaking this way?

Restrict published cuts to a frozen segment grid rather than allowing arbitrary filter composition, and add a minimum-difference rule to your review: no two published figures may be based on respondent sets differing by fewer than k people. Base-size rules alone cannot catch differencing, because each individual view passes.

Are trackers more exposed than one-off studies?

Yes. A tracker reports the same segments repeatedly over a sample that changes by small increments, which is the ideal setup for differencing - a segment that goes from 84 to 85 respondents publishes the 85th respondent's score if both waves report a mean. Report tracker segments as bands, or hold the panel composition fixed within a reporting period.

Related Resources

Related Articles

k-Anonymity for Segment Reporting: How Small Is Too Small to Publish? (2026)

The rule for minimum base size in research reporting, stated exactly: every visible combination of attributes must be shared by at least k respondents - and why generalisation beats suppression.

Every Report Was Safe and the Set Was Not: The Privacy Budget in Research Reporting (2026)

Disclosure controls are applied one report at a time, but privacy is a property of the whole release history. What the Census reconstruction attack proves about aggregate reporting.

Quasi-Identifiers in Research Data: Why Removing Names Does Not Anonymise a Study (2026)

Deleting the name column does not anonymise a study. The identifier is the combination of screener fields you kept - here is how to measure it before you publish.

How to Read Your Koji Research Report: A Section-by-Section Guide

A complete walkthrough of every section in a Koji research report — from the overview and themes to quantitative charts, key quotes, and Insights Chat — so you can extract maximum value from your findings.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

User Research Report Template: How to Present Findings That Drive Action

A complete guide to writing user research reports that stakeholders actually read — with a proven structure, templates for key sections, and how AI-generated reports change the game.