{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-24T23:01:47.424Z"},"content":[{"type":"documentation","id":"3621b1db-6ba1-4f3f-9df0-0988fd8dec93","slug":"k-anonymity-segment-reporting-minimum-base-size","title":"k-Anonymity for Segment Reporting: How Small Is Too Small to Publish? (2026)","url":"https://www.koji.so/docs/k-anonymity-segment-reporting-minimum-base-size","summary":"k-anonymity requires that every sequence of quasi-identifier values in a release appears at least k times. Statistical agencies default to a threshold of 3 for count data; 5 is a reasonable internal-report floor and 10 for external publication or employee research. Generalisation beats suppression: coarsening date of birth from full date to year-and-month cuts population uniqueness from 63.3% to 4.2% with no records hidden. k-anonymity protects the row, not the answer - a group where all k gave the same response discloses it, which is the homogeneity attack l-diversity was designed for.","content":"**TL;DR:** A published segment is safe when every combination of visible attributes in it is shared by at least k respondents - that is Sweeney's k-anonymity requirement, and it is the closest thing research reporting has to a hard rule. You reach it two ways: **generalisation** (coarsen the categories) and **suppression** (hide the cell). Generalisation is almost always the better trade, because coarsening a date of birth from the full date to year-and-month drops population uniqueness from 63.3% to 4.2% at essentially no analytical cost. But k-anonymity alone is not enough: if all k respondents in a group gave the same answer, the group discloses that answer for every one of them. That is the homogeneity attack, and the defence is l-diversity, not a bigger k.\n\n## The rule, stated exactly\n\nLatanya Sweeney's 2002 paper in the *International Journal on Uncertainty, Fuzziness and Knowledge-based Systems* gives the definition in one line:\n\n> **Definition 3. k-anonymity.** Let RT(A1,...,An) be a table and QIRT be the quasi-identifier associated with it. RT is said to satisfy k-anonymity if and only if each sequence of values in RT[QIRT] appears with at least k occurrences in RT[QIRT].\n\nTranslated into research language: take the set of attributes that will be visible in your deliverable - role, company size band, region, plan tier, whatever your cuts are. Group your respondents by that set. If the smallest group has k members, your release is k-anonymous. Nobody can be narrowed down past a crowd of k.\n\nThe elegance of the definition is that it is a property you can *check*, in one query, on the exact artefact you are about to publish. It replaces \"does this feel identifying?\" with a number. That is why it has survived twenty-plus years of criticism from people who can name its limitations - some of which we get to below.\n\n## Choosing k\n\nThere is no universal correct value, but the practice of official statistics converges on small integers, and the reasoning is worth borrowing.\n\nThe US Federal Committee on Statistical Methodology's *Statistical Policy Working Paper 22* describes the family of rules agencies use to flag cells for suppression - \"the (n) threshold rule, (n, k) rule, and the p-percent or pq rules\" - and makes an observation that is directly relevant to survey and interview data: \"since all respondents contribute the same value to a frequency count, the rules default to a threshold rule and the cell is sensitive if it has too few respondents. The p% and pq rules default to a threshold rule of 3 when applied to count data.\"\n\nA threshold of 3 is the floor of official practice for counts. For customer research the honest defaults are:\n\n| Context | Working minimum | Why |\n| --- | --- | --- |\n| Internal analysis, named audience under NDA | k = 3 | Matches the statistical-agency floor for count data |\n| Cross-functional report, wide internal circulation | k = 5 | Standard in health and education reporting; survives forwarding |\n| External publication, benchmark, or marketing content | k = 10 | Assume an adversary with a customer list and a LinkedIn account |\n| Employee research | k = 10 minimum, and see below | The adversary is the respondent's own manager |\n\nEmployee research deserves the special case. In customer research the person trying to re-identify a respondent usually has no motive; in employee research they may be in the reporting line. Anything under a team size of about 10 should be rolled up, and it is worth saying so in the invitation, because a stated floor is also a candour intervention - respondents answer more honestly when the reporting rule is published in advance.\n\n## Generalisation beats suppression\n\nSweeney's paper offers two mechanisms for reaching k-anonymity: generalisation, which replaces a specific value with a broader one, and suppression, which removes the value entirely. Her own assessment of the second is blunt - suppression \"can drastically reduce the quality of the data.\"\n\nThe quantitative case for preferring generalisation comes from Philippe Golle's 2006 replication of Sweeney's uniqueness study on 2000 census data. His Table 1 reports the fraction of the US population uniquely identifiable by {gender, location, date of birth}:\n\n| Location granularity | Year of birth | Year and month | Full date |\n| --- | --- | --- | --- |\n| 5-digit ZIP code | 0.2% | 4.2% | 63.3% |\n| County | 0.0% | 0.2% | 14.8% |\n\nMoving one column left - dropping the day of the month - takes uniqueness from 63.3% to 4.2%. Moving one row down, from ZIP to county, takes it from 63.3% to 14.8%. Doing both leaves 0.2%.\n\nNo cell was hidden. No respondent was dropped. The analytical loss is close to zero, because no study needs a respondent's exact birthday. That is the trade generalisation offers and suppression does not: suppression removes information about the people you were most interested in, while generalisation blurs information you were not using anyway.\n\nApplied to a research screener, the generalisation ladder looks like this. Take a 120-respondent study cut on role (6 values), company size (5), industry (8), country (12) and tenure (4):\n\n| Fields retained in the deliverable | Distinct profiles | Respondents per profile on average |\n| --- | --- | --- |\n| All five | 11,520 | 0.01 |\n| Role, size, industry, country | 2,880 | 0.04 |\n| Role, size, industry | 240 | 0.50 |\n| Role, size, region (2 values) | 60 | 2.00 |\n| Role, size | 30 | 4.00 |\n| Role, region | 12 | 10.00 |\n\nAn average of 4.00 respondents per profile does not mean the release is 4-anonymous - the *minimum* class size is what counts, and with 30 cells and 120 people the smallest cell will usually hold one or two. Average class size is a necessary condition, not a sufficient one. But the ladder tells you where to start: you cannot reach k = 5 on a 120-person study while retaining five segmentation fields, and no amount of careful reviewing will change that arithmetic. You have to coarsen.\n\n## The failure k-anonymity does not catch\n\nHere is where a lot of teams stop, and where they should not.\n\nAshwin Machanavajjhala and colleagues opened their l-diversity paper with a scenario that has become the standard illustration. Alice wants to know why her neighbour Bob went to hospital. She finds a 4-anonymous table of inpatient records. She knows Bob is \"a 31-year-old American male who lives in the zip code 13053,\" so she knows his record is one of four. And then: \"all of those patients have the same medical condition (cancer), and so Alice concludes that Bob has cancer.\"\n\nThe table satisfied k-anonymity perfectly. k-anonymity protects the *row* - which respondent is which - and says nothing about the *answer*. If a group of k respondents all gave the same response to the sensitive question, membership in the group is the disclosure.\n\nThis is not a hypothetical in customer research. Consider a report cell reading:\n\n> Enterprise / EMEA / Admin role (n = 6): 6 of 6 said they are actively evaluating a competitor.\n\nSix is a respectable k. It is also a public statement that every Enterprise EMEA admin in the study is churning, which is a disclosure about each of them individually. The authors' prescription is that \"the sensitive attributes are well-represented in each group\" - a group should contain at least l distinct, reasonably frequent values of the sensitive attribute, not just l distinct people.\n\nIn practice, for research reporting, this reduces to two extra checks after the k check passes:\n\n1. **Response diversity.** For every published cell, is the sensitive answer distribution non-degenerate? A 6/6 or 0/6 split is a disclosure regardless of n.\n2. **Near-degenerate splits.** A 5/6 split is barely better than 6/6 when the reader can guess which one is the exception. If the cell is close to unanimous *and* small, roll it up.\n\nThe second check catches the case l-diversity was designed for and the first catches the obvious one. Neither is expensive; both are routinely skipped.\n\n## The other thing k-anonymity does not do\n\nk-anonymity is a property of a *single release*. It says nothing about what happens when you publish a second table over the same respondents, and that turns out to matter enormously - a suppressed cell can often be recovered exactly by subtraction from a published total, and a sequence of individually-safe reports can compose into a disclosure that none of them contained. Those are the subjects of the [differencing attack](/docs/cell-suppression-differencing-attack-research-reports) and the [privacy budget](/docs/privacy-budget-research-reporting-composition) respectively, and you should read at least the first before you rely on suppression as your main control.\n\n## Making the check routine with Koji\n\nThe reason minimum-base-size rules get violated is almost never ignorance. It is that the check lives in a reviewer's head and the report gets built at 6pm.\n\nKoji is an AI-native research platform - the AI interviewer runs voice or text conversations, asks its own follow-up questions, and generates the analysed report automatically - and two properties of that architecture make the k check a lot cheaper to run.\n\n**The question schema is explicit, so the quasi-identifier set is enumerable.** Koji's six structured question types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` - each carry a stable ID from the interview plan through analysis into report aggregation. A `single_choice` question with five declared options *is* a generalisation, decided at design time and visible in the study configuration. Choosing \"201-500 employees\" over a free-text headcount field is a k-anonymity decision made before the first interview runs, which is the only point at which it is cheap.\n\n**Report cuts are derived from those declared questions, so cell sizes are computable rather than discovered.** When a report aggregates a `scale` question by a `single_choice` segment, the n for each cell is a known quantity at generation time - not something a reader has to notice in a footnote. Traditional survey tools hand you a crosstab and leave the base-size discipline to you.\n\n**Coarsening does not cost you the qualitative depth.** This is the part that makes generalisation politically possible. The usual objection to collapsing \"Director\" and \"VP\" into \"Senior leadership\" is that the nuance disappears. In an AI interview it does not, because the nuance lives in what the person said, and Koji's follow-up probing - up to three follow-ups per question, driven by the actual answer - captures it in the `open_ended` responses. You can report on a coarse grid and still quote the specific, because the specificity moved from the demographic field into the transcript. That is the opposite of the trade a survey tool forces on you, where coarse categories are all you have.\n\nKoji does not make the disclosure determination for you. What it does is make the inputs to the determination - which fields are visible, how many people are in each cell, how the answers are distributed - available before the report leaves the building rather than after.\n\n## A working procedure\n\n1. **Fix the quasi-identifier set** for this deliverable. Not the set you collected - the set that will be *visible*.\n2. **Group and count.** Find the minimum class size. That number is your k.\n3. **If k is below your floor, generalise first.** Collapse the finest-grained field. Re-check. Repeat.\n4. **Only then suppress**, and read the [differencing article](/docs/cell-suppression-differencing-attack-research-reports) before you do, because naive suppression frequently fails.\n5. **Check response diversity** in every surviving cell. Unanimous small cells are disclosures with a passing k.\n6. **Record the floor in the report template.** A rule that is not in the artefact is a rule that will be broken by whoever builds the next one.\n\n## Frequently asked questions\n\n### What is k-anonymity in plain language?\n\nA release is k-anonymous if every respondent is indistinguishable from at least k-1 others on the attributes that are visible. Sweeney's formal version requires that each sequence of quasi-identifier values \"appears with at least k occurrences\" in the released table. In reporting terms: no published combination of segment attributes may describe fewer than k people.\n\n### What is the minimum sample size for reporting a segment?\n\nThere is no single answer, but the statistical-agency floor for count data is 3 - *Statistical Policy Working Paper 22* notes that the common sensitivity rules \"default to a threshold rule of 3 when applied to count data.\" For a widely circulated internal report, 5 is a reasonable working minimum; for anything published externally, 10. For employee research, 10 is a floor rather than a target, because the likely adversary is in the reporting line.\n\n### Should I generalise or suppress?\n\nGeneralise first. Golle's census figures show the leverage: coarsening date of birth from the full date to year-and-month cuts population uniqueness from 63.3% to 4.2% without hiding a single record. Suppression, in Sweeney's words, \"can drastically reduce the quality of the data,\" and it removes information about exactly the small groups you were most curious about.\n\n### Is a 5-anonymous report definitely safe to publish?\n\nNo. k-anonymity protects which respondent is which, not what they said. If all five people in a group gave the same answer to the sensitive question, the group membership discloses that answer for each of them - the homogeneity attack from the l-diversity paper, where a 4-anonymous table let an adversary conclude her neighbour had cancer. Check that each published cell has a genuinely mixed answer distribution as well as enough people in it.\n\n### Does k-anonymity cover multiple reports over the same respondents?\n\nIt does not. The definition applies to one released table. Two individually k-anonymous releases over the same panel can be differenced against each other to isolate individuals, and a long series of safe releases composes into an unsafe whole. That is a separate control, covered in the differencing and privacy-budget articles.\n\n### How do I apply this to open-ended interview quotes?\n\nk-anonymity is defined over structured attributes, so it does not directly govern verbatims - a quote can be unique even in a large, well-mixed cell. Run quote review as a separate control: strip employer names, distinctive event references, and any detail that would identify the respondent to a colleague. The structured check and the verbatim check are complementary, and passing one says nothing about the other.\n\n## Related Resources\n\n- [Quasi-Identifiers in Research Data](/docs/quasi-identifiers-research-data-reidentification) - why the combination of screener fields is the identifier, and how to measure uniqueness on your own respondent table.\n- [The Differencing Attack](/docs/cell-suppression-differencing-attack-research-reports) - what goes wrong when you reach k by suppression and publish the totals anyway.\n- [The Privacy Budget in Research Reporting](/docs/privacy-budget-research-reporting-composition) - why a series of individually safe releases is not a safe series.\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types, and how declared option sets act as design-time generalisation.\n- [Survey Sample Size: How Many Responses Do You Really Need?](/docs/survey-sample-size-guide) - the statistical-power side of base size, which sets a floor for a different reason.\n- [Is 4.1 Good? Internal Benchmarks and Percentile Norms](/docs/internal-benchmarks-percentile-norms) - what to do with segment numbers once they are big enough to publish.","category":"Research Methods","lastModified":"2026-08-24T03:25:26.339056+00:00","metaTitle":"k-Anonymity for Segment Reporting: Minimum Base Size Rules (2026)","metaDescription":"How small can a research segment be before you publish it? The k-anonymity rule, sensible values of k, why generalisation beats suppression, and the homogeneity attack.","keywords":["k-anonymity","minimum base size","segment reporting","l-diversity","generalisation and suppression","minimum sample size to report","disclosure control"],"aiSummary":"k-anonymity requires that every sequence of quasi-identifier values in a release appears at least k times. Statistical agencies default to a threshold of 3 for count data; 5 is a reasonable internal-report floor and 10 for external publication or employee research. Generalisation beats suppression: coarsening date of birth from full date to year-and-month cuts population uniqueness from 63.3% to 4.2% with no records hidden. k-anonymity protects the row, not the answer - a group where all k gave the same response discloses it, which is the homogeneity attack l-diversity was designed for.","aiPrerequisites":["Familiarity with segment crosstabs and base sizes"],"aiLearningOutcomes":["State the k-anonymity requirement and check it with one query","Choose a defensible value of k for a given audience","Prefer generalisation over suppression and quantify the trade","Catch homogeneity failures that a passing k check misses"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}