{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-14T17:21:10.162Z"},"content":[{"type":"documentation","id":"1a3a93a5-fb65-4c42-8db0-a338c10855d5","slug":"research-review-selection-sampling","title":"How to Choose What to Review: Sampling Rules for Tickets, Recordings, and Transcripts You Already Have","url":"https://www.koji.so/docs/research-review-selection-sampling","summary":"There are two sampling problems in research and only recruitment is well covered. The second is selecting what to read from evidence you already own. Auditing formalized this as audit sampling (PCAOB AS 2315), and its central rule is that items examined 100 percent because they were individually significant are removed from the sampling population and cannot be projected - so the vivid cases you deliberately chose may be described but never used to compute a rate. Three further rules follow: every item needs an opportunity to be selected (recency sorting and skipping long threads are named biases), stratify by significance such as ARR or affected accounts rather than one item one vote, and set the tolerable deviation rate before reading begins.","content":"Answer first: there are two completely different sampling problems in research, and the literature only covers one of them. The first is who to recruit, which every sampling guide addresses. The second is what to read from the mountain of evidence you already own - four thousand support tickets, three hundred session recordings, nine hundred open-ended responses, sixty transcripts. Auditing has spent decades formalizing the second problem, and the single most important rule it produced is this: the items you deliberately picked because they looked significant can be described, but they can never be projected to the population.\n\n## The short answer\n\nAuditing calls it audit sampling, and PCAOB AS 2315.01 defines it as \"the application of an audit procedure to less than 100 percent of the items within an account balance or class of transactions for the purpose of evaluating some characteristic of the balance or class.\"\n\nSwap in your own nouns and it is exactly the problem you have on a Monday morning with a support inbox. The standard then supplies four rules that product teams almost never apply.\n\n| Rule | The standard | What it means for your review |\n| --- | --- | --- |\n| Split the census from the sample | AS 2315.21 | Items you chose to examine 100 percent are not part of the sample and cannot be projected |\n| Give every item a chance of selection | AS 2315.24 | Scrolling until something looks interesting is not selection |\n| Stratify by significance, not by count | AS 2315.22 | Weight by revenue or affected accounts, not one ticket one vote |\n| Decide the threshold before you look | AS 2315.34 | Set the deviation rate you would accept in advance |\n\n## This is not the sampling question you have already read about\n\nYour existing sampling guides are about recruitment - who enters the study. [Purposive sampling](/docs/purposive-sampling-guide) is about strategic participant selection. [Qualitative sampling methods](/docs/qualitative-research-sampling-methods), [probability vs non-probability sampling](/docs/probability-vs-non-probability-sampling), [stratified sampling](/docs/stratified-sampling-guide) and [convenience sampling](/docs/convenience-sampling-guide) all answer the same question: how do you choose the people?\n\nThis article assumes the people are gone. The interviews happened. The tickets accumulated. The recordings exist. You now hold more evidence than anyone will ever read, and the question is which slice of it gets human attention. Nothing about recruitment sampling tells you how to make that choice defensibly, and the choice determines what your read-out is entitled to claim.\n\n## Rule 1: the census bucket and the sample bucket are different objects\n\nThis is the most important paragraph in the article.\n\nAS 2315.21 instructs the auditor to \"examine those items for which, in his judgment, acceptance of some sampling risk is not justified\" - the items that individually matter enough that you cannot afford to miss them. Then comes the crucial sentence: \"Any items that the auditor has decided to examine 100 percent are not part of the items subject to sampling.\"\n\nThey are removed from the population. They are reported separately. They are never used to estimate a rate.\n\nTranslate that. You should absolutely read all twelve enterprise churn tickets. You should watch every session where the user abandoned checkout with a full cart. Those are your census bucket, and reading them is good practice. But the moment you say \"and eight of the twelve mentioned the export limit, so about two thirds of churn is export-driven,\" you have projected a deliberately selected set onto a population, and the number is meaningless.\n\nThe fix is structural rather than statistical. Run two buckets:\n\n- **Census bucket.** Defined by a rule you can state in advance - every account above 50,000 dollars ARR, every severity-1 incident, every enterprise churn. Read all of them. Report findings as existence claims: this happens, here is what it looks like, here is a quote.\n- **Sample bucket.** Everything else, selected without regard to how interesting it looks. Report findings as rate claims: in this slice, X percent showed Y.\n\nTwo buckets, two registers, two kinds of sentence. Most read-outs blend them and produce a confident percentage assembled from the items somebody found compelling.\n\n## Rule 2: every item needs an opportunity to be selected\n\nAS 2315.24 states the requirement: \"Sample items should be selected in such a way that the sample can be expected to be representative of the population. Therefore, all items in the population should have an opportunity to be selected. For example, haphazard and random-based selection of items represents two means of obtaining such samples.\"\n\nISA 530 is more specific about what haphazard selection actually demands. It defines it as selecting without following a structured technique, while nonetheless avoiding conscious bias or predictability - explicitly naming the temptations: avoiding items that are difficult to locate, and always choosing or always avoiding the first or last entries on a page. And it adds a hard limit: haphazard selection is not appropriate when using statistical sampling.\n\nRead that list against how review actually happens on most teams. Somebody opens the ticket queue, sorts by most recent, reads until they have enough for the deck, and skips the ones with long threads because they take too long. Every single one of those moves is named in the standard as a bias to avoid. Sorting by recency is predictability. Skipping long threads is avoiding difficult-to-locate items - and long threads are systematically the hard cases.\n\nSo the ordering, from best to worst, is: random selection, then systematic selection with a random start, then genuine haphazard selection, then what most teams do, which is not haphazard at all. It is conscious selection dressed as convenience, and it has a direction: toward the recent, the short, the articulate, and the already-familiar.\n\n## Rule 3: stratify by significance, not by count\n\nAS 2315.22 notes that the auditor \"may be able to reduce the required sample size by separating items subject to sampling into relatively homogeneous groups on the basis of some characteristic related to the specific audit objective,\" and gives recorded or book value as a common basis.\n\nAuditing takes this further than most research does, through value-weighted approaches where an item's chance of selection scales with its monetary size. The logic is simply that a mistake in a large item matters more than a mistake in a small one.\n\nProduct research has an obvious analogue that almost nobody uses. One ticket, one vote is the default, and it is wrong whenever the items differ in weight. Weight selection probability by:\n\n- ARR or expected revenue of the account\n- Number of affected users or accounts behind the report\n- Severity or blocking status\n- Strategic segment, when the roadmap is committed to a segment\n\nThen read the stratum results separately rather than averaging them. A rate computed across an unstratified ticket population is dominated by whichever segment files the most tickets, which is usually the one with the most users and the least revenue.\n\n## Rule 4: decide the threshold before you look\n\nAS 2315.34 requires the auditor to \"determine the maximum rate of deviations from the prescribed control that he would be willing to accept without altering his planned assessed level of control risk. This is the tolerable rate.\" AS 2315.18 does the same for magnitude, calling it tolerable misstatement.\n\nThe point is temporal. The threshold is set before the reading starts, so the reading cannot influence it.\n\nIn review terms: before you open the first recording, write down the number that would change your decision. If more than 15 percent of sessions show users failing to find the filter, we redesign the filter. Without that line, review generates a rate, and the rate gets interpreted against whatever intuition is in the room, which is exactly the flexibility that produces conclusions from noise - the same family of problem covered in [p-hacking and researcher degrees of freedom](/docs/p-hacking-researcher-degrees-of-freedom).\n\nAS 2315.10 names the risk you are managing: \"Sampling risk arises from the possibility that, when a test of controls or a substantive test is restricted to a sample, the auditor's conclusions may be different from the conclusions he would reach if the test were applied in the same way to all items.\"\n\n## Describe or estimate: pick one before you select\n\nThe whole article compresses into one decision made before selection, because the selection method determines which of two things you are allowed to do afterwards.\n\n| You want to | Select by | You may say | You may not say |\n| --- | --- | --- | --- |\n| Describe what can happen | Judgment - pick the extreme, the severe, the strategic | This failure mode exists, here is the mechanism, here is who it hits | How common it is |\n| Estimate how often it happens | Random or systematic over a defined population | X percent of this population showed Y, plus or minus | That the vivid cases you also read are representative |\n\nBoth are legitimate research. Only one of them produces a percentage. The failure that recurs in read-out after read-out is doing the first and reporting the second.\n\n## Where this gets easier\n\nTwo things change the economics of review.\n\nThe first is that selection only matters when reading is scarce. When analysis runs automatically across every response rather than across the slice a human had time for, the census bucket expands until sampling becomes unnecessary for whole classes of question. Koji analyzes each interview as it completes and aggregates themes, quotes and quality scores across the full set, so the thematic layer is a census rather than a sample - see [real-time research insights](/docs/real-time-research-insights). Human review then goes where it is genuinely additive: the disconfirming cases, the outliers, the sessions the model scored as low quality.\n\nThe second is structural. Koji's structured questions remove the sampling problem from anything expressed as a closed question, because every participant answers every one. A scale question gives you a distribution over the whole study, not over the transcripts somebody read. single_choice and multiple_choice give complete frequency counts, ranking gives an average position across all respondents, and yes_no gives a clean denominator. open_ended questions still carry the depth, and those are where selective reading remains a live risk - which is precisely why the closed types are the right baseline to check your qualitative impressions against. See [the structured questions guide](/docs/structured-questions-guide).\n\nThe honest limit: none of this rescues a badly defined population. If your ticket corpus only contains tickets from customers who bother to file them, no selection method inside it will fix that, and you are looking at the sampling frame problem covered in [sampling bias](/docs/sampling-bias-research).\n\n## Common mistakes\n\n**Projecting from the interesting ones.** The single most common error, and the one AS 2315.21 exists to prevent. Selected-because-significant items go in a separate bucket with a separate register.\n\n**Sorting by recency and calling it a sample.** Recency is predictability. Newest-first review systematically over-weights whatever happened after your last release.\n\n**Skipping long threads and long recordings.** They are systematically the complicated cases, which is exactly why they take longer.\n\n**One ticket, one vote.** Weight by what matters - revenue, affected accounts, severity - or your rate describes whoever complains most.\n\n**Setting the threshold after reading.** A rate with no pre-committed threshold gets interpreted against the mood of the room.\n\n**Treating a saturated read as a complete one.** Reading until nothing new appears tells you about the frequent, not about the rare and severe. Keep a census rule for the severe.\n\n## Frequently asked questions\n\n### How is this different from participant sampling?\n\nParticipant sampling decides who enters a study. This decides what you read from evidence that already exists - tickets, recordings, transcripts, open-ended responses. The two are separate problems, and guides on purposive, stratified or convenience sampling address the first, not the second.\n\n### Can I read all the important cases and still report a percentage?\n\nNot from the same pool. Auditing standards remove items examined 100 percent from the sampling population entirely, precisely so they cannot be projected. Read every important case you like, but report those as existence claims and compute rates only from a separately selected sample.\n\n### Is haphazard selection acceptable?\n\nIt is acceptable for non-statistical work if it is genuinely haphazard - no structured technique but also no conscious bias, no avoiding hard-to-locate items, no always taking the first or last entries. It is explicitly not appropriate when using statistical sampling. Sorting by recency and reading until you have enough is neither haphazard nor acceptable.\n\n### How should I weight items of different importance?\n\nStratify by a characteristic tied to your objective - ARR, affected accounts, severity, segment - and report each stratum separately rather than averaging. Auditing commonly stratifies by recorded value on the logic that an error in a large item matters more, and product research has direct analogues.\n\n### What is a tolerable rate and why set it in advance?\n\nIt is the maximum deviation rate you would accept without changing your conclusion. Setting it before you start reading prevents the observed rate from being interpreted against whatever intuition happens to be in the room, which is how selective review turns into confident but unfounded conclusions.\n\n### Does this apply if my tool analyzes every response automatically?\n\nIt applies less, and that is the point. When analysis covers the full set, the thematic layer is a census and sampling is unnecessary for those questions. Selection still matters for whatever a human reads in depth, and it still matters for defining the population in the first place.\n\n## Related Resources\n\n- [Purposive sampling](/docs/purposive-sampling-guide) - the recruitment-side question this one deliberately does not answer\n- [Qualitative sampling methods](/docs/qualitative-research-sampling-methods) - choosing a participant sampling approach\n- [Sampling bias](/docs/sampling-bias-research) - what no selection method inside a bad population can fix\n- [Process controls vs output checks](/docs/research-process-controls-vs-output-checks) - how to pick the slice you reperform\n- [The structured questions guide](/docs/structured-questions-guide) - the six question types and why closed types need no sampling\n- [Product feedback triage](/docs/product-feedback-triage-guide) - turning the reviewed material into a prioritized backlog\n","category":"Analysis & Synthesis","lastModified":"2026-08-14T03:30:37.667939+00:00","metaTitle":"How to Choose What to Review: Sampling Tickets and Transcripts","metaDescription":"Selection rules for evidence you already own. Why items you picked because they looked significant can be described but never projected, and how to stratify review by significance.","keywords":["how to choose what to review","review sampling","support ticket sampling","session recording sample","audit sampling for research","haphazard selection","which transcripts to read"],"aiSummary":"There are two sampling problems in research and only recruitment is well covered. The second is selecting what to read from evidence you already own. Auditing formalized this as audit sampling (PCAOB AS 2315), and its central rule is that items examined 100 percent because they were individually significant are removed from the sampling population and cannot be projected - so the vivid cases you deliberately chose may be described but never used to compute a rate. Three further rules follow: every item needs an opportunity to be selected (recency sorting and skipping long threads are named biases), stratify by significance such as ARR or affected accounts rather than one item one vote, and set the tolerable deviation rate before reading begins.","aiPrerequisites":["Familiarity with basic sampling concepts","Access to a corpus of research artifacts such as tickets, recordings or transcripts"],"aiLearningOutcomes":["Separate a census bucket from a sample bucket and report each in its own register","Recognize conscious selection disguised as convenience review","Stratify review selection by significance rather than by item count","Set a tolerable deviation rate before reviewing begins","Decide whether a review is describing what can happen or estimating how often"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}