Back to docs
Analysis & Synthesis

How to Choose What to Review: Sampling Rules for Tickets, Recordings, and Transcripts You Already Have

Every sampling guide covers who to recruit. This one covers what to read from evidence you already own - and why selecting the interesting items means you can describe but never estimate.

Answer first: there are two completely different sampling problems in research, and the literature only covers one of them. The first is who to recruit, which every sampling guide addresses. The second is what to read from the mountain of evidence you already own - four thousand support tickets, three hundred session recordings, nine hundred open-ended responses, sixty transcripts. Auditing has spent decades formalizing the second problem, and the single most important rule it produced is this: the items you deliberately picked because they looked significant can be described, but they can never be projected to the population.

The short answer

Auditing calls it audit sampling, and PCAOB AS 2315.01 defines it as "the application of an audit procedure to less than 100 percent of the items within an account balance or class of transactions for the purpose of evaluating some characteristic of the balance or class."

Swap in your own nouns and it is exactly the problem you have on a Monday morning with a support inbox. The standard then supplies four rules that product teams almost never apply.

RuleThe standardWhat it means for your review
Split the census from the sampleAS 2315.21Items you chose to examine 100 percent are not part of the sample and cannot be projected
Give every item a chance of selectionAS 2315.24Scrolling until something looks interesting is not selection
Stratify by significance, not by countAS 2315.22Weight by revenue or affected accounts, not one ticket one vote
Decide the threshold before you lookAS 2315.34Set the deviation rate you would accept in advance

This is not the sampling question you have already read about

Your existing sampling guides are about recruitment - who enters the study. Purposive sampling is about strategic participant selection. Qualitative sampling methods, probability vs non-probability sampling, stratified sampling and convenience sampling all answer the same question: how do you choose the people?

This article assumes the people are gone. The interviews happened. The tickets accumulated. The recordings exist. You now hold more evidence than anyone will ever read, and the question is which slice of it gets human attention. Nothing about recruitment sampling tells you how to make that choice defensibly, and the choice determines what your read-out is entitled to claim.

Rule 1: the census bucket and the sample bucket are different objects

This is the most important paragraph in the article.

AS 2315.21 instructs the auditor to "examine those items for which, in his judgment, acceptance of some sampling risk is not justified" - the items that individually matter enough that you cannot afford to miss them. Then comes the crucial sentence: "Any items that the auditor has decided to examine 100 percent are not part of the items subject to sampling."

They are removed from the population. They are reported separately. They are never used to estimate a rate.

Translate that. You should absolutely read all twelve enterprise churn tickets. You should watch every session where the user abandoned checkout with a full cart. Those are your census bucket, and reading them is good practice. But the moment you say "and eight of the twelve mentioned the export limit, so about two thirds of churn is export-driven," you have projected a deliberately selected set onto a population, and the number is meaningless.

The fix is structural rather than statistical. Run two buckets:

  • Census bucket. Defined by a rule you can state in advance - every account above 50,000 dollars ARR, every severity-1 incident, every enterprise churn. Read all of them. Report findings as existence claims: this happens, here is what it looks like, here is a quote.
  • Sample bucket. Everything else, selected without regard to how interesting it looks. Report findings as rate claims: in this slice, X percent showed Y.

Two buckets, two registers, two kinds of sentence. Most read-outs blend them and produce a confident percentage assembled from the items somebody found compelling.

Rule 2: every item needs an opportunity to be selected

AS 2315.24 states the requirement: "Sample items should be selected in such a way that the sample can be expected to be representative of the population. Therefore, all items in the population should have an opportunity to be selected. For example, haphazard and random-based selection of items represents two means of obtaining such samples."

ISA 530 is more specific about what haphazard selection actually demands. It defines it as selecting without following a structured technique, while nonetheless avoiding conscious bias or predictability - explicitly naming the temptations: avoiding items that are difficult to locate, and always choosing or always avoiding the first or last entries on a page. And it adds a hard limit: haphazard selection is not appropriate when using statistical sampling.

Read that list against how review actually happens on most teams. Somebody opens the ticket queue, sorts by most recent, reads until they have enough for the deck, and skips the ones with long threads because they take too long. Every single one of those moves is named in the standard as a bias to avoid. Sorting by recency is predictability. Skipping long threads is avoiding difficult-to-locate items - and long threads are systematically the hard cases.

So the ordering, from best to worst, is: random selection, then systematic selection with a random start, then genuine haphazard selection, then what most teams do, which is not haphazard at all. It is conscious selection dressed as convenience, and it has a direction: toward the recent, the short, the articulate, and the already-familiar.

Rule 3: stratify by significance, not by count

AS 2315.22 notes that the auditor "may be able to reduce the required sample size by separating items subject to sampling into relatively homogeneous groups on the basis of some characteristic related to the specific audit objective," and gives recorded or book value as a common basis.

Auditing takes this further than most research does, through value-weighted approaches where an item's chance of selection scales with its monetary size. The logic is simply that a mistake in a large item matters more than a mistake in a small one.

Product research has an obvious analogue that almost nobody uses. One ticket, one vote is the default, and it is wrong whenever the items differ in weight. Weight selection probability by:

  • ARR or expected revenue of the account
  • Number of affected users or accounts behind the report
  • Severity or blocking status
  • Strategic segment, when the roadmap is committed to a segment

Then read the stratum results separately rather than averaging them. A rate computed across an unstratified ticket population is dominated by whichever segment files the most tickets, which is usually the one with the most users and the least revenue.

Rule 4: decide the threshold before you look

AS 2315.34 requires the auditor to "determine the maximum rate of deviations from the prescribed control that he would be willing to accept without altering his planned assessed level of control risk. This is the tolerable rate." AS 2315.18 does the same for magnitude, calling it tolerable misstatement.

The point is temporal. The threshold is set before the reading starts, so the reading cannot influence it.

In review terms: before you open the first recording, write down the number that would change your decision. If more than 15 percent of sessions show users failing to find the filter, we redesign the filter. Without that line, review generates a rate, and the rate gets interpreted against whatever intuition is in the room, which is exactly the flexibility that produces conclusions from noise - the same family of problem covered in p-hacking and researcher degrees of freedom.

AS 2315.10 names the risk you are managing: "Sampling risk arises from the possibility that, when a test of controls or a substantive test is restricted to a sample, the auditor's conclusions may be different from the conclusions he would reach if the test were applied in the same way to all items."

Describe or estimate: pick one before you select

The whole article compresses into one decision made before selection, because the selection method determines which of two things you are allowed to do afterwards.

You want toSelect byYou may sayYou may not say
Describe what can happenJudgment - pick the extreme, the severe, the strategicThis failure mode exists, here is the mechanism, here is who it hitsHow common it is
Estimate how often it happensRandom or systematic over a defined populationX percent of this population showed Y, plus or minusThat the vivid cases you also read are representative

Both are legitimate research. Only one of them produces a percentage. The failure that recurs in read-out after read-out is doing the first and reporting the second.

Where this gets easier

Two things change the economics of review.

The first is that selection only matters when reading is scarce. When analysis runs automatically across every response rather than across the slice a human had time for, the census bucket expands until sampling becomes unnecessary for whole classes of question. Koji analyzes each interview as it completes and aggregates themes, quotes and quality scores across the full set, so the thematic layer is a census rather than a sample - see real-time research insights. Human review then goes where it is genuinely additive: the disconfirming cases, the outliers, the sessions the model scored as low quality.

The second is structural. Koji's structured questions remove the sampling problem from anything expressed as a closed question, because every participant answers every one. A scale question gives you a distribution over the whole study, not over the transcripts somebody read. single_choice and multiple_choice give complete frequency counts, ranking gives an average position across all respondents, and yes_no gives a clean denominator. open_ended questions still carry the depth, and those are where selective reading remains a live risk - which is precisely why the closed types are the right baseline to check your qualitative impressions against. See the structured questions guide.

The honest limit: none of this rescues a badly defined population. If your ticket corpus only contains tickets from customers who bother to file them, no selection method inside it will fix that, and you are looking at the sampling frame problem covered in sampling bias.

Common mistakes

Projecting from the interesting ones. The single most common error, and the one AS 2315.21 exists to prevent. Selected-because-significant items go in a separate bucket with a separate register.

Sorting by recency and calling it a sample. Recency is predictability. Newest-first review systematically over-weights whatever happened after your last release.

Skipping long threads and long recordings. They are systematically the complicated cases, which is exactly why they take longer.

One ticket, one vote. Weight by what matters - revenue, affected accounts, severity - or your rate describes whoever complains most.

Setting the threshold after reading. A rate with no pre-committed threshold gets interpreted against the mood of the room.

Treating a saturated read as a complete one. Reading until nothing new appears tells you about the frequent, not about the rare and severe. Keep a census rule for the severe.

Frequently asked questions

How is this different from participant sampling?

Participant sampling decides who enters a study. This decides what you read from evidence that already exists - tickets, recordings, transcripts, open-ended responses. The two are separate problems, and guides on purposive, stratified or convenience sampling address the first, not the second.

Can I read all the important cases and still report a percentage?

Not from the same pool. Auditing standards remove items examined 100 percent from the sampling population entirely, precisely so they cannot be projected. Read every important case you like, but report those as existence claims and compute rates only from a separately selected sample.

Is haphazard selection acceptable?

It is acceptable for non-statistical work if it is genuinely haphazard - no structured technique but also no conscious bias, no avoiding hard-to-locate items, no always taking the first or last entries. It is explicitly not appropriate when using statistical sampling. Sorting by recency and reading until you have enough is neither haphazard nor acceptable.

How should I weight items of different importance?

Stratify by a characteristic tied to your objective - ARR, affected accounts, severity, segment - and report each stratum separately rather than averaging. Auditing commonly stratifies by recorded value on the logic that an error in a large item matters more, and product research has direct analogues.

What is a tolerable rate and why set it in advance?

It is the maximum deviation rate you would accept without changing your conclusion. Setting it before you start reading prevents the observed rate from being interpreted against whatever intuition happens to be in the room, which is how selective review turns into confident but unfounded conclusions.

Does this apply if my tool analyzes every response automatically?

It applies less, and that is the point. When analysis covers the full set, the thematic layer is a census and sampling is unnecessary for those questions. Selection still matters for whatever a human reads in depth, and it still matters for defining the population in the first place.

Related Resources

Related Articles

Probability vs Non-Probability Sampling: Methods, Examples & When to Use Each

A clear guide to probability and non-probability sampling — the two families of sampling methods. Learn the types (random, stratified, convenience, purposive, quota, snowball), the trade-off between generalizability and speed, and how to recruit the right participants.

Product Feedback Triage: A Framework for Turning Noise Into a Prioritized Backlog

A practical framework for triaging product feedback at scale — capture, dedupe, tag, route, and validate every request before it ever reaches prioritization. Includes a triage workflow, a severity matrix, and an AI-native approach.

Purposive Sampling: The Complete Guide to Strategic Participant Selection

A complete guide to purposive (purposeful) sampling in qualitative research — covering all major types, when to use each, how to determine sample size, and how AI tools enable purposive sampling at scale.

Qualified, Adverse, or Disclaimed: How to Report a Study That Did Not Go to Plan

Research has one register for five different failures - a bullet in the limitations section. Auditing built a named ladder with explicit triggers. Here is how to port it, and when the honest output is to decline to conclude.

Sampling Methods in Qualitative Research: A Complete Guide for Choosing the Right Approach (2026)

Master the eight sampling methods used in qualitative research — purposive, theoretical, snowball, convenience, quota, criterion, maximum variation, and homogeneous. Learn when to use each, how to combine them, and how to determine sample size.

Process Controls vs Output Checks: How to Earn the Right to Read Fewer Transcripts

Evidence that your research process worked substitutes for evidence about each individual output. The trade auditors formalized, why existence is not operation, and how reperformance proves a control actually ran.

Sampling Bias: Types, Examples, and How to Avoid It

Sampling bias is when some people in your population are systematically more likely to end up in your sample than others — quietly invalidating your findings. Learn the six main types, classic examples, and how to build a representative sample at scale.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.