{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-06T14:10:03.035Z"},"content":[{"type":"documentation","id":"5f299505-11f6-40a0-b9a3-ff4ce8c0a7e6","slug":"theme-co-occurrence-collocation-analysis","title":"Theme Co-Occurrence: Finding the Problems That Travel Together (2026)","url":"https://www.koji.so/docs/theme-co-occurrence-collocation-analysis","summary":"Theme pairs are often more actionable than single themes, but raw co-occurrence counts rank theme popularity rather than theme relationships: two themes in 40 and 30 of 60 interviews co-occur in about 20 by chance alone. Build a 2x2 table over participants or accounts and report observed divided by expected, where 1.0 means independence. Gries 2021 supports the log odds ratio as an association-only measure reported separately from frequency. Apply a support floor, and recompute within segment, because the most common false positive is a shared population rather than a real relationship.","content":"**Bottom line up front:** Themes rarely arrive alone, and the pair that shows up together more often than chance allows is usually closer to the real problem than either theme on its own. \"Slow exports\" plus \"end of month\" is a reporting-deadline problem; \"slow exports\" plus \"large accounts\" is a scaling problem. Do not rank pairs by how often they co-occur, because that ranking is mostly driven by how common each theme is individually. Compare **observed co-occurrence against expected co-occurrence**, keep the association statistic in a separate column from the frequency, and count each pair once per participant so that one person discussing both topics does not manufacture a pattern.\n\n## Why the pair is often the finding\n\nA single theme is usually too coarse to act on. \"Confusing pricing\" is not a task. But \"confusing pricing\" co-occurring with \"annual renewal\" in nineteen of the twenty-two interviews that mention either one is a specific, fixable thing: your pricing page does not explain what happens at renewal.\n\nCorpus linguistics has a long-established name for this. **Collocation** is the tendency of two items to appear together more often than their individual frequencies would predict. Kenneth Church and Patrick Hanks brought the idea into computational lexicography at the 27th Annual Meeting of the Association for Computational Linguistics, using mutual information to identify word pairs that belong together. The same arithmetic works on coded themes, and for feedback analysis it is more useful on themes than on words, because themes are already the unit your roadmap argues about.\n\n## The trap: common themes co-occur by accident\n\nSuppose 60 interviews. Theme X appears in 40 of them. Theme Y appears in 30. How often would you expect X and Y to show up in the same interview purely by chance?\n\nRoughly 40/60 times 30/60 times 60, which is 20 interviews.\n\nSo an observed co-occurrence of 20 is *exactly nothing*. It is what independence predicts. Yet 20 shared interviews out of 60 looks like a striking pattern in a dashboard, and it will be the top row of any table sorted by raw co-occurrence count, simply because X and Y are the two most common themes in the corpus.\n\nThis is the whole problem with co-occurrence counts: **the most frequent themes dominate the top of the list whether or not they are related.** Sorting pairs by shared-interview count produces a ranking of theme popularity wearing the costume of a ranking of relationships.\n\n## The 2x2 table and what to compute from it\n\nFor any theme pair, build the table over your chosen unit (participants, not mentions - more on this below):\n\n| | Has theme Y | No theme Y | Total |\n|---|---|---|---|\n| Has theme X | a | b | a+b |\n| No theme X | c | d | c+d |\n| Total | a+c | b+d | n |\n\nTwo numbers are worth reporting:\n\n**1. Observed over expected.** Expected *a* under independence is (a+b)(a+c)/n. The ratio of observed *a* to that expected value is immediately interpretable: 1.0 means independence, 2.0 means the pair appears twice as often together as chance predicts, below 1.0 means the themes repel each other. This single ratio is enough for most product conversations.\n\n**2. The log odds ratio.** log( (a times d) / (b times c) ). This is the association measure Gries singles out in a 2021 paper in the *Journal of Second Language Studies*, where he examined the field's most widely used association measures and argued that most are not particularly valid, in that they measure an amalgam of a lot of frequency and a little association rather than association itself. His conclusion supports the log odds ratio as a genuine association-only measure, used **separately from frequency** rather than blended with it.\n\nThat last clause is the operational instruction. Report the association and the frequency as two columns. A composite \"strength score\" that folds them together reintroduces exactly the ambiguity you computed the association to escape - the same objection Gries raises against blended dispersion measures.\n\nTwo older measures you will meet, and their biases:\n\n- **Mutual information** rewards rare pairs heavily. Two themes that each appear twice and happen to co-occur both times will score enormously. Useful for discovery, dangerous for ranking.\n- **Log-likelihood** leans the other way and favours frequent pairs, which is the bias you were trying to remove.\n\nUse observed-over-expected plus a minimum support threshold (say, the pair must co-occur in at least five participants) and you will avoid both failure modes without needing to defend a choice of statistic in a roadmap meeting.\n\n## A worked example\n\nSixty interviews, four candidate pairs.\n\n| Pair | Co-occurring interviews | Expected | Obs/Exp | Verdict |\n|---|---|---|---|---|\n| slow exports + end of month | 14 | 4.7 | 3.0 | Real, and specific |\n| confusing pricing + annual renewal | 19 | 8.3 | 2.3 | Real |\n| slow exports + confusing pricing | 20 | 20.0 | 1.0 | Nothing. Both are just common |\n| custom SSO + mobile app | 2 | 0.6 | 3.3 | High ratio, too thin to act on |\n\nSorted by raw co-occurrence, row three wins and you spend a sprint hunting a relationship that does not exist. Sorted by observed over expected, row four wins and you spend a sprint on two customers. Sorted by observed over expected **with a support floor**, rows one and two surface, which is the correct answer.\n\nRow one is the kind of finding that changes a roadmap. Slow exports concentrated at month end is not a performance project, it is a capacity-at-peak project, and the pair is what told you.\n\n## Count once per participant, not once per mention\n\nThis is where theme co-occurrence quietly breaks, and it is the same unit discipline that governs [theme dispersion](/docs/theme-dispersion-vs-mention-frequency).\n\nIf you count co-occurrence at the level of mentions, or within some sliding window of transcript text, you are measuring something close to \"did one person talk about both of these in the same breath\". One participant who spends ten minutes connecting two topics can generate dozens of co-occurrence events. The pair then looks strongly associated on the strength of a single conversation.\n\nCount instead at the level of the unit that matches your decision:\n\n- **Per participant** - the general default. Did this person raise both themes at all?\n- **Per account** - correct for B2B roadmap decisions, because the account is the renewing unit.\n- **Per interview** - acceptable only when each participant gave exactly one interview.\n\nAnd treat **within-interview adjacency** as a separate, qualitative signal. If two themes keep appearing in the same breath, that is a strong hint about mechanism and worth reading, but it is evidence about how people reason, not about how widespread the pairing is. Keep the two questions apart.\n\n## Three readings of a real association\n\nA confirmed pair is a starting point for interpretation, not a conclusion. There are usually three candidate explanations, and distinguishing them requires reading the passages:\n\n1. **Shared cause.** Both themes are symptoms of one underlying thing. Fix the cause and both disappear. The most valuable case.\n2. **Sequence.** One theme leads to the other - a failed import causes a support ticket. The intervention point is upstream, and the second theme is a consequence you should expect to see fall.\n3. **Shared population.** Both themes belong to the same group of people for unrelated reasons. Enterprise customers mention SSO and procurement because they are enterprise customers, not because the two are connected. This is the most common false positive, and the giveaway is that the association vanishes when you condition on segment.\n\nReading three is why a co-occurrence table should always be checked within segment before it is believed. If a pair is strong overall but absent inside every individual segment, you have found a segment, not a relationship.\n\n## How Koji handles this\n\nTheme co-occurrence needs three things: themes extracted consistently, participants kept attached to their themes, and clean attributes to condition on. Koji supplies all three.\n\n- **Consistent theme extraction across the whole corpus.** Koji's analysis assigns themes using the same logic for every interview, so a pair's association reflects the participants rather than which analyst coded which batch on which day. Inconsistent human coding is a major source of phantom associations, because coders drift toward pairing concepts they have recently seen.\n- **Stable question IDs preserve the unit.** Koji's study questions carry stable identifiers that maintain traceability from the interview plan, through the AI interviewer, into analysis and report aggregation. Each theme stays attached to its participant and the question that produced it, so switching your co-occurrence unit from interview to participant to account is a grouping decision rather than a reconstruction effort.\n- **Structured questions give you the conditioning variables.** Koji's six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - deliver plan, segment, role and tenure as structured values. That is precisely what you need to run the within-segment check that kills the shared-population false positive. The [structured questions guide](/docs/structured-questions-guide) explains how each type is asked and aggregated.\n- **AI follow-up questions surface the mechanism, not just the pair.** When a participant raises two related themes, Koji's interviewer probes the connection in the moment. A statistic can tell you two themes travel together; a follow-up question asked at the right second tells you why, and that is the difference between a correlation and an actionable finding.\n- **Insights chat for the reading step.** Confirming which of the three readings applies means reading the passages where both themes appear. Koji lets you ask for exactly those passages conversationally rather than grepping transcripts.\n\n## Common mistakes\n\n- **Ranking pairs by raw co-occurrence count.** Produces a popularity ranking, not a relationship ranking.\n- **Using mutual information to rank.** It rewards rare pairs, so your top row will be two themes that appeared twice.\n- **Blending association and frequency into one score.** Gries's objection: you get mostly frequency in disguise.\n- **Counting co-occurrence per mention.** One talkative participant becomes a pattern.\n- **Skipping the within-segment check.** Most strong pairs are a segment, not a relationship.\n- **Treating a pair as causal.** Shared cause, sequence and shared population all look identical in the table.\n\n## Frequently asked questions\n\n### What is theme co-occurrence analysis?\n\nTheme co-occurrence analysis identifies pairs of themes that appear together in the same interviews, accounts or segments more often than their individual frequencies would predict. It is the applied form of collocation analysis from corpus linguistics, where Church and Hanks established the approach for word pairs. The value is that a pair is often more specific and more actionable than either theme alone, because it narrows a vague complaint into a particular situation.\n\n### Why can I not just count how often two themes appear together?\n\nBecause that count is dominated by how common each theme is individually. If one theme appears in 40 of 60 interviews and another in 30, chance alone puts them together in about 20 interviews, so an observed count of 20 means nothing at all. Ranking pairs by raw co-occurrence therefore ranks theme popularity rather than theme relationships, and your two most common themes will always top the list.\n\n### Which association measure should I use for theme pairs?\n\nFor most product decisions, observed co-occurrence divided by expected co-occurrence is enough, because it reads directly as a multiple of chance: 1.0 is independence and 2.0 is twice chance. If you want a formal statistic, Gries's 2021 review of association measures supports the log odds ratio as a true association-only measure, provided you report it separately from frequency rather than blending the two. Avoid ranking by mutual information, which heavily rewards rare pairs.\n\n### Should I count co-occurrence per mention or per participant?\n\nPer participant, as a default, and per account for B2B roadmap decisions. Counting per mention lets one person who discusses two topics at length generate dozens of co-occurrence events, which makes a single conversation look like a widespread pattern. Within-interview adjacency is still worth reading as a qualitative hint about mechanism, but it should be kept separate from any estimate of how common the pairing is.\n\n### How do I tell whether a theme pair is causal?\n\nYou cannot tell from the table, and that is the point. A real association has three common explanations: a shared underlying cause, a sequence where one theme produces the other, or a shared population where both themes simply belong to the same group for unrelated reasons. The shared-population case is the most frequent false positive, and the diagnostic is to recompute the association within each segment. If the pair is strong overall but absent inside every segment, you have found a segment rather than a relationship.\n\n### Can Koji find theme pairs automatically?\n\nKoji extracts themes consistently across the entire corpus using the same logic for every interview, and keeps each theme attached to its participant and originating question through stable question IDs, which is what makes a reliable co-occurrence table possible. Koji's structured question types then supply the segment and plan attributes needed for the within-segment check that separates real pairs from segment artefacts, and insights chat retrieves the passages where both themes appear so you can judge which of the three readings applies.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types that supply the attributes you condition on to kill false pairs\n- [Theme Dispersion vs Mention Frequency](/docs/theme-dispersion-vs-mention-frequency) - the unit discipline that co-occurrence analysis depends on\n- [Keyness: Distinctive Segment Language](/docs/keyness-analysis-distinctive-segment-language) - the term-level counterpart, comparing two groups rather than two themes\n- [Affinity Mapping](/docs/affinity-mapping) - the manual, qualitative route to the same groupings\n- [Customer Feedback Categorization](/docs/customer-feedback-categorization) - building the taxonomy whose themes you are pairing\n- [Coding Qualitative Data](/docs/coding-qualitative-data) - how the coded units that feed this analysis are produced\n","category":"Analysis & Synthesis","lastModified":"2026-10-06T07:51:48.571184+00:00","metaTitle":"Theme Co-Occurrence and Collocation Analysis (2026)","metaDescription":"Two common themes co-occur by chance alone. Rank theme pairs by observed over expected with a support floor, and count once per participant.","keywords":["theme co-occurrence","feedback themes that appear together","collocation analysis","theme correlation interviews","log odds ratio association","co-occurrence analysis customer feedback"],"aiSummary":"Theme pairs are often more actionable than single themes, but raw co-occurrence counts rank theme popularity rather than theme relationships: two themes in 40 and 30 of 60 interviews co-occur in about 20 by chance alone. Build a 2x2 table over participants or accounts and report observed divided by expected, where 1.0 means independence. Gries 2021 supports the log odds ratio as an association-only measure reported separately from frequency. Apply a support floor, and recompute within segment, because the most common false positive is a shared population rather than a real relationship.","aiPrerequisites":["A corpus with themes coded per participant","Segment or plan attributes to condition on"],"aiLearningOutcomes":["Compute expected co-occurrence and the observed-over-expected ratio","Choose between observed/expected, log odds ratio and mutual information","Count co-occurrence at the correct unit","Distinguish shared cause, sequence and shared population"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}