{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-09T12:08:29.788Z"},"content":[{"type":"documentation","id":"779b548c-b4d8-4531-9b5b-c9192059002d","slug":"evidence-synthesis-research-findings","title":"Evidence Synthesis: How to Combine Findings Across Multiple Research Studies (2026)","url":"https://www.koji.so/docs/evidence-synthesis-research-findings","summary":"Evidence synthesis combines findings from multiple studies into one conclusion with an explicit confidence rating. It is distinct from triangulation (methods within one question) and from repository storage (retrieval). GRADE rates a body of evidence at four certainty levels with five downgrade domains - risk of bias, inconsistency, indirectness, imprecision, publication bias - and three upgrade criteria. The main obstacle in product research is instrument drift, where the same question was operationalised differently each time so nothing pools; the fix is a small standing instrument reused verbatim across studies in a domain, plus organising the evidence base by standing decision question rather than by study.","content":"Evidence synthesis is the practice of combining findings from multiple separate studies into one conclusion, with an explicit rating of how much confidence that conclusion deserves. It is not the same as storing studies in a repository, and it is not the same as triangulating methods within a single study. Medicine formalised it decades ago through systematic review, the PRISMA reporting standard (Page et al., *BMJ*, 2021, 372:n71, a 27-item checklist) and the GRADE framework for rating certainty of evidence (Guyatt et al., *BMJ*, 2008, 336:924-926). Product research has the studies and almost never has the synthesis, which is why teams with sixty studies in a repository still answer strategic questions from whichever one someone remembers.\n\n**Key takeaways**\n\n- A repository answers \"what did study 47 find\". A synthesis answers \"what do we collectively believe about activation, and how sure are we\". Most teams have built the first and assume it delivers the second.\n- GRADE rates a body of evidence at one of four certainty levels and specifies five reasons to downgrade - risk of bias, inconsistency, indirectness, imprecision and publication bias - plus three reasons to upgrade.\n- Publication bias is one of GRADE's five downgrade domains, which means a synthesis is only as trustworthy as the completeness of the record it draws on.\n- The reason internal synthesis usually fails is not analytical. It is instrument drift: every study asked the question slightly differently, so nothing pools.\n- Organise the evidence base by standing decision question, not by study, and the synthesis becomes a maintained artifact rather than a quarterly heroic effort.\n\n## Three things that get called synthesis\n\nThe word is used for three genuinely different operations, and conflating them is why teams think they are doing this already.\n\n| Operation | Question it answers | Scope | When to use |\n|---|---|---|---|\n| **Triangulation** | Do different methods agree about this one question, right now? | Multiple methods, one study or programme, one point in time | Validating a single finding before acting on it |\n| **Conflict resolution** | Two datasets disagree - which is right, and why? | Two specific sources | A qual finding contradicts a quant one |\n| **Evidence synthesis** | What does everything we have ever learned about this question add up to, and how confident should we be? | Many studies, many methods, across time | Strategy, prioritisation, onboarding, any decision bigger than one study |\n\nOur guide to [triangulation in research](/docs/triangulation-in-research-guide) covers the first, and [conflicting research findings](/docs/conflicting-research-findings) covers the second. This guide covers the third, which is the one almost nobody does deliberately.\n\nThe distinction matters because the three have different failure modes. Triangulation fails when methods share a bias. Conflict resolution fails when you pick the source that agrees with you. **Synthesis fails when the corpus you are synthesising was assembled by selection rather than by design** - and that failure is invisible from inside the corpus.\n\n## Why a repository is not a synthesis\n\nResearch repositories solve retrieval. They are genuinely valuable, and our guides to [building a research repository](/docs/research-repository-guide) and [insight repository methodology](/docs/insight-repository-methodology) cover how to do it well. But retrieval and synthesis are different products of different work.\n\nA repository is organised by study: here is what we ran, when, with whom, and what it found. A synthesis is organised by question: here is what we believe about why trial users churn in week two, assembled from nine studies over three years, rated moderate confidence, with the two studies that disagree flagged and explained.\n\nThe gap between them is judgment, and judgment does not accumulate automatically. A repository with sixty studies and no synthesis layer produces a specific and recognisable pathology: every planning cycle, somebody searches the repository, finds four relevant studies, reads the two with the best titles, and forms a view. The other fifty-six might as well not exist. **Storage without synthesis does not preserve institutional knowledge, it preserves the raw materials of institutional knowledge and quietly transfers the assembly cost to whoever is in the most hurry.**\n\n## The instrument drift problem\n\nBefore the method, the obstacle - because this is the one that actually stops teams, and it is rarely named.\n\nMedicine can pool trials because they measure the same outcomes in comparable ways. Product research usually cannot, because the same question was operationalised differently every time it was asked. Consider a team that has studied onboarding four times:\n\n- 2024 Q1: five-point satisfaction scale on the setup experience\n- 2024 Q3: open-ended \"what was hardest about getting started\"\n- 2025 Q2: task completion rate in a usability test\n- 2026 Q1: seven-point ease-of-use scale, different wording, different anchors\n\nEvery one of these is a reasonable study. Collectively they cannot be pooled on any single metric, because there is no metric they share. The team does not have four studies of onboarding. It has four studies of four different things that all mention onboarding.\n\nThe fix is a standing instrument: a small set of measures that every study touching a given domain includes verbatim, regardless of what else it asks. Two or three items is enough. The cost is a few extra questions per study; the return is that the fifth study can be compared to the first four, which is the entire precondition for synthesis. Our guide to [survey design best practices](/docs/survey-design-best-practices) covers the item-writing side, and the discipline of reusing exact wording matters more here than the elegance of the wording itself.\n\n## The confidence ledger: GRADE adapted for product research\n\nGRADE is the most widely adopted framework for rating a body of evidence, and its core structure transfers to commercial research almost unchanged. It works by assigning a starting level based on study design, then adjusting.\n\n**The four certainty levels**, restated for product decisions:\n\n- **High** - further research is unlikely to change the conclusion. Act on it, including on irreversible decisions.\n- **Moderate** - further research could change the conclusion. Act, but instrument the outcome and be prepared to revise.\n- **Low** - further research is likely to change the conclusion. Use it to choose what to try, not what to commit to.\n- **Very low** - the conclusion is highly uncertain. It is a hypothesis, and should be labelled as one in every document that cites it.\n\n**The five reasons to rate down.** Each is a published GRADE guideline in its own right, and each has a direct product-research reading:\n\n| GRADE domain | The product research version |\n|---|---|\n| Risk of bias | Leading questions, convenience samples, moderator effects, respondents recruited from your happiest cohort |\n| Inconsistency | The studies disagree with each other and you cannot explain why from their design |\n| Indirectness | You studied a proxy population, a proxy behaviour or a prototype rather than the real thing |\n| Imprecision | Small samples, wide intervals, or nulls that were never equivalence-tested |\n| Publication bias | You cannot demonstrate that the studies you are pooling are all the studies that were run |\n\nNote the last row carefully. Publication bias is not an afterthought in GRADE; it is one of the five load-bearing reasons to distrust a body of evidence. If your organisation has no register of commissioned studies, then by GRADE's own logic every internal synthesis you produce should be rated down at least one level on that domain alone. That is the practical link between this guide and [publication bias in product research](/docs/publication-bias-product-research), and it is not a rhetorical one.\n\nThe imprecision row carries a similar dependency. A body of evidence containing three flat results is only reassuring if those flat results were genuinely null rather than merely underpowered - a distinction that requires the methods in [equivalence testing](/docs/equivalence-testing-no-difference). Pooling uninformative nulls as if they were evidence of absence is one of the most common ways an internal synthesis reaches a confident wrong answer.\n\n**The three reasons to rate up**, which apply when the evidence is observational but unusually compelling: a large magnitude of effect, a dose-response gradient, and situations where all plausible confounding would work against the effect you observed and it appeared anyway. The third is the most useful in commercial settings, because it describes the common case where every incentive in the study design pushed toward a flattering answer and the unflattering answer showed up regardless.\n\n## A working procedure\n\nSystematic review methodology is heavy by design, because the stakes in medicine justify it. The transferable core is lighter and fits a week rather than a year.\n\n**1. Frame one standing question.** Not \"what do we know about onboarding\" but \"what causes new teams to abandon setup before inviting a second user\". A synthesis needs a question specific enough to judge relevance against.\n\n**2. Search the register, not your memory.** List every study that touched the question, including the ones nobody presented. This step is the difference between synthesis and a literature review of your own greatest hits. If the register does not exist, note that the search was incomplete and rate down for publication bias.\n\n**3. Record what each study can and cannot speak to.** Population, method, sample size, instrument, date. Most disagreements between internal studies dissolve at this step: two studies \"contradicting\" each other frequently turn out to have sampled different segments a year apart.\n\n**4. Pool where the instruments allow it, and narrate where they do not.** Quantitative synthesis is preferable when it is available. Where it is not - and it usually is not - a structured narrative synthesis that states each study's contribution and weight is still vastly better than an unstated one.\n\n**5. Rate the body of evidence.** Assign one of the four levels using the domains above, and write down the reason for each downgrade. The reasons are more valuable than the level, because they are a research agenda: a conclusion rated down for imprecision tells you the next study should be bigger, not different.\n\n**6. Record the dissent.** Every synthesis should name the studies that disagree with its conclusion and say why they were weighted as they were. A synthesis with no dissenting evidence listed is usually a synthesis that stopped searching early.\n\n**7. Date it and set a review trigger.** A synthesis is a snapshot. Tie its expiry to an event - the next major release, a pricing change, entry into a new segment - rather than to a calendar interval nobody will honour.\n\n## How Koji makes synthesis a by-product rather than a project\n\n**A shared vocabulary across studies.** The instrument drift problem is structural, and it is solved structurally. Koji computes themes across studies in a common vocabulary, so a theme surfaced in a discovery interview study in March is directly comparable with the same theme in a validation study in September. Pooling qualitative findings stops requiring a human to read both reports and decide they are talking about the same thing.\n\n**Reusable structured instruments.** The six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - documented in our [structured questions guide](/docs/structured-questions-guide), make it trivial to carry the same two or three standing items into every study in a domain. Identical question type plus identical wording produces identical data shape, which is the only thing pooling actually requires. Legacy tools like SurveyMonkey let you duplicate a survey, but they do not carry an analysis vocabulary across studies, so synthesis remains a manual reading exercise.\n\n**Every study leaves a record.** Synthesis fails most often because the corpus is incomplete, and the corpus is incomplete because unremarkable studies never got written up. When reports are generated automatically as responses arrive, the flat study is in the corpus by default. This is the single largest improvement available to most teams, and it is not an analytical one.\n\n**Filling the gap is a days-long task.** The most useful output of a synthesis is usually a specific hole: we have four studies on this question and none of them sampled churned users. Under traditional research economics that gap stays open for a quarter. When an AI-moderated study can be fielded and analysed in days at roughly $20 per voice conversation - against $20,000 to $50,000 for a comparable manual programme - the synthesis stops being a retrospective document and becomes the thing that commissions the next study.\n\n**Research literacy stops being the bottleneck.** GRADE-style rating is not hard, but it has historically required someone with methods training to run it. When quality scoring, sample composition and theme prevalence are computed and visible on every study, a product manager can apply the five downgrade domains without a background in research methods. Our guide to [research democratization](/docs/research-democratization-playbook) covers the broader shift.\n\n## Common failure modes\n\n**Vote counting.** Concluding that four studies found an effect and two did not, therefore the effect is real. Study quality and sample size vary enormously; counting treats a 12-person session and a 900-person survey as equal votes. Weight, or at minimum describe, rather than count.\n\n**Synthesising only the studies that are easy to find.** In practice this means the ones with good decks. It is publication bias operating one level up, and it produces syntheses that are systematically more confident than the underlying evidence.\n\n**Treating recency as quality.** Newer studies are often more relevant, but not always better. A 2023 study of 400 users beats a 2026 study of 15, unless the product changed in a way that makes the older population irrelevant - and that judgment should be written down rather than assumed.\n\n**Confusing consistency with truth.** Studies that agree may share a common flaw. If four studies all recruited from the in-app feedback widget, their agreement tells you about that channel, not about your users. See [sampling bias](/docs/sampling-bias-research) for the mechanism.\n\n**Stopping at the rating.** A synthesis that produces \"moderate confidence\" and nothing else has done half the job. The downgrade reasons are the deliverable, because they say exactly what the next study should fix.\n\n## Frequently asked questions\n\n### How many studies do I need before synthesis is worth doing?\nTwo. The value is not in statistical pooling, which needs many more, but in the discipline of stating what the studies collectively support and where they diverge. Teams tend to wait for a corpus large enough to feel like a systematic review, by which point the habit has never formed. Synthesise at two studies and maintain it as the third and fourth arrive.\n\n### Can I do quantitative meta-analysis on internal studies?\nOccasionally, and only where instruments genuinely match. Formal meta-analysis needs comparable effect sizes and their standard errors from studies measuring the same outcome, which a typical product research corpus does not provide. Structured narrative synthesis with explicit weighting is the realistic method for almost all internal work, and it captures most of the benefit.\n\n### How is this different from triangulation?\nTriangulation combines methods to check one finding at one point in time. Synthesis combines studies to answer a standing question across time. A single study can be triangulated; a body of work is synthesised. They also fail differently, since triangulation is vulnerable to methods sharing a bias while synthesis is vulnerable to the corpus having been selectively assembled.\n\n### Should qualitative and quantitative studies go into the same synthesis?\nYes, with their roles distinguished. Qualitative studies establish what the mechanisms and explanations are; quantitative studies establish how widespread each one is. A synthesis that reports a mechanism from interviews and its prevalence from a survey is stronger than either alone, provided it does not present the interview finding as a prevalence claim.\n\n### Who should own the synthesis?\nWhoever owns the standing question, which is usually a product lead rather than a researcher. Ownership by the research team tends to produce a document that is methodologically excellent and consulted rarely. Ownership by the decision-maker, with methods support, produces a document that gets updated because someone needs it.\n\n### What do I do when the synthesis says low confidence and a decision is due tomorrow?\nMake the decision and label the evidence honestly. Low confidence is not a reason to freeze; it is a reason to choose the reversible option, to instrument the outcome, and to schedule the study that would raise the rating. The failure is not deciding under low confidence - it is deciding under low confidence while describing the evidence as strong.\n\n## Related Resources\n\n- [Publication Bias in Product Research](/docs/publication-bias-product-research) - why the corpus you synthesise may be incomplete\n- [Equivalence Testing: How to Prove There Is No Difference](/docs/equivalence-testing-no-difference) - what a null entry in your evidence base actually means\n- [Triangulation in Research](/docs/triangulation-in-research-guide) - combining methods within a single question\n- [Conflicting Research Findings](/docs/conflicting-research-findings) - what to do when two datasets disagree\n- [Building a UX Research Repository](/docs/research-repository-guide) - the storage layer synthesis sits on top of\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types and when to use each\n- [Insight Repository Methodology](/docs/insight-repository-methodology) - tagging and activating a research library\n- [Thematic Analysis](/docs/thematic-analysis-guide) - the coding method that makes qualitative findings poolable","category":"Analysis & Synthesis","lastModified":"2026-08-09T03:23:08.839248+00:00","metaTitle":"Evidence Synthesis: Combining Findings Across Research Studies","metaDescription":"A repository tells you what study 47 found. A synthesis tells you what you collectively know and how sure you should be. Adapt GRADE and systematic review practice to product research in a week, not a year.","keywords":["evidence synthesis","combining research findings","cross-study synthesis","weight of evidence","GRADE certainty of evidence","systematic review","research confidence rating","narrative synthesis","instrument drift","standing research question"],"aiSummary":"Evidence synthesis combines findings from multiple studies into one conclusion with an explicit confidence rating. It is distinct from triangulation (methods within one question) and from repository storage (retrieval). GRADE rates a body of evidence at four certainty levels with five downgrade domains - risk of bias, inconsistency, indirectness, imprecision, publication bias - and three upgrade criteria. The main obstacle in product research is instrument drift, where the same question was operationalised differently each time so nothing pools; the fix is a small standing instrument reused verbatim across studies in a domain, plus organising the evidence base by standing decision question rather than by study.","aiPrerequisites":["Experience running or commissioning multiple research studies","Familiarity with qualitative and quantitative research outputs"],"aiLearningOutcomes":["Distinguish evidence synthesis from triangulation and from conflict resolution","Diagnose instrument drift and install a standing instrument that makes studies poolable","Rate a body of internal evidence using the four GRADE certainty levels","Apply the five downgrade domains and three upgrade criteria to a product research corpus","Run a seven-step synthesis on a standing decision question","Avoid vote counting, recency bias and false consistency when pooling findings"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 min"}],"pagination":{"total":1,"returned":1,"offset":0}}