{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-07-30T18:00:16.510Z"},"content":[{"type":"documentation","id":"a17bf0ca-237c-4433-ad48-42a91abc045a","slug":"conflicting-research-findings","title":"Conflicting Research Findings: What to Do When Qualitative and Quantitative Data Disagree (2026)","url":"https://www.koji.so/docs/conflicting-research-findings","summary":"When qualitative and quantitative findings disagree, the conflict is usually a finding rather than an error. Five causes explain nearly all real-world conflicts: sample mismatch, timeframe mismatch, construct mismatch, response bias in the qualitative source, and aggregation hiding a segment split. Farmer et al. (2006) provide the vocabulary — agreement, partial agreement, dissonance, silence. The intention-behavior gap (intentions explain ~28% of behavioral variance) explains why say-do divergence is expected rather than anomalous. The resolution protocol is: verify both results, classify the relationship, align populations then timeframes then constructs, state competing explanations, collect a targeted tiebreaker, and report the conflict itself. AI-moderated platforms collapse the cost of the tiebreaker step from weeks to hours.","content":"## What do you do when research findings conflict?\n\nWhen qualitative and quantitative findings disagree, do not average them, do not defer to the larger sample, and do not quietly drop the inconvenient one. Treat the conflict as a finding in its own right and run a diagnostic: confirm both results are real, check whether the two sources actually measured the same construct on the same people in the same window, and then explain the gap. In most cases the disagreement is not an error — it is the single most informative thing in your dataset, because it marks the exact point where your mental model of the user is wrong.\n\nThis guide gives you a repeatable protocol for that diagnosis, the five causes that account for nearly every real-world conflict, and the three responses that destroy research credibility fastest.\n\n## Why conflict is the normal case, not the exception\n\nEvery method measures a different construct and carries a different blind spot. Surveys measure stated attitudes. Interviews measure articulated reasoning and memory. Analytics measure logged behavior. These are three different things, and there is no law of nature requiring them to agree.\n\nThe gap between what people say they will do and what they actually do is one of the best-quantified effects in behavioral science. A synthesis of ten meta-analyses covering 422 studies found that intentions accounted for roughly **28% of the variance in subsequent behavior** — meaning nearly three-quarters of behavioral variance is driven by something other than what people told you they intended. A separate meta-analysis of experimental studies found that a medium-to-large change in intention (d = 0.66) produced only a small-to-medium change in behavior (d = 0.36), a result Sheeran and Webb formalized as the *intention-behavior gap*.\n\nThe practical implication for research teams: if your interview participants enthusiastically describe a workflow they never actually complete in the product, you have not caught them lying. You have observed a well-replicated psychological phenomenon, and your job is to explain which mechanism produced it.\n\n> \"Diversifying user research methods ensures more reliable, valid results by considering multiple ways of collecting and interpreting data. Using triangulation helps tell a consistent and cohesive story with multiple sources of data, to avoid stakeholder temptation to cherry-pick data that supports preexisting assumptions.\" — Nielsen Norman Group\n\nCherry-picking is the failure mode conflict invites. A protocol is what prevents it.\n\n## The four relationships between any two data sources\n\nThe most useful vocabulary for this comes from Farmer, Robinson, Elliott and Eyles (2006), whose triangulation protocol for qualitative health research introduced a *convergence coding matrix*. For each theme, you classify how the two sources relate:\n\n| Relationship | What it means | What to do |\n| --- | --- | --- |\n| **Agreement** | Both sources point the same direction with the same emphasis | Raise confidence; this is your most defensible finding |\n| **Partial agreement** | Same direction, different magnitude or scope | Usually a segment or timeframe effect — split the data |\n| **Dissonance** | The sources genuinely contradict each other | Run the diagnostic below; do not resolve by fiat |\n| **Silence** | One source raises a theme the other never touches | Not a conflict — a coverage gap. Often the highest-value finding |\n\nNaming the relationship before you argue about the answer is what keeps the conversation analytical. Most stakeholder fights labelled \"the research contradicts the data\" turn out, on inspection, to be *partial agreement* or *silence* rather than true dissonance.\n\n## The five causes of nearly every real conflict\n\nWork these in order. The first three explain the large majority of cases and are cheap to check.\n\n### 1. Sample mismatch (different people)\n\nYour interviews recruited engaged power users who answered a research invite. Your analytics cover everyone, including the 60% who churned in week one. These populations behave differently, so of course the numbers differ. **Diagnostic:** re-run the quantitative cut restricted to the exact population you interviewed. If the conflict evaporates, it was never a conflict.\n\n### 2. Timeframe mismatch (different moments)\n\nInterviews describe the present and a reconstructed past; analytics aggregate a window that may span a redesign, a pricing change, or a seasonal peak. **Diagnostic:** narrow the analytics window to the weeks your participants are actually describing.\n\n### 3. Construct mismatch (different questions)\n\nThis is the subtle one. \"Would you use this?\" and \"did you use this?\" are different constructs, and so are \"satisfaction\" and \"retention\". Teams routinely treat a stated-preference measure and a revealed-preference measure as interchangeable evidence about the same thing. **Diagnostic:** write both measures out as literal sentences and ask whether a single person could honestly satisfy one and not the other. If yes, you have a construct mismatch, not a contradiction.\n\n### 4. Response bias in the qualitative source\n\nSocial desirability, acquiescence, and courtesy bias all push interview answers toward the flattering. A participant who tells a moderator the onboarding was \"pretty clear\" may be managing the social situation rather than reporting an experience. **Diagnostic:** look for hedging language and check whether the positive statements contain any specifics. Vague praise plus concrete complaints is the signature of courtesy bias.\n\n### 5. Aggregation hiding a real segment split\n\nThe quantitative average is flat; the interviews are polarized. Both are correct — the mean is concealing a bimodal distribution where one segment loves the feature and another abandons it. **Diagnostic:** stop looking at the mean and plot the distribution. This is the most commonly missed cause and the one with the most product value, because it usually reveals a segmentation your roadmap does not yet reflect.\n\n## The resolution protocol\n\n1. **Verify both results before theorising.** Re-run the query, re-read the transcripts. A surprising share of \"conflicts\" are a broken filter or a misread chart.\n2. **Classify the relationship** using the four categories above. Only genuine dissonance needs the full diagnostic.\n3. **Align the populations, then the timeframes, then the constructs** — in that order, because each is cheaper than the next.\n4. **State the competing explanations explicitly.** Write them down as testable claims, not vibes.\n5. **Collect the tiebreaker.** This is the step teams skip, because historically it meant another three-week study.\n6. **Report the conflict, not just the resolution.** Stakeholders trust research that shows its work far more than research that arrives suspiciously tidy.\n\n## What not to do\n\n- **Do not average them.** The mean of a behavioral measure and an attitudinal measure is a number that describes nothing real.\n- **Do not automatically defer to the bigger sample.** Sample size fixes sampling error. It does nothing for construct mismatch — a 50,000-response survey measuring the wrong thing is precisely as wrong as a 50-response one, only more persuasive.\n- **Do not quietly drop the inconvenient source.** This is how research programs lose credibility. If you discard a source, say so in the report and say why.\n\n## How Koji resolves conflicts in hours instead of weeks\n\nStep 5 — collect the tiebreaker — is where most conflict resolution dies. In a traditional setup, testing a competing explanation means writing a screener, booking recruitment, scheduling moderated sessions, and waiting weeks. By the time the answer arrives, the decision has been made without it.\n\nAI-moderated research changes the economics of that step specifically:\n\n- **Targeted follow-up in a day, not a quarter.** Because Koji's AI moderator runs interviews asynchronously and around the clock, you can put the exact diagnostic question to the exact disputed segment and have coded answers back the same day.\n- **Population alignment by design.** Import the precise cohort from your CRM or product analytics — the users the disputed metric describes — so the sample-mismatch cause is eliminated rather than argued about.\n- **Attitudinal and behavioral measures in one instrument.** Koji's six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let a single study capture a rating *and* the reasoning behind it, so the quantitative and qualitative signals come from the same people in the same session. Construct mismatch becomes visible inside one dataset instead of across two vendors.\n- **Distribution, not just the average.** Automatic aggregation charts scale and choice responses as distributions with the supporting verbatims attached, which is exactly what surfaces cause #5 — the hidden segment split behind a flat mean.\n- **Consistent moderation removes one variable.** Interviewer effects vary between human moderators and between sessions; an AI moderator asks the agreed question the agreed way every time, so a divergence between waves is more likely to be signal than moderator drift.\n\nWhile legacy survey platforms like SurveyMonkey or Qualtrics can tell you *that* 40% disagree, and analytics tools like Amplitude or Mixpanel can tell you *that* usage fell, neither can ask the follow-up question that explains why. Closing that loop in a single platform is what turns a stalled stakeholder argument into a decision.\n\n## A worked example\n\nA subscription team sees NPS rise four points while renewals fall. Averaging says \"flat\"; deferring to the larger sample says \"customers are happy\". Both are wrong.\n\nRunning the protocol: populations align, timeframes align, but the constructs do not — NPS measures advocacy among *responders*, renewal measures behavior among *everyone*. Plotting the distribution reveals a bimodal split. Fifteen targeted AI interviews with non-responding accounts surface the mechanism: a workflow the product removed in the last release, which delighted the vocal power users who requested the simplification and quietly stranded the administrators who depended on it.\n\nThe conflict was not noise. It was the finding.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — how the six question types capture quantitative and qualitative signal in one session\n- [Triangulation in Research](/docs/triangulation-in-research-guide) — designing studies with complementary blind spots from the start\n- [Product Analytics vs. User Research](/docs/product-analytics-vs-user-research) — which source answers which question\n- [Mixed Methods Research](/docs/mixed-methods-research-guide) — combining qualitative and quantitative data by design\n- [Qualitative Research Validity and Reliability](/docs/qualitative-research-validity) — building studies you can defend\n- [Research Bias: The Complete Guide](/docs/research-bias-guide) — the biases behind cause #4\n\n## Frequently asked questions\n\n**Which should I trust, interviews or analytics?**\nNeither by default. Analytics tell you what happened with high reliability and no explanation; interviews tell you why with rich explanation and low generalizability. Trust each for what it measures well, and treat a conflict between them as a signal that you have mislabelled what one of them measures.\n\n**Is it ever correct to just pick one source?**\nYes — when one is demonstrably invalid for the question, for example a survey item that turned out to be double-barrelled, or an analytics event that was firing incorrectly. State the reason in the report. What is never correct is picking the one that matches the roadmap.\n\n**How much extra data do I need to resolve a conflict?**\nFar less than a fresh study. Diagnostic follow-ups are narrow by design — you are testing two or three named explanations, not exploring. Ten to fifteen targeted interviews with the disputed segment is usually decisive.","category":"Research Methods","lastModified":"2026-07-28T03:18:25.454577+00:00","metaTitle":"Conflicting Research Findings: When Qual and Quant Disagree (2026) | Koji","metaDescription":"When interviews and analytics contradict each other, averaging them is the worst move. Learn the 5 causes of conflicting research findings, a 6-step resolution protocol, and how AI interviews deliver the tiebreaker in hours.","keywords":["conflicting research findings","qualitative and quantitative data disagree","when user research contradicts analytics","say-do gap","intention-behavior gap","divergent findings","triangulation","conflicting user research results","research conflict resolution","Koji"],"aiSummary":"When qualitative and quantitative findings disagree, the conflict is usually a finding rather than an error. Five causes explain nearly all real-world conflicts: sample mismatch, timeframe mismatch, construct mismatch, response bias in the qualitative source, and aggregation hiding a segment split. Farmer et al. (2006) provide the vocabulary — agreement, partial agreement, dissonance, silence. The intention-behavior gap (intentions explain ~28% of behavioral variance) explains why say-do divergence is expected rather than anomalous. The resolution protocol is: verify both results, classify the relationship, align populations then timeframes then constructs, state competing explanations, collect a targeted tiebreaker, and report the conflict itself. AI-moderated platforms collapse the cost of the tiebreaker step from weeks to hours.","aiPrerequisites":["triangulation-in-research-guide","qualitative-vs-quantitative-research"],"aiLearningOutcomes":["Diagnose why qualitative and quantitative findings disagree","Classify source relationships as agreement, partial agreement, dissonance, or silence","Apply the five common causes of research conflict to a real dataset","Run a six-step conflict resolution protocol","Avoid the three responses that destroy research credibility","Use AI-moderated follow-ups to collect a tiebreaker in hours"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}