Back to docs
Research Methods

Conflicting Research Findings: What to Do When Qualitative and Quantitative Data Disagree (2026)

When your interviews say one thing and your analytics say another, averaging them is the worst possible move. A step-by-step protocol for diagnosing and resolving conflicting research findings.

What do you do when research findings conflict?

When qualitative and quantitative findings disagree, do not average them, do not defer to the larger sample, and do not quietly drop the inconvenient one. Treat the conflict as a finding in its own right and run a diagnostic: confirm both results are real, check whether the two sources actually measured the same construct on the same people in the same window, and then explain the gap. In most cases the disagreement is not an error — it is the single most informative thing in your dataset, because it marks the exact point where your mental model of the user is wrong.

This guide gives you a repeatable protocol for that diagnosis, the five causes that account for nearly every real-world conflict, and the three responses that destroy research credibility fastest.

Why conflict is the normal case, not the exception

Every method measures a different construct and carries a different blind spot. Surveys measure stated attitudes. Interviews measure articulated reasoning and memory. Analytics measure logged behavior. These are three different things, and there is no law of nature requiring them to agree.

The gap between what people say they will do and what they actually do is one of the best-quantified effects in behavioral science. A synthesis of ten meta-analyses covering 422 studies found that intentions accounted for roughly 28% of the variance in subsequent behavior — meaning nearly three-quarters of behavioral variance is driven by something other than what people told you they intended. A separate meta-analysis of experimental studies found that a medium-to-large change in intention (d = 0.66) produced only a small-to-medium change in behavior (d = 0.36), a result Sheeran and Webb formalized as the intention-behavior gap.

The practical implication for research teams: if your interview participants enthusiastically describe a workflow they never actually complete in the product, you have not caught them lying. You have observed a well-replicated psychological phenomenon, and your job is to explain which mechanism produced it.

"Diversifying user research methods ensures more reliable, valid results by considering multiple ways of collecting and interpreting data. Using triangulation helps tell a consistent and cohesive story with multiple sources of data, to avoid stakeholder temptation to cherry-pick data that supports preexisting assumptions." — Nielsen Norman Group

Cherry-picking is the failure mode conflict invites. A protocol is what prevents it.

The four relationships between any two data sources

The most useful vocabulary for this comes from Farmer, Robinson, Elliott and Eyles (2006), whose triangulation protocol for qualitative health research introduced a convergence coding matrix. For each theme, you classify how the two sources relate:

RelationshipWhat it meansWhat to do
AgreementBoth sources point the same direction with the same emphasisRaise confidence; this is your most defensible finding
Partial agreementSame direction, different magnitude or scopeUsually a segment or timeframe effect — split the data
DissonanceThe sources genuinely contradict each otherRun the diagnostic below; do not resolve by fiat
SilenceOne source raises a theme the other never touchesNot a conflict — a coverage gap. Often the highest-value finding

Naming the relationship before you argue about the answer is what keeps the conversation analytical. Most stakeholder fights labelled "the research contradicts the data" turn out, on inspection, to be partial agreement or silence rather than true dissonance.

The five causes of nearly every real conflict

Work these in order. The first three explain the large majority of cases and are cheap to check.

1. Sample mismatch (different people)

Your interviews recruited engaged power users who answered a research invite. Your analytics cover everyone, including the 60% who churned in week one. These populations behave differently, so of course the numbers differ. Diagnostic: re-run the quantitative cut restricted to the exact population you interviewed. If the conflict evaporates, it was never a conflict.

2. Timeframe mismatch (different moments)

Interviews describe the present and a reconstructed past; analytics aggregate a window that may span a redesign, a pricing change, or a seasonal peak. Diagnostic: narrow the analytics window to the weeks your participants are actually describing.

3. Construct mismatch (different questions)

This is the subtle one. "Would you use this?" and "did you use this?" are different constructs, and so are "satisfaction" and "retention". Teams routinely treat a stated-preference measure and a revealed-preference measure as interchangeable evidence about the same thing. Diagnostic: write both measures out as literal sentences and ask whether a single person could honestly satisfy one and not the other. If yes, you have a construct mismatch, not a contradiction.

4. Response bias in the qualitative source

Social desirability, acquiescence, and courtesy bias all push interview answers toward the flattering. A participant who tells a moderator the onboarding was "pretty clear" may be managing the social situation rather than reporting an experience. Diagnostic: look for hedging language and check whether the positive statements contain any specifics. Vague praise plus concrete complaints is the signature of courtesy bias.

5. Aggregation hiding a real segment split

The quantitative average is flat; the interviews are polarized. Both are correct — the mean is concealing a bimodal distribution where one segment loves the feature and another abandons it. Diagnostic: stop looking at the mean and plot the distribution. This is the most commonly missed cause and the one with the most product value, because it usually reveals a segmentation your roadmap does not yet reflect.

The resolution protocol

  1. Verify both results before theorising. Re-run the query, re-read the transcripts. A surprising share of "conflicts" are a broken filter or a misread chart.
  2. Classify the relationship using the four categories above. Only genuine dissonance needs the full diagnostic.
  3. Align the populations, then the timeframes, then the constructs — in that order, because each is cheaper than the next.
  4. State the competing explanations explicitly. Write them down as testable claims, not vibes.
  5. Collect the tiebreaker. This is the step teams skip, because historically it meant another three-week study.
  6. Report the conflict, not just the resolution. Stakeholders trust research that shows its work far more than research that arrives suspiciously tidy.

What not to do

  • Do not average them. The mean of a behavioral measure and an attitudinal measure is a number that describes nothing real.
  • Do not automatically defer to the bigger sample. Sample size fixes sampling error. It does nothing for construct mismatch — a 50,000-response survey measuring the wrong thing is precisely as wrong as a 50-response one, only more persuasive.
  • Do not quietly drop the inconvenient source. This is how research programs lose credibility. If you discard a source, say so in the report and say why.

How Koji resolves conflicts in hours instead of weeks

Step 5 — collect the tiebreaker — is where most conflict resolution dies. In a traditional setup, testing a competing explanation means writing a screener, booking recruitment, scheduling moderated sessions, and waiting weeks. By the time the answer arrives, the decision has been made without it.

AI-moderated research changes the economics of that step specifically:

  • Targeted follow-up in a day, not a quarter. Because Koji's AI moderator runs interviews asynchronously and around the clock, you can put the exact diagnostic question to the exact disputed segment and have coded answers back the same day.
  • Population alignment by design. Import the precise cohort from your CRM or product analytics — the users the disputed metric describes — so the sample-mismatch cause is eliminated rather than argued about.
  • Attitudinal and behavioral measures in one instrument. Koji's six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let a single study capture a rating and the reasoning behind it, so the quantitative and qualitative signals come from the same people in the same session. Construct mismatch becomes visible inside one dataset instead of across two vendors.
  • Distribution, not just the average. Automatic aggregation charts scale and choice responses as distributions with the supporting verbatims attached, which is exactly what surfaces cause #5 — the hidden segment split behind a flat mean.
  • Consistent moderation removes one variable. Interviewer effects vary between human moderators and between sessions; an AI moderator asks the agreed question the agreed way every time, so a divergence between waves is more likely to be signal than moderator drift.

While legacy survey platforms like SurveyMonkey or Qualtrics can tell you that 40% disagree, and analytics tools like Amplitude or Mixpanel can tell you that usage fell, neither can ask the follow-up question that explains why. Closing that loop in a single platform is what turns a stalled stakeholder argument into a decision.

A worked example

A subscription team sees NPS rise four points while renewals fall. Averaging says "flat"; deferring to the larger sample says "customers are happy". Both are wrong.

Running the protocol: populations align, timeframes align, but the constructs do not — NPS measures advocacy among responders, renewal measures behavior among everyone. Plotting the distribution reveals a bimodal split. Fifteen targeted AI interviews with non-responding accounts surface the mechanism: a workflow the product removed in the last release, which delighted the vocal power users who requested the simplification and quietly stranded the administrators who depended on it.

The conflict was not noise. It was the finding.

Related Resources

Frequently asked questions

Which should I trust, interviews or analytics? Neither by default. Analytics tell you what happened with high reliability and no explanation; interviews tell you why with rich explanation and low generalizability. Trust each for what it measures well, and treat a conflict between them as a signal that you have mislabelled what one of them measures.

Is it ever correct to just pick one source? Yes — when one is demonstrably invalid for the question, for example a survey item that turned out to be double-barrelled, or an analytics event that was firing incorrectly. State the reason in the report. What is never correct is picking the one that matches the roadmap.

How much extra data do I need to resolve a conflict? Far less than a fresh study. Diagnostic follow-ups are narrow by design — you are testing two or three named explanations, not exploring. Ten to fifteen targeted interviews with the disputed segment is usually decisive.

Related Articles

Mixed Methods Research: How to Combine Qualitative and Quantitative Data

Learn how to design and run mixed methods research that combines the statistical power of quantitative data with the depth of qualitative insight — including how AI interview platforms like Koji make mixed methods accessible to every research team.

Product Analytics vs. User Research: When to Use Each (2026 Guide)

Product analytics tells you what users do; user research tells you why. Learn when to use each, how they combine, and how to get the why at analytics speed.

Qualitative Research Validity and Reliability: How to Build Studies You Can Trust

A practical guide to Lincoln and Guba's trustworthiness framework — credibility, transferability, dependability, and confirmability — and how to build each into your qualitative research studies.

Research Bias: The Complete Guide to Cognitive Biases That Corrupt User Research

A comprehensive guide to the 9 most damaging cognitive biases in user research — from confirmation bias to social desirability bias — with practical strategies to detect and eliminate them before they corrupt your findings.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Triangulation in Research: Combining Methods for Stronger, More Credible Insights (2026)

Triangulation is the practice of using multiple data sources, methods, researchers, or theories to validate a finding. Learn Denzin's four types, when to use each, and how AI-native research platforms make multi-method studies practical instead of aspirational.