Back to blog
Research

Customer Research for Data and Analytics Teams: How to Answer the "Why" Behind Your Metrics (2026)

Your warehouse can tell you what changed, for whom, and by how much. It cannot tell you why, because motivation was never emitted as an event. A practical playbook for analytics engineers, analysts and data scientists on adding qualitative evidence to a quantitative practice, without hiring researchers.

Koji

Koji Team

Research · · 11 min read

Short answer: the "why" is the only input a data team cannot backfill from the warehouse. Every other missing field can be joined, modelled, backfilled, or instrumented in the next sprint. Motivation cannot, because it was never emitted as an event. That is a structural hole in every dashboard you have ever built, and in 2026 it is finally cheap to fill.

This is a playbook for data and analytics teams — analytics engineers, analysts, data scientists, and the leaders who run them — on how to add qualitative evidence to a quantitative practice without hiring a research team.

The problem is not that you lack data

Data teams are not short of numbers. They are short of explanations, and they are being asked for explanations at an accelerating rate.

The 2026 State of Analytics Engineering Report from dbt Labs, released 14 April 2026, is the clearest read on where the pressure is coming from. Its headline finding is a trust crisis created by speed:

  • Trust in data as an organisational priority jumped from 66% to 83% year over year — the steepest single-year increase of any objective the survey measures.
  • "Shipping data products faster" climbed from 50% to 71%.
  • 71% of data professionals cite incorrect or hallucinated outputs reaching stakeholders as a top concern.
  • 72% now prioritise AI-assisted coding, while only 24% prioritise AI-assisted pipeline management, including testing and observability.
  • 57% report increased warehouse and compute spend, against just 36% reporting increased team budgets.
  • Ambiguous data ownership remains a top governance obstacle at 41%.

Read the third and fourth bullets together. Teams have adopted AI overwhelmingly to produce faster and only marginally to validate. Output is accelerating; the checks on that output are not. That is the exact shape of the trust problem the report names.

The prior year set the baseline. The 2025 edition, based on 459 data practitioners surveyed between 8 October and 27 December 2024 (70% individual contributors, 30% managers), found 56% citing poor data quality as their most frequent challenge, 57% spending most of their workdays maintaining or organising data sets, and 80% already using AI in day-to-day work, up from 30% a year earlier. Nearly 65% said that enabling non-technical business users to create governed data sets would improve their organisation's data value.

So: more output, more spend, flat headcount, and a stakeholder base that trusts the numbers less than it used to.

Why more pipeline work will not fix this

Here is the uncomfortable diagnosis. A large share of what stakeholders call a "data quality problem" is not a defect in the pipeline at all. It is a missing explanation.

The dashboard is correct. Activation fell 9% in March. The model is right, the tests pass, the lineage is clean. And the question that comes back is: why?

That question has no answer in the warehouse. Not because the warehouse is badly built, but because motivation is not an event. Nobody fires intent_to_churn with a reason property. You can observe that a cohort stopped, and you can observe everything they did before they stopped, and you still cannot observe what they were trying to do and why it was not worth continuing.

When the data team cannot answer why, the organisation does not stop asking. It defaults to the loudest available explanation — usually the most senior person's hypothesis, occasionally a plausible-sounding correlation someone found in the same dashboard. Then the roadmap gets built on it.

That is how a numerically rigorous organisation makes decisions with no evidence at all.

The three questions your data can and cannot answer

QuestionWarehouse can answerNeeds customers
What changed?Yes — preciselyNo
Where and for whom?Yes — down to the cohortNo
How much did it cost?YesNo
Why did they do it?NoYes
What did they expect instead?NoYes
What did they try before giving up?Partially — only what you instrumentedYes
What would have changed their mind?NoYes

The bottom four rows are where roadmap decisions actually get made. Our guides to qualitative vs quantitative research and mixed methods research cover the methodological framing; what follows is the operational version for a team that already owns the numbers.

Your unfair advantage: you already have the sampling frame

Most people who run customer research struggle with recruitment. They need "users who activated but never returned," and they have to approximate it with a screener question and hope respondents self-report accurately — which, as anyone who has run a screener knows, they frequently do not.

You do not have that problem. You can define the cohort exactly.

That is the single biggest structural advantage a data team brings to qualitative research, and almost nobody exploits it. You can build a sampling frame that is not an approximation:

  • Users who hit the activation event and then went dormant for 21 days
  • Accounts whose seat utilisation dropped by half in a quarter but renewed anyway
  • The 4% of users who trigger the power feature, versus a matched control who never did
  • Everyone who abandoned at step 3 of the flow you shipped last month
  • Trials that converted despite a support ticket in week one

Each of those is a SELECT you can already write. Turn it into an interview list and you have a study whose sampling frame is exact rather than self-reported — a level of precision most professional researchers never get. Koji's Amplitude and Mixpanel integrations exist to close this loop directly: trigger interviews from product events, and pipe the resulting insight back as properties on the same profiles.

The metric-anomaly triage loop

The practical pattern that makes this routine rather than heroic:

1. The metric moves. Your anomaly detection or a stakeholder flags it.

2. You answer what, where, and how much. This is the part you are already excellent at, and it usually takes an afternoon. Segment the change, confirm it is not an instrumentation artefact, size it.

3. You define the affected cohort precisely. The output of step 2 is a cohort definition. This is the handoff most teams never make.

4. You field a short study against that exact cohort — same week. Ten to thirty conversations, six to eight questions, open-ended with structured measures attached. Not a quarter-long research project. A same-cycle instrument.

5. You deliver the why alongside the what. One artefact, both halves.

The reason this loop rarely existed before is straightforward: step 4 used to cost weeks of scheduling and a trained moderator, so it could never run on the same cadence as step 2. When the qualitative half takes as long as a sprint and the quantitative half takes an afternoon, nobody waits — they guess. Collapse step 4 to hours and the loop closes.

What good looks like: five studies worth standing up

Start with these. They map onto questions your stakeholders already ask you, and each one has a natural cohort definition you can write in SQL today.

  1. Activation drop-off. Cohort: reached step n, never reached n+1, within 14 days. Ask what they were trying to accomplish and what they did instead. Pairs with aha moment research and user onboarding research.
  2. Silent churn. Cohort: usage decayed 70%+ over two months, still subscribed. These accounts are invisible in your churn number until renewal, and they are the cheapest to save. See customer renewal interviews.
  3. The unexplained good cohort. Cohort: users who over-perform your model's prediction. Everyone studies failure; the mechanism behind your best cohort is usually reproducible and nobody has looked. Related: account expansion and upsell research.
  4. The flat experiment. Cohort: participants in a test that came back null. A null result means the metric did not move, not that nothing happened — and the qualitative read usually explains whether the change was invisible, irrelevant, or actively confusing. See A/B testing vs user research.
  5. Trials that did not convert. Cohort: trial expired, no purchase, no support contact. The silent non-converters, who never told you anything. See free trial conversion research.

Making qualitative evidence survive contact with a data team

Analysts are trained to be sceptical of small-n evidence, and rightly so. Four rules keep qualitative findings credible in a numerate room:

Report qual as coverage, not as proportion. "7 of 22 interviewees raised billing confusion unprompted" is a defensible statement. "32% of users are confused by billing" is not — you did not sample for prevalence. When you need prevalence, attach structured questions and field to a larger sample.

Pre-register the question, not the answer. Write down what you expect to hear before you field. It is the same discipline you apply to an experiment, and it is what stops thematic analysis from becoming confirmation of the hypothesis you started with.

Treat divergence as the finding. When the interviews contradict the dashboard, that is signal, not noise — most often it means the metric is measuring something adjacent to the behaviour you named it after. Conflicting research findings walks through how to adjudicate this properly.

Say what the sample cannot support. The fastest way to lose a numerate audience is to over-claim from 12 conversations. The fastest way to win one is to state the limit before they do.

Where Koji fits

Koji was built for teams who need the why on the same cadence as the what, without hiring researchers.

AI-moderated voice and text interviews run in parallel and on the respondent's own schedule, so a study you launch on Tuesday returns transcripts by Thursday rather than in three weeks. There is no moderator to book, and no moderator bias to control for.

Structured questions are the part data teams care about most. Koji supports six types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — so one study can carry a countable backbone alongside the open-ended depth. That means you get something you can actually join back to the warehouse, not just a folder of quotes. See structured questions in AI interviews.

Automatic thematic analysis reads every transcript, not the six someone had time for. One-click reports turn a finished study into a shareable artefact the same day — which matters when the stakeholder asking "why did activation drop" is waiting on your answer, not on a research readout next month.

And because it is credit-based rather than per-seat, the analyst who needs one study this quarter is not blocked by a licence. A text conversation costs 1 credit, a voice conversation 3, and Koji only charges for conversations that clear its quality bar, so poor-quality responses cost you nothing — a gating model that will feel familiar to anyone who has built a data quality test.

For the broader organisational pattern, see research democratization and our blog guides on scaling insights beyond the research team and the research team of one. Teams doing this alongside a support function will also want customer research for support and CX teams, and PLG organisations should read customer research for product-led growth.

Start with the metric you cannot explain

You already know which one it is. It is the number in your weekly review that someone asks about every week and nobody can answer — the one where the conversation ends in "we think it is probably..."

Write the cohort definition. Field ten conversations against it. Bring the answer to the next review.

Koji gives you 10 free credits when you sign up — enough to run a real study against a real cohort and see whether the explanation you have been assuming is the right one. No research expertise required, and from question to insight in hours rather than weeks.

Frequently Asked Questions

We already have product analytics. Why do we need interviews?

Product analytics tells you what happened, to whom, and how much. It cannot tell you what the user was trying to do, what they expected instead, or what would have changed their mind — because motivation is never emitted as an event. Those are exactly the inputs roadmap decisions turn on, which is why organisations that lack them tend to default to the most senior person's hypothesis.

How many interviews does a data team need to make a call?

For explaining a specific anomaly in a defined cohort, 10 to 20 conversations usually surfaces the dominant mechanisms. Report the result as coverage — "7 of 22 raised this unprompted" — rather than as a percentage of your user base, because you sampled for explanation, not prevalence. If you need prevalence, attach structured questions and field to a larger sample.

Can we recruit participants straight from a cohort we defined in the warehouse?

Yes, and this is the biggest advantage data teams have over most research functions. Instead of approximating a segment with self-reported screener questions, you can define it exactly — users who hit the activation event and went dormant for 21 days, for instance — and invite that list directly. Koji's Amplitude and Mixpanel integrations let you trigger interviews from product events and write insights back as profile properties.

How do we keep qualitative findings credible with a sceptical analytics audience?

Three habits: report coverage rather than proportions, write down what you expect to hear before you field so thematic analysis is not just confirmation, and state explicitly what your sample cannot support before someone else does. When interviews contradict the dashboard, treat the divergence as the finding — it usually means the metric measures something adjacent to what it is named after.

Does this replace our A/B testing programme?

No, it explains it. Experiments are excellent at measuring whether a change moved a metric and useless at explaining a null result. Pairing a flat experiment with a short study of the participants typically reveals whether the change was invisible, irrelevant, or actively confusing — three very different next steps that look identical in the experiment readout.

What is the realistic time cost for an analyst?

Roughly an hour to write six to eight good questions and define the cohort, then the study runs itself. Interviews happen in parallel on participants' own schedules, thematic analysis is automatic, and a shareable report is one click. The bottleneck for most teams is deciding what to ask, not running the study — which is the inverse of how traditional research works.

Run your first AI-moderated study in 10 minutes

10 free credits on signup. No credit card required.

GDPR compliantEU or US data residencyNo AI training on your data
Koji

Koji Team

Research

Share this article

Keep reading