Back to blog
Use Cases

AI Feature Adoption Research: A Product Manager Playbook for Finding Out Why Users Do Not Trust or Use Your AI Feature

Half of software makers say fewer than a quarter of customers use the AI features they built. Here is how a PM can find out why, using interviews split by what users actually did.

Koji

Koji Team

Product · · 12

You shipped the AI feature. The launch post went out, the demo got applause, and now the dashboard shows a flat line. AI feature adoption research is the work of finding out why: which users tried it, which ones quietly went back to the old way, and what each group needed to see before they would rely on it. Run as AI-moderated interviews, you can talk to 30 or more of those users in a few days, split by what they did rather than what they say, and walk into your next roadmap review with reasons instead of guesses.

Why AI feature adoption research matters now

Most teams have shipped AI into an existing product by now. Far fewer know whether anyone uses it.

  • In Banyan Software's 2026 benchmark of vertical software operators (more than 260 founder and CEO responses), 51% said fewer than one in four of their customers use the AI features they built. CIO reported the survey in August 2026. Max Risen, Banyan's president of M&A, told IT Brew the likely cause is companies "building for the sake of building".
  • Userflow's 2026 survey of 107 senior product leaders found that 44% name trusting the output as the place users struggle most. That ranks above understanding the feature (26%), fitting it into a workflow (25%) and discovering it exists (21%). The same survey says only 33% of teams measure outcomes such as task completion or time saved, and 30% have no reliable way to tell whether the AI helps at all. Details are on Userflow's write-up. Note that this is what product leaders believe about their users, not what users said.
  • The problem is not new. In 2024, Mind the Product and Pendo reported that only 15% of the 300+ product leaders they surveyed said their users were embracing AI features.

Read those three together. Trust is the leading barrier, and the metrics most teams watch (clicks, sessions, weekly active users) cannot see trust. A PM who is measured on adoption, retention and support load is flying with the one instrument that does not read the main variable.

The cost of not finding out is specific. You keep funding a feature most of your base ignores, or you cut one that a small segment depends on, and you will not know which until someone leaves.

What good looks like

Teresa Torres, who wrote Continuous Discovery Habits, defines the practice as the team building the product talking to customers at least weekly, in small research activities tied to a desired outcome. Her story-based approach is the right base here: ask for a specific time the user did something, not their opinion of the feature. Everything below builds on that.

1. Sample by behaviour, not by plan tier. Split users into four groups:

  • Never tried it. They saw it, or never saw it.
  • Tried once and stopped. This is usually your richest group.
  • Use it, but check everything it does. Heavy verification is a trust signal.
  • Rely on it. They tell you what a good outcome looks like.

Pull each group from your analytics. If you use a product analytics tool, a cohort for each is a ten-minute job.

2. Interview enough to see each group's reasons repeat. Hennink, Kaiser and Marconi (2017) found that new codes stopped appearing at around 9 interviews, but understanding the meaning behind them took 16 to 24. Nielsen Norman Group makes a similar point: interview studies vary more than usability tests, so five is often too few. With four groups, aim for 6 to 10 per group. That is about 30 to 40 conversations, which is where an AI moderator earns its place.

3. Ask about the last real task, not the feature. "Walk me through the last time you had to do X" beats "What do you think of the summary feature?" Then ask what they did next: did they use the output as is, edit it, or redo it by hand?

4. Probe trust at the moment of use. Ask what they looked at to decide whether the output was right. Ask what would have to be visible for them to stop checking. Trust problems hide in answers like "I just glance at it", which a good probe turns into specifics.

5. Get a number, then ask about the number. A 1 to 10 confidence score is useful only when followed by "why that number and not two points higher?"

6. Map every finding to an outcome metric. Pick one before you start: task completion, time on task, tickets avoided. If a finding cannot connect to it, park it.

Common mistakes

  • Interviewing only the people who use the feature. Non-adopters hold the biggest barriers.
  • Asking "Did you like it?" You get politeness.
  • Treating low usage as a discoverability problem when it is a trust problem, or the reverse. The interview should tell you which.
  • Running one study and calling it done. Trust shifts after model changes, new data sources and bad outputs, so repeat after any meaningful change.
  • Letting a loud example win. One dramatic quote does not outweigh a theme from 12 interviews.

How to set it up in Koji

You can run this on any plan for text interviews. Voice interviews need the Interviews or Enterprise plan.

  1. Start a study from a template. Create a new study and choose the Post-Launch Feedback Interviews template. It is written for exactly this situation: it asks about the need before the feature, how the person found it, what they did the first time, and why they came back or did not. Rename it for the feature you are researching.
  2. Shape the research brief. Fill in the decision this study informs ("keep investing, redesign the trust cues, or cut"), your current hypothesis (for example "adoption is low because users cannot tell when the output is wrong"), and who counts as a participant. The brief assistant can change any part of the brief on request, including screening questions and follow-up depth, and tells you what it changed.
  3. Add screening questions for your four groups. A single-choice question such as "Which best describes your use of the AI assistant in the last 30 days?" can qualify the groups you want and screen out the rest. Screening questions are checked before the interview starts. They apply to studies created from September 5, 2026, so create a fresh study if you are updating an older one.
  4. Add structured questions where you need a number or a choice. Koji supports open-ended, scale, single choice, multiple choice, ranking and yes/no questions. Use a scale for confidence and a ranking for "what would make you trust it more". Use open-ended questions for everything else and let the interviewer follow up.
  5. Preview it. The Preview tab lets you take your own interview by text or voice. Previews are free, do not count as responses and never appear in the report. Fix anything that sounds like a leading question before you invite real users.
  6. Choose text or voice, then invite people. A text interview costs 1 credit and a voice interview costs 3. Share the interview link, or use personalized links to send each user their own. If you want to find users inside your analytics, the Mixpanel and PostHog integration guides describe how to trigger interviews from product events. If you need people who are not your users (for example, to understand why prospects distrust AI in your category), panel recruitment is available on paid plans.
  7. Read the report. Koji produces a report with themes and quotes, and charts for your structured questions. Only interviews that score 3 or higher on the quality score use credits or enter the report.

Example questions for this study

  • "Think about the last time you needed to [task the AI feature helps with]. Walk me through what you did, from the start."
  • "Did you use the AI suggestion at any point? What did you do with what it gave you?" (probe: used as is, edited, discarded)
  • "When it gave you a result, how did you decide whether it was right?" (probe: what did you look at?)
  • "On a scale of 1 to 10, how much do you trust it to do this task without you checking?" (scale, then probe the number)
  • "What would have to be true for you to stop double-checking its work?"
  • "Rank these in order of what would make you rely on it more: showing its sources, a way to undo, seeing how it reached the answer, a track record, an option to review before it acts." (ranking, adapt the options to your product)

Plans and cost

Free includes a one-time grant of 10 credits and text interviews. Insights is €29 a month with 29 credits, enough for about 29 text interviews. Interviews is €79 a month with 79 credits, which covers 79 text interviews or 26 voice interviews. Check the pricing page for the current figures before you budget, because plans change.

What you get and how a PM uses it the next day

Within a few days you have a report that separates the four groups and shows where they diverge. The next-day uses are concrete:

  • In the roadmap review: "Of the 'tried once and stopped' group, most could not tell when the summary was wrong. That is a design problem, not a marketing one." That sentence changes the conversation.
  • In the backlog: two or three trust cues with evidence behind them, such as showing sources or adding a review step.
  • In the metrics plan: the outcome metric you will use to judge the redesign, chosen from what users said mattered.
  • In the go or no-go on cutting the feature: a clear view of who depends on it. For the other side of that decision, see the guide to feature sunset research.

You can also chat with the study data to ask follow-up questions across all the interviews, such as "which non-adopters mentioned data privacy?"

The same playbook elsewhere

The structure carries over to other roles. Customer success teams run the same four-group split after a new onboarding flow. HR and people ops teams use it to understand why employees avoid an internal AI tool; Koji's guide to employee AI adoption research covers that. Founders use a lighter version in the first month after launching a feature to a handful of early customers. In all of them, the move is the same: segment by behaviour, ask about the last real task, and probe trust.

Why Koji for AI feature adoption research

What a PM weighsAnalytics aloneSurvey or in-app promptManual interviewsKoji
Explains whyNoThinYesYes, with follow-up probing
Time to resultsImmediateDaysWeeks to scheduleDays
Reach across 4 behaviour groupsYes, but no reasonsBiased to respondersLimited by PM hours30+ interviews in parallel
Consistency across interviewsNot applicableHighVaries by moderatorSame brief for every interview
Candour about a feature your team builtNot applicableSocial desirabilitySocial desirabilityParticipants talk to an AI, not the builder
Analysis effortLowLowHighThemes, quotes and charts generated
CostTool costTool costPM and researcher timeFrom €29 a month

The reasons, each tied to something you can check in the product:

  • It probes. Each open-ended question can have follow-ups, so "I just glance at it" turns into what they glance at.
  • It mixes numbers and conversation. Scale, choice and ranking questions sit inside the interview and show up as charts in the report.
  • It screens. You can send one link and let screening questions sort people into your four groups.
  • It protects your budget. Low-effort interviews are filtered out by the quality score and do not use credits.
  • You can rehearse. Free previews mean you hear the interview before your users do.

Questions you might have about AI interviews

Will users be honest with an AI about a feature my team built? Some people criticise more freely when nobody from the team is listening live, though that is not guaranteed. Frame the opening as "we want to hear what does not work", and use the preview to hear how it sounds before you invite anyone.

Can an AI moderator probe as well as an experienced researcher? It follows the depth you set per question and asks follow-ups based on what the person said. Read the first five transcripts and adjust the brief, as you would after a pilot with a new human interviewer. Koji's guide on AI interview quality explains the safeguards.

Where does the data go? The privacy and security guide covers how interview data is handled, and the GDPR guide covers consent. If your company has strict requirements, read them before you invite customers.

When Koji may not be the right fit

If you need to watch people struggle in the interface, with clicks and eye movement, you want a usability test or session replay, and the interview is the follow-up. If your user base is only a few dozen people, a handful of calls you run yourself may serve you better. And if the real question is whether the AI output is accurate, that is an evaluation problem for your engineering team; interviews tell you whether users believe it is.

Metrics to track

  • Share of eligible users who tried the feature, and who used it more than once
  • Task completion or time on task for users of the feature versus a comparable group
  • Trust score (1 to 10) by behaviour group, before and after changes
  • Support tickets that mention the feature
  • Proportion of interviews that mention each trust barrier, so you can see which ones your changes move

Where to start

Create a study from the Post-Launch Feedback template, add the four-group screening question, and preview it. If you want a wider view of the role, see Koji for product managers. For related reading, the Product Manager's Guide to Customer Discovery with AI, the 2026 playbook for researching AI products and the Continuous Discovery Handbook pair well with this one. For the general method, see feature adoption research.

Run your first AI-moderated study in 10 minutes

10 free credits on signup. No credit card required.

GDPR compliantEU or US data residencyNo AI training on your data
Koji

Koji Team

Product

Share this article