Back to docs
Analysis & Synthesis

Why Adding One Option Can Reverse Your Ranking Results (2026)

Average rank, the default summary for ranking questions, depends on which other options are in the list. A worked example of a reversal no respondent caused, and the first-place and pairwise summaries that stay stable.

Why Adding One Option Can Reverse Your Ranking Results (2026)

Answer first: Average rank, the default summary for ranking questions in almost every survey tool, depends on which other options are in the list. Add or remove one option that nobody cares about and two items can swap places, even though not a single respondent changed their mind about those two items. Two summaries do not have this problem: first-place share (what fraction put each item first) and pairwise preference (for each pair, what fraction ranked one above the other). Report those alongside average rank, keep the option list fixed across waves, and never compare average ranks between lists of different lengths. Koji's ranking question type collects full orderings and can ask each participant why their top choice beat the runner-up, so you get the priority order and the reasoning behind it.

The worked example

Five customers rank three features, A, B and C. Three of them rank A first, B second, C third. Two of them rank B first, C second, A third.

SummaryABC
Average rank (lower is better)1.81.62.6
First-place share60%40%0%
Ranked above the other in the A-vs-B pair3 of 52 of 5-

Average rank says B is the top priority. First-place share and the head-to-head comparison both say A.

Now the team adds a fourth option, D, to the next wave. Nobody changes their view of A, B or C. The first three customers slot D in second (A, D, B, C); the other two put it last (B, C, A, D).

SummaryABCD
Average rank1.82.23.22.8
First-place share60%40%0%0%
A-vs-B head-to-head3 of 52 of 5--

Average rank now says A is the top priority. The roadmap review will see A overtook B. Nothing about A or B changed. The only thing that changed is that D pushed B down one position for three respondents, while A, ranked first by those same respondents, could not be pushed.

First-place share and the head-to-head count gave the same answer in both waves. Average rank flipped.

Why average rank behaves this way

Average rank is the same arithmetic as a Borda count, the points-based voting scheme, and it inherits a property that social-choice theory has studied for more than two centuries: the result between two options can depend on a third, irrelevant option. Every option you add takes up a position, and it takes that position from whichever items respondents rank below it. Items that respondents love are shielded because they sit above the newcomer. Items in the middle absorb the displacement.

Two further properties follow from the arithmetic.

Ranks are ipsative, meaning every respondent's ranks add up to the same total. With five options, each respondent's ranks sum to 15, so the average of the five average ranks is always exactly 3.0. With eight options it is always 4.5. One item can only improve if another gets worse. An item's average rank of 2.9 in a five-item list and 4.1 in an eight-item list says nothing about whether it gained or lost ground. Only the relative order within one fixed list is interpretable.

Partial rankings distort it further. If respondents rank only their top three of eight, an item that appears in few rankings but always near the top can post an excellent average rank from a tiny base. Always check how many respondents actually placed each item before reading its average.

The three summaries to report instead of one

SummaryWhat it answersStable when options change?Weakness
First-place shareWhich item do the most people prioritise above everything else?Yes, unless the new option is itself someone's first choiceIgnores everything below first place
Pairwise preference matrixFor each pair, what share ranked one above the other?Yes; adding D never changes an A-vs-B comparisonA table of k(k-1)/2 pairs is harder to present
Average rankRoughly, where does each item sit overall?NoDepends on the rest of the list; hides polarisation

A useful pattern for a readout: lead with first-place share, support it with the pairwise matrix for the top few items, and show average rank only as an overall ordering within this study. If an item wins every one of its pairwise comparisons (a Condorcet winner, in social-choice terms), you can call it the top priority with confidence, whatever its average rank.

Polarisation is the other thing average rank hides. An item ranked first by 40 percent and last by 40 percent will post a middling average, much like an item everyone ranks in the middle. Those are opposite findings: one is a segment-defining feature, the other a nice-to-have. First-place share and a look at last-place share separate them immediately.

Design rules that prevent ranking artefacts

  1. Fix the option list for any tracked ranking. If you must add an option, report the old list and the new list separately for one wave so the shift is visible.
  2. Keep lists short. Ranking more than about seven items is hard work for respondents, and the bottom half of a long list is mostly noise. For long lists, use MaxDiff or a top tasks vote instead.
  3. Ask for a full ranking, or be explicit about partial ones. If you collect top-three rankings, report first-place share and the number of respondents who placed each item, not average rank.
  4. Do not use ranking when you need absolute levels. A ranking tells you the order, never whether the top item is actually wanted. Pair it with a scale or yes_no question when absolute appeal matters. The trade-offs are covered in ranking vs rating questions.
  5. Ask why the top choice beat the runner-up. The gap between first and second is the most decision-relevant fact in a ranking, and the ranking alone cannot tell you whether it is a chasm or a coin flip.

How Koji handles ranking questions

Koji's ranking type is one of six structured question types, alongside open_ended, scale, single_choice, multiple_choice and yes_no. In text interviews respondents drag options into order; in voice interviews the AI interviewer asks for the order conversationally. Either way the AI can probe after the answer, which is where rule 5 becomes automatic: participants are asked what separates their first choice from their second, and those explanations are themed across the whole study.

The Koji report shows the average position of each item, with every data point linked back to the conversation it came from. Each respondent's full ordering is preserved in the study data, so first-place share and the pairwise matrix can be computed directly from a data export or through Koji's MCP server. Because an AI moderator runs every interview, collecting full rankings from a larger sample costs the same researcher effort as a small one, which makes the more stable summaries practical rather than a luxury.

Ranking data is ordinal within each respondent, which places it in the ordinal row of the levels of measurement table: medians and shares are safe, and arithmetic on positions needs care.

Frequently asked questions

How do you analyse ranking question data?

Report three things: first-place share (the percentage who ranked each item first), a pairwise preference matrix (for each pair, the share who ranked one above the other), and average rank as an overall ordering within this study only. First-place share and pairwise preferences are stable when options are added or removed; average rank is not.

Why is average rank misleading?

Because it depends on which other options are in the list. In a worked example with five respondents, B led A on average rank (1.6 against 1.8) until a fourth option was added, after which A led (1.8 against 2.2), even though no respondent changed their view of A or B. Average rank also hides polarisation: an item loved by some and disliked by others can look like one everyone ranks in the middle.

Can I compare average ranks across two surveys?

Only if the option list is identical. Ranks within a respondent always add up to the same total, so the average of all items' average ranks is fixed by the length of the list: 3.0 for five options, 4.5 for eight. An item's average rank moving from 2.9 to 4.1 across lists of different lengths tells you nothing about whether it gained or lost ground.

What is a pairwise preference matrix?

A table that shows, for every pair of options, what share of respondents ranked one above the other. It is computed directly from full rankings, it never changes when an unrelated option is added, and an option that beats every other option head-to-head can be called the top priority with confidence.

How many items should a ranking question have?

Ideally seven or fewer. Beyond that, respondents struggle to order the middle and bottom of the list and those positions become noisy. For longer lists, a MaxDiff exercise or a top tasks vote collects priorities more reliably.

How does Koji collect and report ranking questions?

Respondents drag options into order in text interviews or give the order conversationally in voice interviews, and the AI interviewer can probe on why their top choice beat the runner-up. The report shows each item's average position with links to the underlying conversations, and full per-respondent orderings are kept in the study data for first-place share and pairwise analysis.

Related Resources

Related Articles

Choice and Ranking Questions in AI Interviews: Capture Preference Data at Scale

Learn how to use single choice, multiple choice, ranking, and yes/no questions in Koji AI interviews — with automatic report charts that show preference distributions across all your participants.

Levels of Measurement: Which Statistics Each Survey Question Type Allows (2026)

Nominal, ordinal, interval and ratio data explained for customer research: the summaries each level supports, how Koji's six structured question types map onto them, and a relabelling test that catches meaningless statistics.

MaxDiff Analysis: The Complete Guide to Maximum Difference Scaling (2026)

Learn how MaxDiff (Maximum Difference Scaling) produces sharper feature and message prioritization than rating scales — and how to pair it with conversational AI interviews to capture the why behind every score.

When Relabelling the Scale Reverses Which Group Scores Higher (2026)

Comparing two groups by average rating assumes the scale points are equally spaced. When the groups' answer distributions cross, an equally valid scoring reverses the result. The cumulative dominance check tells you in advance.

Ranking vs. Rating Questions: Which to Use and When

Rating questions score each item independently and scale easily; ranking questions force trade-offs and reveal true priorities. Learn the strengths, weaknesses, and biases of each, and how to choose the right format for clean, decision-ready data.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.