{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-01T04:33:05.009Z"},"content":[{"type":"documentation","id":"bda6dc05-1b8a-4b1f-92ed-428c654766c9","slug":"network-scale-up-method-prevalence","title":"The Network Scale-Up Method: Estimating How Many, by Asking Who They Know (2026)","url":"https://www.koji.so/docs/network-scale-up-method-prevalence","summary":"The network scale-up method estimates the size of a group that will not identify itself by asking ordinary respondents how many members they know and scaling that up by their total personal network size. The relation is m/c = e/t, so e = t x (m/c), where m is people known in the hidden group, c is personal network size, e is the hidden group size and t is the known population. Personal network size is measured with the known population method, using groups whose size is already known; one study used eighteen scaling variables and found an average network size of 604.03. In a worked example, 50,000 employees, a network size of 220 and 3.1 known members give an estimate near 700. Three assumptions govern it: transmission effects, barrier effects, and recall accuracy, and the first makes every estimate a lower bound.","content":"When a group will not identify itself, you cannot sample it. The network scale-up method sidesteps that by never asking anyone about themselves: it asks ordinary respondents how many people they know in the target group, then scales that answer up using an estimate of how many people they know in total.\n\n## The short answer\n\n[Randomized response](/docs/randomized-response-technique-research) and [list experiments](/docs/list-experiment-item-count-research) both estimate what share of your respondents do something. Neither helps when the people you care about will not enter your study at all.\n\nThe network scale-up method attacks that problem from outside. Its premise, as stated in the methodological literature, is that the proportion of people in an average individual's personal network who are members of a given subpopulation is indicative of the relative size of that subpopulation to the general population as a whole.\n\nIn other words: if the average person knows 220 people, and knows 3.1 people in the target group, then the target group is roughly 3.1/220 of the population. Multiply by the population size and you have an estimate of how many there are.\n\n## The idea in one equation\n\n### What the equation says\n\nThe core relation is:\n\nm / c = e / t\n\nwhere m is the number of people a respondent knows in the hidden group, c is the respondent's total personal network size, e is the size of the hidden group, and t is the total population. You measure m and c from your survey, you know t, and you solve for e:\n\ne = t x (m / c)\n\n### Why this works when direct questions do not\n\nNobody is ever asked to disclose anything about themselves. A respondent reporting that they know three people who have quietly abandoned a tool is not confessing to anything, is not implicating any named individual, and has no incentive to shade the number. The sensitivity that wrecks a direct question is simply absent from the question being asked.\n\nThat also means the method reaches groups that are structurally unreachable by sampling. You do not need a single member of the hidden group to take part.\n\n## The two numbers you need\n\n### The known population method for network size\n\nThe weak point is c, the respondent's total network size. People do not know how many people they know, and asking them produces a guess.\n\nThe standard solution is the known population method. You ask about several groups whose size you already know, and work backwards. If a group makes up 1.2% of the population and the respondent knows 2.64 of them, their network is about 2.64 / 0.012 = 220 people.\n\n### The scaling groups that make it work\n\nOne known group gives a noisy estimate, so real studies use many and average them. In the study that produced the estimator refinements cited here, eighteen scaling variables were used - twelve first names and six professions - and the average personal network size came out at 604.03 under the original estimator without weights, falling to 464.28 once the authors applied their recursive refinement. Two estimates of the same quantity, the larger about 30% above the smaller, is a fair warning about how much weight c can bear.\n\nFor a study inside a single company you have an unusual advantage: you know the exact size of many internal groups. Headcount by office, by department, by tenure band, by job family. These make far better scaling groups than first names, because the sizes are exact rather than estimated from census data.\n\n## A worked example\n\nYou want to know how many people at a 50,000-person customer have stopped using your product while keeping their seat. They will not tell you, and their manager will not tell you either.\n\n### Step 1: estimate network size\n\nYou ask a sample of employees how many people they know in a department you know has 600 people. They report knowing 2.64 on average. That department is 600/50,000 = 1.2% of the company, so:\n\nc = 2.64 / 0.012 = 220\n\n### Step 2: ask about the hidden group\n\nYou ask the same respondents: how many people do you know at this company who have effectively stopped using the tool? They report 3.1 on average.\n\n### Step 3: scale up\n\n| Quantity | Value |\n| --- | --- |\n| Total population (t) | 50,000 |\n| Known in hidden group (m) | 3.1 |\n| Personal network size (c) | 220 |\n| Share of network in group | 1.41% |\n| Estimated group size (e) | about 700 |\n\n50,000 x (3.1 / 220) = 704.5, so roughly 700 people - about 1.4% of the account.\n\n### Reading the result with appropriate humility\n\nThis is an order-of-magnitude instrument, not a precise one. The honest readout is \"several hundred, likely between 400 and 1,000\", not \"704\". The value is that you moved from no number at all to a defensible range, which is usually the difference between a conversation that can happen and one that cannot.\n\n## The three assumptions, and how each one breaks\n\nEach assumption below is what the method requires. The corresponding bias is what you get when it fails.\n\n### Transmission effects\n\nThe method assumes everyone is fully aware of the characteristics that define a given subpopulation. As the source puts it, if someone in their personal network did go to prison but the respondent is unaware, the result would be an underestimate.\n\nThis is the dominant bias for most product questions, and it always runs the same direction. People do not announce that they have quietly stopped using a tool. Your estimate is a floor, and you should present it as one.\n\n### Barrier effects\n\nThe method assumes everyone in the larger population has an equal probability of knowing someone in a given subpopulation. The source illustrates this with geography: respondents who live in Montana may have different probabilities of knowing a shark attack victim than a respondent living in Florida.\n\nInside a company, the barriers are departmental and hierarchical. Engineers mostly know engineers. If the behavior you are estimating clusters in one function and your respondents cluster in another, your estimate is wrong in a direction you cannot sign without knowing the network structure. Spread your sample across functions deliberately.\n\n### Recall and estimation error\n\nThe method assumes respondents can correctly recall the number of people that they truly know in the subpopulation, and the source notes that different survey modes could be expected to result in different recall effects.\n\nIn practice people estimate rather than count, and they compress large numbers. This is why the known population method matters: it measures network size through the same faulty recall process that produces m, so some of the error cancels.\n\n## When to reach for this\n\n### Good fits\n\nGroups defined by something visible to colleagues but not to you. Silent abandonment, unofficial workarounds, people doing a job the org chart does not name. Also any question where you need a population size rather than a rate, and where no sampling frame for the group exists.\n\n### Bad fits\n\nAnything genuinely invisible to a respondent's network - private opinions, individual beliefs, feelings about your pricing. If people would not know it about their colleagues, the transmission assumption fails completely and the method returns noise. For those, use a list experiment instead.\n\nIt is also a poor fit when you need precision. This method will not distinguish 4% from 5%.\n\n## Running it with Koji\n\nThe network scale-up method needs several numeric estimates from each respondent, and it needs them to be considered rather than guessed.\n\nKoji's structured questions cover six types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` - and each \"how many people do you know who...\" item is captured as a discrete numeric value, so the scale-up arithmetic runs on clean data instead of transcribed prose.\n\nThe part that a static survey handles badly is recall quality. A respondent who answers \"about 5\" to every scaling question has told you nothing. Because Koji's AI interviewer is conversational, it can ask the respondent to think through a specific group before answering, and follow up when an answer looks like a reflex rather than an estimate. That is a genuine improvement on a form field, and it directly attacks the recall error that is this method's largest weakness.\n\nRunning the design across several scaling groups also means a longer interview than most teams would field by hand. Koji runs them in parallel and aggregates the numeric answers automatically, which is what makes an eighteen-question scaling battery practical rather than theoretical.\n\nKoji cannot repair a broken transmission assumption. If colleagues genuinely cannot observe the behavior, no amount of interviewing quality will rescue the estimate.\n\n## A working procedure\n\n1. Confirm the behavior is observable by colleagues. If it is not, stop and use a different method.\n2. Pick at least five scaling groups whose exact size you know from internal records.\n3. Write each item as \"how many people do you know who...\", using a consistent definition of \"know\".\n4. Field to a sample spread deliberately across departments and levels.\n5. Compute c from each scaling group, average them, then apply e = t x (m / c).\n6. Report a range, label it a floor, and name the transmission assumption in the readout.\n\n## Frequently asked questions\n\n### What does \"know\" mean in a network scale-up question?\n\nYou must define it explicitly and use the same definition throughout, because c and m have to be measured on the same scale. A common choice is that you know them by name and have interacted in the past two years. Any definition works as long as it is constant.\n\n### Why is my estimate always too low?\n\nBecause of transmission error. Respondents cannot report people whose membership in the group they are unaware of, and that failure only ever removes people from the count. A network scale-up estimate should be read as a lower bound.\n\n### How large a sample does the network scale-up method need?\n\nIt needs fewer respondents than a list experiment, because every respondent contributes information about many people rather than one. A few hundred well-spread respondents is a reasonable starting point for an account-level estimate.\n\n### Can I use this to size a group outside my own customers?\n\nYes, and that is its original use in public health, where it estimates the size of populations no registry covers. You need a defensible figure for the total population t and scaling groups whose sizes you genuinely know.\n\n### How is this different from a key informant interview?\n\nA key informant gives you a considered judgment about a group. The network scale-up method gives you an arithmetic estimate built from many people's mechanical counts, none of whom is asked to form a judgment at all.\n\n### Can Koji run a network scale-up study?\n\nYes. Capture each count as a structured numeric question, let the AI interviewer press for a considered answer rather than a reflex, and apply the scale-up arithmetic to the aggregated results.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types, including the numeric capture this method depends on\n- [Randomized Response](/docs/randomized-response-technique-research) - when the people you need are inside your sample but will not answer\n- [The List Experiment](/docs/list-experiment-item-count-research) - the better choice for behavior colleagues cannot observe\n- [Key Informant Interviews](/docs/key-informant-interviews-research) - why one person cannot speak for a whole company\n- [Snowball Sampling](/docs/snowball-sampling-guide) - reaching hard-to-find participants by referral instead of by estimate\n- [Hard-to-Reach Audiences](/docs/hard-to-reach-participants-research) - recruiting the people this method only counts\n- [Why Complaint Counts Cannot Become Rates](/docs/complaint-counts-cannot-be-rates) - the missing denominator this method sometimes lets you estimate\n","category":"Research Methods","lastModified":"2026-09-30T03:37:48.624028+00:00","metaTitle":"Network Scale-Up Method for Estimating Hidden Groups","metaDescription":"Estimate a hidden group size by asking how many members people know. The equation, a worked example, and the three assumptions.","keywords":["network scale-up method","NSUM","estimating hidden population size","how many people do you know method","prevalence estimation","sizing a group that will not self-identify"],"aiSummary":"The network scale-up method estimates the size of a group that will not identify itself by asking ordinary respondents how many members they know and scaling that up by their total personal network size. The relation is m/c = e/t, so e = t x (m/c), where m is people known in the hidden group, c is personal network size, e is the hidden group size and t is the known population. Personal network size is measured with the known population method, using groups whose size is already known; one study used eighteen scaling variables and found an average network size of 604.03. In a worked example, 50,000 employees, a network size of 220 and 3.1 known members give an estimate near 700. Three assumptions govern it: transmission effects, barrier effects, and recall accuracy, and the first makes every estimate a lower bound.","aiDifficulty":"advanced","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}