Survey Weighting: How to Correct a Skewed Sample
A practical guide to survey weighting — post-stratification, raking, and propensity weighting — plus how to calculate design effect and effective sample size, and when weighting cannot save your data.
Weighting makes your sample look like your population by multiplying each respondent by a factor. It fixes composition — who answered — and nothing else. If 10% of your respondents are enterprise admins but 20% of your customer base is, each admin respondent gets a weight near 2 and each over-represented casual user gets a weight below 1. What weighting cannot do is tell you how the customers who never answered would have replied. That distinction is the whole subject.
This guide covers the three weighting families you will actually use, the arithmetic for design effect and effective sample size, when to trim, and — most usefully — how to design studies that need less weighting in the first place.
When your sample actually needs weighting
Weighting is worth the complexity when three things are all true:
- You know the true population distribution. For customer research this is usually straightforward: your own CRM or product database tells you the real split by plan, region, tenure, or role. This is a large advantage over public-opinion polling, where the population totals themselves are estimated.
- Your sample is meaningfully skewed on that variable. A 2-point gap is noise. A 15-point gap changes conclusions.
- The variable is related to what you are measuring. Weighting on a variable that has no relationship to your outcome adds variance and corrects nothing.
If you cannot satisfy point 1, you are not weighting — you are guessing with extra steps. If you fail point 3, you are paying precision for decoration.
| Situation | Weight? | Better first move |
|---|---|---|
| Enterprise customers are 3x over-represented | Yes | Weight down, and check whether the enterprise skew also biased the questions asked |
| 60% of respondents came from one email blast | Usually | Fix the recruiting channel mix on the next wave |
| Sample matches population within 2-3 points | No | Report unweighted; note the composition |
| You have 14 respondents in a segment of interest | No | Do not weight a cell up from 14. Field more interviews |
| Population distribution is unknown | No | Report the sample composition explicitly and scope your claims to it |
That fourth row is the one teams get wrong most often. Weighting a tiny cell up by a factor of 6 does not create data; it just makes six people's opinions load-bearing for a company decision.
The three weighting methods
1. Post-stratification (cell weighting)
You know the population count for every combination of your variables, so you weight each full cell directly.
Weight for a cell = (population share of that cell) / (sample share of that cell)
Say you research three user roles:
| Role | Population share | Sample (n=200) | Sample share | Weight |
|---|---|---|---|---|
| Admin | 20% | 80 | 40% | 0.50 |
| Power user | 30% | 70 | 35% | 0.86 |
| Casual user | 50% | 50 | 25% | 2.00 |
After weighting, the sum of weights is still 200 (40 + 60 + 100), but the weighted composition is 20 / 30 / 50 — exactly the population.
Post-stratification is the cleanest method and preserves relationships between variables. Its limitation is combinatorial: with role (3) x region (4) x plan tier (3) you need reliable population counts for 36 cells, and some of those cells will contain two respondents.
2. Raking (RIM weighting)
Raking is the workhorse. It needs only the marginal totals — the population split by role, and separately by region, and separately by plan — not the joint distribution.
The algorithm, iterative proportional fitting, is simple: adjust the weights so role matches, then adjust so region matches (which breaks role slightly), then adjust so plan matches, then loop back to role. Repeat until every margin is within tolerance, usually in five to twenty passes.
Use raking when you have three or more weighting variables, or when your cross-tab has cells too thin to weight directly. It is the default in every serious survey package for good reason.
3. Propensity weighting and matching
You model each respondent's probability of responding using a reference dataset, then weight by the inverse of that propensity. This is the method reached for when the skew is behavioural rather than demographic — for instance, when heavy product users answer research invitations far more readily than light ones.
It is also the method with the weakest evidence behind it relative to its complexity. Pew Research Center's benchmark study of online opt-in samples found that more complex statistical methods never reduced average estimated bias by more than 0.3 percentage points beyond raking, the simplest technique tested. Machine-learning adjustment did not rescue a bad variable set.
What weighting can and cannot fix
This is the section to send to the stakeholder who says "just weight it."
Pew's study measured 24 benchmarks against high-quality government data. Unweighted online opt-in samples showed an average bias of 8.4 percentage points. The most effective adjustment strategy brought that down to about 6 points — removing roughly 30% of the original bias. Seventy percent survived weighting.
Two findings from that work should change how you weight:
- Variables matter more than method. For political-engagement benchmarks with an average unweighted bias of 22.3 points, weighting on demographics alone cut bias by 2.9 points; adding politically relevant variables cut a further 8.8. Same arithmetic, radically different result, purely from variable choice. The customer-research equivalent: weighting on region and company size will do far less than weighting on tenure, plan tier, or usage decile, because those are what actually predict your outcome.
- More sample does not fix a biased sample. Going from n=2,000 to n=8,000 improved average bias by 0.2 points. Bias and sample size are independent problems, and only one of them is solved by spending more money.
The practical rule: weight on the variables that predict your outcome, not the variables that are easy to find. If you are measuring willingness to pay, tenure and current plan tier are the weighting variables that matter. Job title is decoration.
Design effect: what weighting costs you
Weights are not free. Unequal weights inflate the variance of your estimates, which is measured by the design effect and expressed as effective sample size using Kish's formula:
Effective n = (sum of weights)^2 / (sum of squared weights)
Take the role example above:
- Sum of weights = (80 x 0.50) + (70 x 0.86) + (50 x 2.00) = 200
- Sum of squared weights = (80 x 0.25) + (70 x 0.735) + (50 x 4.00) = 271.4
- Effective n = 200^2 / 271.4 = 147
Your 200 interviews now carry the statistical precision of 147. The design effect is 200 / 147 = 1.36, and the margin of error on a 50% estimate widens from about +/- 6.9 points to about +/- 8.1 points.
| Weight spread | Approx. design effect | 200 interviews are worth |
|---|---|---|
| Nearly equal (0.9-1.1) | 1.01 | 198 |
| Moderate (0.5-2.0) | 1.3-1.4 | 145-155 |
| Wide (0.3-4.0) | 1.8-2.2 | 90-110 |
| Extreme (0.1-10) | 3.0+ | Under 70 |
Report effective sample size, not raw sample size, whenever you present weighted results. A deck claiming "n=200" from a sample with a design effect of 2.2 is overstating its own confidence by about 50%.
Trimming weights
When a handful of respondents carry weights of 8 or 12, a single unusual answer can swing a headline number. Trimming caps the extremes:
- Cap at a multiple of the mean — 3x to 5x is conventional.
- Cap at percentiles — trim to the 1st and 99th percentile of the weight distribution.
- Redistribute and re-rake — after trimming, re-run the raking so margins still balance approximately.
Trimming deliberately reintroduces a little bias to buy a lot of stability, and it is almost always the right trade at the extremes. Two rules: decide the trimming threshold before you look at the outcome variable, and always disclose it. Choosing a trim point after seeing which one produces the nicer number is not analysis.
Reporting weighted results honestly
Every weighted deliverable should state:
- Which variables were used for weighting and the population source for the targets
- The method (post-stratification or raking) and the number of iterations to convergence
- The weight range after trimming, and the trimming rule
- Effective sample size and design effect, alongside raw n
- Any segment whose unweighted base is under 30, flagged as indicative only
If a segment's unweighted base is under 30, do not report a weighted percentage for it at all. Report the count.
Designing studies that need less weighting
Here is the shift that matters more than any of the arithmetic above. Weighting is a repair applied after fieldwork, and every repair has a cost. The alternative is to make fieldwork cheap enough that you simply keep going until the thin cells are full.
This is where AI-moderated research changes the economics. A traditional moderated interview costs a scheduled hour of a researcher's time, so once you have 200 sessions you stop and weight whatever you got. Koji's AI interviewer runs voice and text conversations concurrently and around the clock, with no moderator to schedule — so extending fieldwork to fill an under-represented segment costs credits, not calendar weeks.
Concretely:
- Screen and quota at the door. Koji's structured questions — six types: open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — let you capture role, tenure, plan, and region as machine-readable fields at the start of every interview rather than inferring them afterwards. Those fields are what your weighting or quota logic runs on, and having them structured is the difference between a five-minute weighting job and a week of transcript coding.
- Watch composition live. Because analysis runs continuously rather than at the end, you can see at n=60 that casual users are under-filled, and re-target recruiting while the study is still open. That is a quota fix, which costs nothing statistically, instead of a weighting fix, which costs precision.
- Keep the qualitative depth while you do it. The AI interviewer asks its own follow-up questions on open-ended items, so filling a quota cell does not mean dropping to a tick-box survey to do it cheaply. You get the structured variable and the reasoning behind it from the same conversation.
- Weight only the residual. After quota management, the leftover skew is usually small — weights in the 0.8 to 1.3 range, a design effect near 1.05, and no awkward conversation about effective sample size.
Traditional survey tools like SurveyMonkey, Typeform, or Qualtrics will happily field a skewed sample and hand you a spreadsheet to weight. Platforms like Koji automate the part that actually improves accuracy: seeing the skew early and closing it with real conversations.
Common weighting mistakes
- Weighting on variables unrelated to the outcome. Adds design effect, corrects nothing.
- Weighting up microscopic cells. A weight of 9 on 11 respondents is not a finding.
- Forgetting to re-weight after cleaning. If you remove speeders and straight-liners after building weights, the margins no longer hold. Clean first, weight second.
- Reporting raw n on weighted results. Overstates confidence, sometimes badly.
- Treating weighting as a substitute for representative recruiting. It removes about 30% of bias. Recruiting removes the cause.
- Weighting a qualitative study. Twelve interviews do not become representative because you multiplied four of them by 2.3. Report them as what they are.
Related Resources
- Non-Response Bias — the problem weighting only partly solves
- Quota Sampling Guide — filling cells during fieldwork instead of weighting after it
- Stratified Sampling Guide — building representativeness into the design
- Survey Margin of Error Guide — how effective sample size changes your error bars
- Survey Sample Size Guide — how many responses you need before weighting is even sensible
- Structured Questions Guide — capturing the weighting variables as clean, machine-readable fields
- Survey Data Quality Guide — cleaning before you weight
Frequently asked questions
What is survey weighting in simple terms? Survey weighting assigns each respondent a multiplier so that the sample's composition matches the population you are trying to describe. If 20% of your customers are enterprise admins but only 10% of your respondents are, each admin respondent counts roughly twice in the weighted results.
What is the difference between post-stratification and raking? Post-stratification weights on the full cross-tabulation of your variables — you need the population count for every combination (enterprise admins in EMEA, SMB casual users in North America, and so on). Raking, also called RIM weighting, only needs the marginal totals for each variable separately and iterates until all margins match simultaneously. Raking is used far more often because marginal totals are much easier to obtain.
Does weighting fix non-response bias? Only partially. Weighting corrects the composition of who answered on the variables you weight on. It cannot correct differences between responders and non-responders within a weighting cell. Pew Research Center found that even the most effective adjustment strategy removed only about 30% of measured bias in online opt-in samples, reducing average bias from 8.4 percentage points to about 6.
How do I calculate effective sample size? Use Kish's formula: effective n equals the square of the sum of the weights divided by the sum of the squared weights. Divide your raw sample size by the effective sample size to get the design effect. A design effect of 1.4 means your 200 interviews carry roughly the precision of 143 unweighted ones.
When should I trim survey weights? Trim when a small number of respondents carry very large weights — a common rule is to cap weights at 3 to 5 times the mean, or at the 1st and 99th percentiles. Trimming deliberately trades a small amount of bias for a large gain in stability. Always report that you trimmed and at what threshold.
Can Koji weight results automatically? Koji reports the composition of your respondent pool against the screener and structured-question data it collects, so you can see which segments are under-filled while the study is still running. Because AI-moderated interviews run around the clock at a fraction of the cost of a moderated session, the better move is usually to keep fielding until the thin segments fill rather than to weight a skewed sample after the fact.
Related Articles
Nonresponse Bias: How Missing Respondents Skew Your Data
Nonresponse bias occurs when the people who do not answer your survey differ systematically from those who do. Learn why a low response rate is not the same as bias, how to detect it, and how to reduce it.
Quota Sampling: A Practical Guide to Getting a Representative Sample
What quota sampling is, when to use it, how to set quotas, and how it differs from stratified and convenience sampling. Includes a step-by-step workflow and how to enforce quotas with screeners and structured questions.
Statistical Significance in Survey Research: A Plain-English Guide (2026)
A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.
Stratified Sampling: How to Get Precise, Representative Results (2026)
A complete guide to stratified random sampling — how it works, proportionate vs disproportionate strata, the calculation, why it beats simple random sampling on precision, and how to run it in modern research.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Data Quality: How to Detect and Prevent Bad Responses (2026)
The threats that corrupt survey data — straightlining, speeding, bots, fraud, and inattentive respondents — how to detect and prevent each, and why conversational AI interviews are structurally resistant to the junk that plagues panel surveys.
Margin of Error in Surveys: What It Means and How to Calculate It (2026)
A plain-English guide to survey margin of error — the formula, a worked example, what changes it, common misreadings, and why AI-moderated interviews sidestep the breadth-vs-depth trade-off entirely.
Survey Sample Size: How Many Responses Do You Really Need? (2026 Guide)
A practical guide to survey sample size — formulas, calculators, real benchmarks by use case, and why AI-moderated interviews change the qual-vs-quant tradeoff entirely.