Sampling & Bias: Why Polls Get It Wrong
A small fair sample beats a huge unfair one
A small fair sample beats a huge unfair one
The dots below are a school of 1,000 students. Blue means they prefer pizza, orange means they prefer tacos. The true split is 60% pizza, 40% tacos. We're going to "poll" some of them and see how close our estimate gets to the truth. Move the slider to change sample size, click "Draw Sample" to pick a fresh random group.
Adjust the slider, then draw a sample to see your estimate.
Now the same population, but you don't get to pick randomly. You can only sample from one region of the playground. Try each spot. The sample size is fixed at 100. Notice how the answer changes wildly depending on where you ask.
Random sampling gives you the truth. Biased zones give you a wrong answer with full confidence.
In 1936, a magazine called the Literary Digest mailed 10 million ballots to predict the US presidential election between Roosevelt and Landon. They got 2.4 million responses and confidently predicted Landon by a landslide. Roosevelt won 46 states out of 48. How did they get it so wrong? They mailed only to people on telephone lists, magazine subscribers, and car owners. In 1936, that was rich people. Poor people didn't have phones.
Hover or tap the bars to see which sample matched reality.
Selection bias has many forms. Each card describes a real way that real surveys go wrong. When you see a poll or stat, ask: who got asked, and who got missed?
Surveying whoever is easy to reach. Cheap, but the answers reflect that group, not the whole population.
People with strong opinions are more likely to volunteer. The result skews toward extremes, not the middle.
You only see the survivors. The people who tried the same things and failed never show up. The lesson you draw from survivors is incomplete.
The ones who answer might not look like the ones who don't. Busy people, angry people, and shy people drop out at different rates.
Your sampling method excludes whole groups: people without phones, people not on the app, kids who don't go to that school. The missing groups can be the answer.
Even a perfect random sample can be biased by leading questions. Wording can shift answers by 20% or more. Pollsters test wording carefully (good ones, anyway).
During World War II, the US Air Force studied bullet holes in returning bombers to figure out where to add armor. The intuitive answer: armor the spots with the most holes. The right answer, from statistician Abraham Wald: armor the spots with the fewest holes. Why? Those were the spots that, when hit, made planes crash. The bombers that returned were the survivors. The full sample of crashed planes was missing from the data. That's survivor bias in one famous example.
You've learned the core ideas behind every survey, poll, and study. A sample is a small group used to estimate a big group. Bigger samples are more accurate, but only if they're drawn fairly. A biased sample, no matter how big, will lie. Random sampling is the engine behind honest polling and good science.
You can't ask every kid in the country what flavor ice cream they like. Instead you ask a smaller group (a sample) and use their answers to guess the truth about everyone.
A sample of 10 might be off by 30%. A sample of 1,000 is usually within 3%. Doubling sample size cuts error by about 30%, not in half. There are diminishing returns.
Who you ask changes the answer. If you only survey kids in the ice cream line, your "everyone loves ice cream" finding is meaningless. The sample wasn't fair.
Pick people randomly from the entire population. Every person has an equal chance of being chosen. A random sample of 1,000 beats a biased sample of 10 million.
Put your new knowledge into practice!