MisleadingCharts
Back to the gallery

The ±1-point poll that missed by 61

Showing the misleading chart

Be the first to star this exhibit

In January 1976 Ann Landers reported that 70% of parents would not have children if they had it to do over again — more than 10,000 replies, a sample so large that the textbook margin of error comes to ±0.9 points. Five months later a national random sample of 1,373 parents answered the same question and 91% said yes. The interval was pricing sampling noise, and there was hardly any. The error was in who chose to write in.

01The claim

Seven in ten American parents would not have children if they had it to do over again. The question went out in an advice column read in more than a thousand newspapers, more than 10,000 people answered it, and every reply was counted — no weighting, no adjustment, no estimation, nothing set aside. Draw that on a bar chart and the drawing is beyond reproach: a zero-based axis running the full 0% to 100%, no log scale, no second axis, and the 95% confidence interval printed on both bars, because at a sample of ten thousand the interval comes to ±0.9 points and the arithmetic leaves almost no room for doubt. Ten thousand replies is seven times the sample a national opinion poll would use. Whatever else this survey is, it is not small.

02The trick

Everything on that slide is true and the chart is drawn honestly. The problem arrives before the drawing, in how the sample was assembled: nobody chose these ten thousand people — they chose themselves. A write-in poll samples the population of readers moved to write in, and what moves a person to write in about parenthood is, fairly reliably, how parenthood went. Landers worked it out herself in a follow-up column: “the hurt, angry and disenchanted tend to write more readily than the contented.” That makes the bias directional, and it makes it immune to volume. Ten thousand replies is ten thousand draws from the wrong population, so the estimate gets sharper and stays wrong. Which is exactly what the ±0.9 fails to warn you about. The margin-of-error formula has one input, the sample size, and one assumption, that the sample was drawn at random. It prices sampling noise — how much a random draw of this size would wobble — and there is almost none at n = 10,000, so it reports a very small number and looks like a guarantee. It is not measuring the thing that went wrong, and no interval ever measures its own assumptions. The scale of the miss is the whole exhibit: Newsday commissioned a national random sample that same June and 91% of its 1,373 respondents said yes, against the column’s 30%. Sixty-one points, from a poll that published an interval of nine-tenths of one. And the demonstration runs the other way too — the Kansas City Star drew 409 people at random and got 94%, three points from the national figure on a sample one-twenty-fourth the size of the mailbag. (This exhibit is our own demonstration, drawn in the style of a 1976 newsroom graphic. The figures are the ones the four surveys reported.)

03The fix

Put all four samples on one zero-based axis and sort them by how the respondents got there, and the picture explains itself. The two write-in surveys — Landers at 30% yes, and Good Housekeeping, which reprinted the question for its own readers and reported 95% — sit 65 points apart. The two random samples sit 3 points apart, at 91% and 94%. Selection, not sample size, is what separates the pairs. Then draw the interval and the error on the same scale, because seeing a 0.9-point sliver against a 61-point span is the fastest way to learn that precision and accuracy are different quantities and only one of them improves when you collect more. Two honest caveats belong on the chart and are on ours. The write-in pair differed in wording and audience as well as in who replied — the letter Landers printed recounted friends who resented their children, the magazine’s sidebar opened by quoting her “horror”, and it addressed mothers — so their 65-point gap is a warning about opt-in data rather than a clean measurement of self-selection alone; the random pair, asked plainly, is the anchor. And Good Housekeeping never published how many people responded, which is worth noticing in its own right. The habit to keep is a question you can ask of any dataset in ten seconds: what did a person have to do to end up in here, and does doing it correlate with the answer? Then look for the response rate rather than the sample size, because the response rate is the number that bounds the damage, and if nobody published it, that is the finding. The mailbag never went away — it just moved into the rating prompt, the NPS pop-up, the post-incident survey and the exit interview, all of them answered by whoever felt like answering. Where you cannot draw a sample, label the chart with the population you actually have, and run a few hundred people chosen at random alongside it; the gap between the two will be the most useful number of your year.