The word cloud that pays by the letter
Showing the misleading chart
Every word of the federal wine-labelling regulation, counted, then drawn as a word cloud with each word sized so its area would match how often it appears. Measure the shapes it actually drew and they track frequency times letter count instead: “TTB” occurs 126 times and “designation” 74, and the cloud draws “designation” two and a half times bigger. Of forty words, two kept their place.
01The claim
Here is what the federal wine-labelling rulebook is really about. Every word of 27 CFR Part 4 — 22,434 of them — has been counted, common English words taken out, and the rest drawn as a word cloud. And the slide can say something most charts cannot: there is no axis on it to truncate, no baseline to move, no window to crop, no colour scale to gerrymander. The whole part is in the picture, every word of it, and hue carries no meaning at all — only size encodes the data. Better still, the size rule is the careful one: font size set to the square root of the count, so that each word’s area, not its height, rises in proportion to how often it appears. Read the picture and the finding is plain. “Statement” is the second-largest word, “designation” the third, “information” the fifth, “appropriate” the sixth. This is the vocabulary of a disclosure regime — a rulebook preoccupied with what has to be stated. Meanwhile the agency’s own names, TTB and ATF, sit among the smallest words on the slide. Procedure, apparently, is not where this regulation spends its attention.
02The trick
The counts are right, the text is complete, and the sizing rule is the better of the two in common use. The trouble is that the rule can only reach one of the two dimensions the reader is measuring. Setting font size to the square root of the count controls the em square — the box one character sits in. But a word is not one character. Its footprint on the page is that square laid down once per letter, so the area the eye actually takes in comes out as letters × count. Length is doing half the work, and length is not data about anything; it is spelling. Measure the forty rectangles the cloud drew and they correlate at r = 0.99 with letters × count, and only 0.89 with the count on its own. What that does to the ranking is not subtle. “TTB” appears 126 times, the fourth-commonest word in the part; three letters put it 28th by area. “ATF” appears 87 times, eighth-commonest; it lands 37th of 40, the fourth-smallest thing on the slide. Coming the other way, “statements” appears 44 times — 35th — and lands 10th; “appellation” appears 43 times, 36th, and lands 12th. Of the forty words, two kept their place, twelve moved ten places or more, and rank agreement between the two orderings comes to ρ = 0.53. So the read-out is backwards on its own evidence. The rulebook says “TTB” and “ATF” 213 times between them and says “designation” 74 times; the picture reports the reverse, because one set of words is short and the other is long. The same arithmetic quietly flattens the leader: “wine” occurs 3.8 times as often as “statement”, 515 against 136, and covers only 1.8 times as much of the page. Everything else on the slide is true, and none of it helps — there is indeed no axis to truncate, and that is precisely the difficulty, because with no axis there is nothing to check the sizes against. (This slide is our own demonstration, drawn in the manner of a consulting deck. The word counts are real, computed from the regulation’s published text.)
03The fix
Put the counts on a zero-based axis, sorted, one length per number, and the picture stops arguing with itself: “wine” at 515 with everything else in a long flat tail — which is what free text almost always looks like, and what a cloud can never show, because it has no axis on which a tail could be flat. The slopegraph is worth drawing once for the education: rank the words by frequency down one side, by measured area down the other, join them up, and watch the three-letter words dive while the eleven-letter words climb. Then print the numbers, because a cloud’s central problem is that it makes a quantitative claim and supplies nothing to check it with. Say which stop-word list you used, too. Ours removed “shall”, “may” and “must” as ordinary English — in a regulation those are the three words that do the work, and taking them out is an editorial act, not a technical one. “Shall” alone occurs 170 times in this part, more often than any word left in the picture except “wine”; the list deleted the second-biggest word in the cloud before the cloud was drawn. Rotation, the layout seed and the colours are the same kind of undeclared choice: rotate a word and it reads smaller, re-run the layout and the picture rearranges, and the hues here mean nothing at all, which readers reliably refuse to believe. The habit worth keeping is a ten-second check that works on any cloud you meet. Find the longest word in it and look up its count. Then find the shortest word you can still read, and look up that one. If the picture has been ranking by letters, those two numbers will tell you straight away. And if the real question is which words travel together rather than how often each appears, it was never a frequency question — reach for phrases, co-occurrence, or a dozen quoted responses. Where a cloud is wanted as ornament, let it be ornament, sitting beside the chart that carries the numbers rather than standing in for it.