
Framing & context
The rate at the wrong grain
Ninety-nine per cent per word is eighty-six per cent per sentence.
a.k.a. per-word accuracy · word error rate · the fine-grained denominator · per-item versus per-episode · uptime nines · five nines · parts per million · the compounding rate · accuracy per unit · the small denominator
A rate is a fraction, and the thing in its denominator is a choice. Pick a small enough unit — a word, a character, a packet, a minute, a part — and any error rate becomes a number just under 100%, because the smaller the unit, the less of anything can go wrong inside one of them. But nobody consumes the small unit. A reader reads sentences, a viewer watches programmes, a customer places an order, a patient has a stay, and each of those is made of many small units in a row, every one of which has to come out right. So the rate the reader cares about is the per-word rate compounded across the length of the thing, and that number can be worlds away from the headline while both are computed from the same measurements with no error anywhere. It is not a distortion of the data and it survives every audit of the drawing: the axis can start at zero and run to a hundred, every case can be present, nothing smoothed or dropped, and the picture is still answering a question nobody asked. Two features make it hard to catch. A rate near 100% has almost no room left to move, so real differences arrive as hundredths of a point and read as noise, while the same differences in the complement — the error rate — are whole multiples of each other. And the compounding is exponential, so the gap between the quoted rate and the experienced one widens with the length of the thing, fastest exactly where the stakes are highest.
How to spot it
- Read the denominator out loud. “Per word”, “per character”, “per packet”, “per part”, “per minute”, “per component” — then ask what unit the person affected actually receives, and how many of the small ones are inside it.
- Do the compounding in your head. A rate p per unit across n units leaves (1 − p)ⁿ intact, which is roughly 1 − np while np is small: 1% per word over a 15-word sentence is about 14% of sentences touched, and over a 100-step process it is 63%.
- Any rate quoted with a 9 in front of it. 99%, 99.9%, “four nines”, “six sigma”, “parts per million” — the nines are a sign that the unit was chosen small enough to produce them, and the useful question is always how many of that unit make up one of the things that has to work.
- Plot the complement instead. A series moving from 99.2% to 99.6% is a flat line at the top of any zero-based axis; the error rate behind it halved, which is the largest change the measure can record.
- Watch for a threshold set in the fine unit and a consequence felt in the coarse one. “98% accurate” as a pass mark says nothing about how often a viewer, a reader or an operator meets the failure, and the two can be moving in opposite directions.
- Check whether the errors are weighted. A score that discounts small errors and counts serious ones in full is not counting errors at all; its complement is a weighted quantity, and the number of actual mistakes is larger than it by however much the discount was.
- Ask what fell out of the denominator on the way in, because the two failures ride together: a per-word accuracy computed over the words that were delivered has no opinion about the words that were dropped. That second half belongs to the late denominator, and a rate can be at the wrong grain and start counting late at the same time.
- A familiar grain is the fastest test there is: convert the same percentage into time. 99% of a year is 3 days 15 hours of outage, 98% is 7 days 7 hours, and a figure that sounds like an A in one unit sounds like a crisis in another without a single number changing.
The fix
Quote the rate at the grain of the thing that has to come out right, and name the grain the way you print a unit. Where the natural unit is small because that is how the measurement is made — and often it is, since a word error rate is exactly the right instrument for comparing two recognisers — keep it, and publish the compounded figure beside it: per word and per sentence, per step and per run, per component and per assembly. The second number is one line of arithmetic and it is the one the reader was trying to work out. Plot the complement rather than the rate, on a zero-based axis, because errors per hundred words has room to halve where 99.2% has nowhere to go; the two are the same measurement and only one of them can show a change. Put a count next to the ratio, since a count is already at a grain a person can hold — fifty errors in a ten-minute clip settles what 98.55% leaves open. Be explicit about what the compounding assumes, because independence is rarely true: real errors cluster, which packs more of them into fewer items and makes the naive calculation an upper bound on how many items are touched, so give the bound and say which way it leans rather than pretending to a precision the assumption cannot carry. And when a threshold is being set, set it in the unit the consequence lands in. A pass mark written in the fine grain is a pass mark nobody has costed, and the whole trick lives in the gap between the unit that was measured and the unit that was lived in.
In the gallery

