MisleadingCharts
Back to the gallery

The 98.55% pass that is an error every twelve seconds

Showing the misleading chart

Be the first to star this exhibit

A quality scorecard we drew takes the regulator’s final measurement of live television subtitling — 84 ten-minute clips, five broadcasters, three genres, 14 hours, none dropped — and plots the accuracy rates on a linear axis from zero to 100%. Four rounds, rising every time, 98.55% at the end of it, and 80% of programmes clear the 98% acceptability threshold. The board reads a pass and carries it forward unchanged. Nothing on the slide is bent. The denominator is one word of subtitle, and nobody watches television one word at a time: at the sample’s own 133 words a minute the missing 1.45% is 1.93 error points a minute, about 19 points in a ten-minute clip. And a point is not an error — the model discounts a minor one to 0.25 and a standard one to 0.5, and the same reviewers published the mix, 56% minor, 39% standard, 5% serious. Put the discount back and the pass is about 50 errors in that clip, one every twelve seconds, with an error in one subtitle in six.

01The claim

Live subtitling passed, and it is getting better. Every programme in the regulator’s final sampling exercise is on the slide — 84 ten-minute clips drawn from five broadcasters across news, chat shows and entertainment, 14 hours of television, about 145,000 words and almost 23,000 subtitles, none dropped. Each was scored on the same published accuracy model and re-scored independently by a university review team, whose figures differed from the broadcasters’ own by 0.07 points on average. The accuracy axis is linear, starts at zero and runs to 100%; nothing is truncated, broken or clipped. Four rounds of measurement, rising every time: 98.28%, then 98.34%, then 98.37%, and now 98.55% — the highest the project recorded. News bulletins rose in step, 98.49 to 98.62 to 98.86 to 99.02. Eighty per cent of programmes cleared the 98% acceptability threshold, against 76, 74 and 77 in the rounds before. Every genre is above the line: news 99.02%, entertainment 98.31%, chat shows 98.26%. The thinnest margin on the card is more than three times the average gap between the broadcasters’ own scoring and the reviewers’. Read-out for the board: quality is good and improving on every measure the model reports. Recommendation: keep the present measurement régime, report the headline figure each quarter, and hold the next review to the same model. Exceptions raised: none. Actions arising: none. Carried forward to next quarter unchanged.

02The trick

The drawing is not doing anything. One model, one unit, one population; a percentage axis that is linear and starts at zero; the threshold drawn exactly where the reviewers set it; every programme measured present in the average and nothing smoothed, logged, indexed or rescaled. The slide survives any audit of its geometry, because what is wrong with it is not in the picture — it is in the denominator. The score is the NER model, (N − E − R) / N, where N is the number of words in the subtitles and E and R are editing and recognition errors. The unit is one word. Nobody watches television one word at a time. Take the missing 1.45% and price it in a unit a viewer actually receives. At the sample’s own average subtitling speed of 133 words a minute, 1.45% of words is 1.93 error points every minute — about 19 points in each ten-minute clip. And a point is not a mistake: the model weights errors by severity, 0.25 for a minor one, 0.5 for a standard one, 1 for a serious one, so 19 points could be anything from 19 mistakes to 77. It does not have to be a range, because the same reviewers counted the mix and published it: 56% of the errors in this sample were minor, 39% standard and 5% serious, which averages 0.385 of a point each. Put the discount back and the 19.3 points are about 50 errors — one every twelve seconds. The discount lives in the score and never in the viewing. Compound it across the thing being read and the headline comes apart further. A subtitle in this sample is about 4.86 words — 23,000 subtitles across 14 hours at 133 words a minute — so at the published mix 83.0% of subtitles have nothing wrong in them, which puts an error in one subtitle in six. Across a fifteen-word sentence it is 56.2% clean: nearly one sentence in two carries an error. The bounds either side are the score’s own, 93.2% and 74.8% for a subtitle, 80.3% and 40.8% for a sentence, and they are what it allows if every error were of a single severity. The same arithmetic in a unit everybody already has a feel for: 1.45% of a year is five days and seven hours, and the 98% pass mark is seven days and seven hours. Nobody grades a service 98.55% available and calls it an A. The genre ranking moves too, because each genre is subtitled at its own speed: news is the most accurate at 99.02% and is also delivered fastest at 148.8 words a minute, so it buys an error every 15.8 seconds, while chat shows at 98.26% and 127.1 wpm buy one every 10.4. The gap between best and worst on the scorecard is 0.76 points; the gap between them in error points per hundred words is 0.98 against 1.74, which is nearly twice as many. That is the second half of the trick, and it is what makes a rate near 100% such a poor instrument: it has almost no room left to move, so a near-doubling of the error arrives as three-quarters of a point and reads as a rounding difference. Three more things the denominator is quietly doing. N counts the words that were subtitled, not the words that were spoken — about 112,000 of the sample’s 145,000 — and correct editing is not an error, so when 27.5% of chat show speech and 32.1% of entertainment speech is edited out, those words leave the numerator and the denominator together and the score has no opinion about them. That half of it belongs to the late denominator, and the two ride together here. Latency is not in N at all: subtitles ran an average of 5.6 seconds behind the speech against a 3-second guideline, 7 to 8 seconds through overlapping speech, with separate peaks of 10 to 21 seconds that the reviewers put down to technical faults. Nor is speed: 92% of programmes contained bursts above 200 words a minute, too fast for many subtitle users to read, and news averaged 35 such bursts per ten-minute clip. The regulator, to its credit, measured that last one and published the answer. Counting rapid subtitles as standard errors takes the average from 98.55% to 97.93% — below the threshold — and moves the share of programmes under the line from 20% to 50%, and the share above 99% from 29% to 5%. It is the same sample, the same model and the same axis; only the grain changed, twice. (The slide is our demonstration in the manner of an internal quality scorecard: the board, the grades, the status line, the exceptions and the read-out are invented, and every figure behind them is the regulator’s own. Nothing here is a criticism of the subtitlers, whose work the same reviewers called great under a hard constraint — the argument is with the number, not the work it describes.)

03The fix

Quote the rate at the grain of the thing that has to come out right, and say which grain you chose. There is nothing wrong with a per-word accuracy as an instrument — it is exactly what you want for comparing two subtitling systems, the same way a word error rate is what you want for comparing two recognisers — and the fix is not to throw it away but to publish beside it the figure the reader was trying to work out anyway: per word and per sentence, per step and per run, per component and per assembly. It is one line of arithmetic. Then plot the complement rather than the rate. Errors per hundred words runs 0.98 for news against 1.74 for chat shows on a zero-based axis, and the difference is plain; the same fact as 99.02% against 98.26% is two bars of identical height. A measure with no room left above it cannot show a change, and its complement can show every change it has — which is why a series creeping from 99.2% to 99.6% deserves to be drawn as an error rate that halved. Put a count next to the ratio, because a count already lives at a grain a person can hold: fifty errors in a ten-minute clip settles what 98.55% leaves open, and a reader who is told fifty will never mistake the slide for a pass. Publish the severity mix while you are at it, since these reviewers did and it is what turns a bracket of 19 to 77 into a number. Then be explicit about what the compounding assumes and which way it leans, since independence is rarely true — real errors cluster, a burst of them lands in one stretch of one programme, and clustering packs more errors into fewer sentences, so the naive calculation is an upper bound on how many sentences are touched rather than an estimate of it. Give the bound and name it as one; a stated bound is worth more than a false point estimate. Say what the denominator never saw, too, because a rate computed over what was delivered has no opinion about what was dropped: when a third of the speech is edited out before N is counted, that belongs in the caption beside the score. And when a threshold is being set, set it in the unit the consequence lands in. Ninety-eight per cent as a pass mark was chosen honestly and confirmed across four rounds of assessment, but as a promise to a viewer it is an error every nine seconds, and as a promise about a service it is seven days a year of not working. That is the sentence worth carrying away, because it outlives subtitles: the unit a thing is measured in and the unit it is lived in are two different choices, and a percentage says nothing at all about which one it was quoted in.