The canary that deployed nothing and still won on five paths out of 24
Showing the misleading chart
A canary-analysis slide we drew from one round of real RIPE Atlas pings: a “new network stack” on half of 14,149 probes is faster to five root servers, every gap significant, m-root by 1.67 ms at p = 0.0018. The halves were drawn by a random number generator and nothing was deployed. The slide shows the five paths that cleared p < 0.05 out of 24 it tested, and 11 of 20 coin-toss splits of the same fleet find a winner of their own.
01The claim
The new network stack makes the fleet faster. A canary-analysis slide of the kind a platform team might put in a release review, which we drew from real measurements: stack 4.18 goes to a canary half of a 14,149-probe measurement fleet, the other half keeps the old stack, and in the same four-minute round each probe pings the DNS root servers three times, over IPv4 and, for the half of the fleet that has it, over IPv6. The canary half reaches five root servers faster than the baseline, every gap significant on a two-sided test. m-root, the biggest, is 1.67 ms faster, 28.94 against 30.61 ms (5.4%), at p = 0.0018. All five point the same way and none goes against the canary, on a zero-based axis with every median printed. Recommendation: promote it to the whole fleet.
02The trick
Nothing was deployed. We split the probes into two halves with a random number generator, drawn before any test was run, so “stack 4.18” is a label on a coin toss. Every round-trip time is real, and every p-value is correct for its own comparison. The mistake is in which comparisons reached the slide. The review tested 24 paths, twelve root servers over IPv4 and over IPv6, and drew the five that cleared p < 0.05. A test at that threshold calls a difference that isn’t there one time in twenty, so 24 tests on two identical halves expect about 1.2 hits, and better-than-even odds of at least one. The small print says “24 paths compared”, but the bars show only the winners. The five agreeing with each other is not extra evidence either. They are the same probes measured five times: by luck of the draw the canary half came out with slightly lower round-trip times to all twelve IPv4 roots (a-root by just 0.01 ms), so one coincidence produced five hits. That is what correlated metrics do: they don’t raise false alarms more often, they bunch them, so a split either finds nothing or finds several at once. Five or more hits turned up in 4.5% of 1,000 coin-toss splits of the same fleet, about seven times what 24 independent tests would give. On IPv6 the same split is a coin toss, canary faster to six roots and slower to six, none significant. This split, the first and only one we drew, happened to be a loud one.
03The fix
Run the same analysis where nothing changed. Draw nineteen more coin-toss splits of the same 14,149 probes, run the same 24 tests on each, and lay all 480 results out in one grid, the misses next to the hits. That gives 23 hits (4.8%, against the 24 chance expects), and 11 of the 20 splits flag at least one path, scattered across roots, address families and both directions. Across 1,000 splits the average is 1.19 hits, 55.8% find at least one, and each single test fires 4.9% of the time, exactly the rate it promises. A correction helps but doesn’t settle it. m-root’s 0.0018 even clears a Bonferroni bar of 0.05/24 = 0.0021, which 4.0% of the nothing-changed splits also cleared, because a correction caps how often a family of tests raises a false alarm and doesn’t stop it happening. What tests a stack is naming one primary metric before the rollout and reporting it first, counting every comparison shown, calibrating against an A/A run on the same fleet, and checking a surprising win on a fresh split of new hosts, where a real effect tends to turn up again, usually smaller, and a lucky draw tends not to.