The 16.9% lift on day three that was gone by day thirty
Showing the misleading chart
An experiment dashboard we drew runs a two-arm test on every street-defect request New York’s 311 line took in June 2025 — 7,243 of them, with the city’s own logging and closing times — and asks whether arm B closes more of them within 24 hours. Assignment is one published random bit per request; nothing is smoothed, dropped or rescaled, and both axes start at zero. On the third morning the dashboard turns green: 51.71% against 44.25%, a lift of 16.9%, p = 0.035, so the desk stops the test and ships the winner. There was no variant. Both arms are the same requests, split by a coin flip, so the true difference is zero — and run to the planned end, the test reads 44.43% against 44.29% at p = 0.90. The only thing the desk chose was when to look, and it looked every morning: split the same requests 10,000 more times and 22.85% of the tests produce a winner somewhere in the month, against 5.25% for the same test read once, at the end.
01The claim
Arm B closes 16.9% more of them. Call it. Two arms, one metric: the share of street-defect requests closed within 24 hours. Every request and every closing time is the city’s own record — 7,243 of them in the month this test was planned to run — and assignment is a published random bit per request, so the arms are a coin flip apart. Nothing is smoothed, binned or dropped, both axes start at zero, and the interval is the ordinary two-proportion one. The desk reads the dashboard every morning. On the third morning it turned green: arm B at 51.71%, arm A at 44.25%, a gap of 7.46 points on 801 requests, an interval of +0.56 to +14.36 points that clears zero, and a two-sided p of 0.035. B is ahead and the difference is significant at the 5% level; holding the test open now only spends requests on the worse arm. Recommendation: stop the test, ship arm B’s routing everywhere, and book the 16.9% as the expected gain in the quarterly plan.
02The trick
Nothing on the dashboard is drawn wrong, and that is not a figure of speech: one metric, one definition, one population, zero-based axes, no request dropped or reweighted, and a p-value that is correctly computed for the look it was computed at. Audit the picture as hard as you like and it comes back clean, because the choice that produced the winner is not in the picture. The desk chose when to look — every morning — and it chose to stop at the first look that pleased it. Here is why that is fatal. The difference between two arms of a fair test does not sit quietly at zero while the sample grows; it wanders, widely when the counts are small and less widely later, and on any given morning it has a fresh chance to be extreme. A test read once, at a sample size fixed before the test opened, raises a false winner 5% of the time when there is nothing there, and that 5% is the number everybody has agreed to live with. A test that may be read on any of 28 mornings and is answered the first time it says yes is a different instrument, because a rule that stops at the first crossing takes its reading from the upper tail of that wander. Armitage, McPherson and Rowe printed the arithmetic in 1969: repeated significance tests on accumulating data push the probability of a significant result above the nominal level, and it keeps climbing with the number of looks. This exhibit is the version where the truth is knowable, because there was no variant. Arm A and arm B are the same 7,243 street-defect requests, split by one bit per request from a published random beacon — an A/A test, the thing experiment platforms run to check themselves — so the real difference between the arms is zero by construction and every gap on the first slide is noise being read as a finding. The day-3 gap itself is perfectly real: 212 of 410 against 173 of 391, 7.46 points, and the exact test agrees with the dashboard’s (Fisher p = 0.040). What p = 0.035 says is that a gap this wide or wider turns up 3.5% of the time when the two arms are the same — at one look fixed in advance. This dashboard offered 28 of them, and the desk took the third. Run the same arithmetic 10,000 times — fresh random splits of the same requests, peeked at daily — and 22.85% of the tests cross p < 0.05 somewhere in the month, against 5.25% for the same test read once at the planned end, which is the 5% the test promised and the sign that nothing else here is broken. The guard rails people reach for help and do not save you: waiting until each arm holds 1,000 requests takes 22.85% down to 17.13%, and looking weekly rather than daily takes it to 12.88%. Neither is near 5%, because the cost is set by how many chances the test is given, not by how sensible each chance feels. And the early stop does a second kind of damage, which outlives the first: stopping on a threshold selects for a gap that is wide, so the number you keep is biased high. Across the 2,285 tests that stopped, the median gap on the morning they stopped was 10.8% in one direction or the other, and the widest was 30.4%, all of it from arms that were identical — so even where a real effect exists, the number that gets written into the plan, and that the next test gets powered off, sits above it. The near relative in this guide is the cherry-picked window, and the difference is the part worth holding on to. A window is chosen inside data that already exists, so a reader can ask to see the rest of it. This chooses when the data stops arriving: on day 3 the other 6,442 requests of the planned month had not been raised yet, and stopping meant they were never counted — in a live test they never exist at all, and they exist here only because the month had already happened when we drew it, so there is no truncated axis, no missing series and no line to look up. Two things the metric is not, neither of which touches a comparison of two arms drawn from the same pile: a 311 closing time is an administrative status change — repaired, referred, duplicated, or found to need no action — rather than a mended street; and each morning’s read uses every request’s eventual closing time, so the dashboard is drawn as though outcomes were known at once, which a live one could not be. (The slide is our demonstration in the manner of an experiment dashboard — the test, the desk, the read-out and the recommendation are invented, and every figure behind them is real. The beacon gave 35 blocks of bits long enough to split this month; 5 of them crossed p < 0.05 somewhere in the daily peeking, and the one drawn here is the second block — pulses 1921610 to 1921624, the first that crossed. That is the same hunt the page is about, run in the open: the desk was not unlucky, it was sampling from the same 22.85%. The dataset was retrieved on 19 September 2026; the arithmetic of repeated looks is Armitage, McPherson and Rowe, JRSS A 132(2):235–244, 1969, and the continuously readable alternative is Johari, Pekelis and Walsh, arXiv:1512.04922.)
03The fix
Run the test to the horizon it was given and read it once. Drawn that way the month is dull, which is the correct answer: after the third day the two lines close on each other and stay closed, ending at 44.43% for arm A and 44.29% for arm B — a gap of 0.14 points the wrong way, a lift of −0.3%, p = 0.90. The p-value panel is the tell that generalises. Twenty-eight mornings, one dip: 0.035 on day 3, 0.067 on day 4, and never under 0.05 again, finishing at 0.903. A single dip is what a fair test looks like when somebody watches it long enough, and a rule that stops at the first dip will find one about a quarter of the time. So fix the sample size before the test opens. It is not a matter of taste — it falls out of the smallest difference worth acting on, the baseline rate and the power you want — and writing it into the brief costs one line and settles the whole problem. Where a test genuinely has to be watched, because a harmful arm should be stopped or the business cannot wait a month to learn it broke the checkout, the watching is a solved problem with a literature and a name. A group-sequential design with a spending boundary — Pocock, O’Brien–Fleming, Lan–DeMets — fixes how many looks there will be and raises the bar at each one, so the error rate across the whole test is still the one you promised. Always-valid sequential p-values and confidence intervals, now standard in commercial experiment platforms, are built to be read continuously and to stay honest whenever you choose to stop. Both cost something: an early stop needs a wider gap to earn it, and the maths is a library call rather than nothing. That cost is the price of the looking, which was always being paid — just not declared. Then report the estimate knowing what the stopping rule did to it. A number that crossed a threshold was selected for being on the high side, so say so beside it, and do not power the next test off it. Print the stopping rule the way you print a unit: how many looks were taken, when they were taken, and what would have ended the test. A reader who knows the rule can discount the result, and a reader who does not cannot. Keep the log, too — every test started, with its horizon and its fate — because the tests that never crossed are the ones that make the crossers legible, which is the file drawer’s discipline applied inside a single experiment rather than across a literature. One last thing worth saying plainly, since it is the sentence that survives all the arithmetic: a p-value is a statement about a procedure, not about a gap. Change the procedure — by looking again, by extending a week “to be sure”, by stopping the moment the dashboard turns green — and the number on the screen keeps its old name and quietly stops meaning it.