The head-to-head trial that was really a photo finish
Showing the misleading chart
A psychedelic-clinic briefing pits psilocybin against a daily antidepressant on depression scores — an 8.0-point drop to the SSRI’s 6.0 — and starts the axis at 5.0, so the winner’s bar towers three times over. Ground the axis at zero and put back the error bars the deck dropped, and the rout is a photo finish: the 2-point gap’s own 95% confidence interval runs from −5.0 to +0.9, straight through zero (p = 0.17), on 30 patients against 29.
01The claim
Psilocybin doesn’t just treat depression — it laps the drugs that do. In a head-to-head trial, two guided doses shed a full eight points off the symptom score in six weeks; the old daily antidepressant managed six, barely off the line. Three times the drop, a fraction of the pills, and the contest is over — ask your provider about the fast lane.
02The trick
Every number on the chart is the trial’s own, printed honestly beneath its bar — psilocybin’s mean 8.0-point drop on the QIDS-SR-16 depression scale, escitalopram’s 6.0 — and the deck still can’t tell the truth, because it pulls two tricks at once and the numbers do nothing to stop it. First the axis: a bar encodes value as length, so its baseline has to be zero, but this one starts at 5.0 and stops at 8.5, keeping only the top three-and-a-half points of an eight-point fall. Each bar is drawn as its value minus five, so psilocybin’s 3.0 towers over escitalopram’s 1.0 and the picture reads 3× — for a real gap of 8 versus 6, about 1.33×. Our tax-rate cliff does this with a 34% axis floor; this one hides the bottom five points of a depression score. Then the error bars vanish, and that is the deeper crime, because the gap the axis is inflating isn’t even real: with 30 patients on psilocybin and 29 on escitalopram, each arm’s mean change carries a ±1.0-point standard error, and the difference between them — 2.0 points — has a 95% confidence interval running from −5.0 to +0.9, straight through zero (p = 0.17). On its own registered primary outcome the trial found no significant difference at all, and its authors said plainly that no conclusion about which treatment is better can be drawn. A crisp bar looks equally confident resting on a firm finding or on noise; our Steam leaderboard pairs these same two tricks on a storefront, and the lesson travels: a two-bar “winner” with no interval is a headline, not a result. (This exhibit is our own demonstration in the house style of a psychedelic clinic’s outcomes briefing, drawn from the trial’s published figures rather than from any real clinic’s deck.)
03The fix
Ground the axis at zero and put the error bars back, and the blowout becomes a photo finish. On a zero-based scale the two bars stand close — 8.0 against 6.0 — and their 95% intervals, roughly ±2 points each, overlap across most of their length; the 2-point gap between them is exactly the kind of difference that a study of 59 people cannot separate from chance. Both treatments, it’s worth saying, moved the needle by a lot: each arm shed six to eight points on a 0–27 scale over six weeks, and both groups received the same intensive psychological support alongside the drug, so the honest headline is “two treatments, similar six-week improvement, difference too small to call,” not “one lapped the other.” The numbers the deck built its victory on — response of 70% versus 48%, remission of 57% versus 28% — really did favor psilocybin, but they were secondary outcomes the authors deliberately left uncorrected for multiple comparisons, precisely so no one would read them as proof of superiority; the two arms also started at different baseline severities (14.5 versus 16.4), which a raw change-score race quietly ignores. The tell is a comparison with a declared winner but no error bars and an axis that doesn’t begin at zero: ask how many people, and ask for the interval. A real effect survives being drawn on an honest axis with its uncertainty attached; a 2-point edge on 59 patients did not — which settles the chart, not the science. What the trial actually calls for is a bigger one, exactly as its own authors and the outside statisticians who reviewed it concluded.