The 2 mmHg look and the 57 mmHg look, drawn the same height
Showing the misleading chart
A theatre quality-review slide we drew takes one dashboard tile — the last ten minutes of the arterial line’s mean pressure — and screenshots it four times across a single anaesthetic, plotting VitalDB’s readings at the recorder’s own two-second interval: 300 to a screenshot, nothing smoothed, binned or dropped. All four come back as mountain ranges, so the desk reports a case that never settled. Three of them move by 6 mmHg, then 2, then 4. The 45–55 look moves by 57. The tile’s axis is fitted to whatever is in it, so a millimetre of mercury is drawn 28.5 times taller in the third look than in the first, and the 8.53 mmHg the pressure fell between the second look and the third is on neither of them.
01The claim
Unstable every time we looked. This is one dashboard tile showing the last ten minutes of arterial pressure, screenshotted four times across the 4 h 23 min the recorder ran. Every reading is VitalDB’s own — mean pressure off a Solar 8000 monitor, one every two seconds, 300 to a screenshot, nothing smoothed or dropped. The tile’s vertical axis is auto-ranged to what is in it, which is what a panel does unless it is told otherwise. Four looks, four times the same shape: a mountain range that never settles. Read-out for the desk: the trace does not settle. Whatever moment you catch the tile at, the pressure is climbing or falling, and the pattern holds from the start of surgery to the end of it — so this is not one bad stretch, it is how the whole case ran. Recommendation: narrow the alarm band on mean pressure, add a stability index to the theatre dashboard, and put this case on the list for haemodynamic review.
02The trick
Nothing on the slide is drawn wrong. It is one tile, one series and one unit throughout. No axis has been truncated by hand, nothing is indexed, rebased, logged or smoothed, no reading has been dropped, every screenshot is the same size, the ticks are round numbers, and each panel prints in grey type the range its axis was fitted to. The drawing is not doing anything. What the four screenshots have in common is that each was handed a scale computed from its own ten minutes — and an axis fitted to a window is an axis that window cannot be measured against. Read the ranges the slide itself prints and the four stop being four of a kind. The 45–55 look spans 69 to 126 mmHg: 57 mmHg of movement. The 85–95 look spans 81 to 87: six. The 135–145 look spans 74 to 76: two. The 205–215 look spans 72 to 76: four. Three of the four are as close to a flat line as a real measurement gets, and one is a genuine event. All four were drawn edge to edge, because that is what fitting an axis to its contents means. The consequence is arithmetic. Every chart has an exchange rate between a unit of data and a millimetre of screen, and on a fixed range that rate is a constant the reader can learn once. Here it is recomputed every time the tile is looked at: the whole height is worth 57 mmHg in the first look and 2 mmHg in the third, so one millimetre of mercury is drawn 28.5 times taller in the third than in the first. That ratio is the entire difference between the pictures, it appears nowhere on the chart, and it changes the moment the window moves. It is worth being precise about how this differs from the axis crimes it resembles, because the difference is what makes it hard to catch. A truncated axis chops the bottom off the scale and announces itself by starting at 94.6; the number is printed and a reader can mentally restore it. A broken axis takes a slice out of the middle and marks the theft with a zigzag. A clipped axis pins a maximum somebody typed, and the tell is a row of flat tops. All three are decisions, made once, by a person, and all three hold still — which is what lets you notice them, and what lets two charts drawn the same way be compared. The nearest relative is not any of those but the free-scale panel grid, where several different series share a layout and each panel is given a ruler of its own; that is the same rule applied across space. This is the same rule applied across time, which is the half that has no grid to give it away: one series, one tile, one place on the dashboard, and a different ruler every time anybody looks at it. Two further consequences follow, and both are doing work here. The first is that the axis is recentred as well as resized, so an auto-ranged panel discards the level along with the amplitude. Mean pressure averaged 83.06 mmHg across the second look and 74.52 across the third — a fall of 8.53 mmHg, which is more than four times the entire vertical extent of the third panel — and neither can show it, because each is drawn around its own middle. Someone flicking between two screenshots sees two equally jagged traces and has no way to learn that the patient spent the second half of the case at a materially lower pressure than the first. On a review dashboard, that step is far likelier to be the finding than the wiggle either panel is shouting about. The second is that an auto-range is set by exactly two numbers: the largest and the smallest reading in view. That makes it maximally sensitive to precisely the readings least worth trusting. Ten minutes of this record, from minute 190 to minute 200, contain 291 readings, 278 of which sit between 72 and 87 mmHg. The other 13 are a line flush at 197.5 minutes — 24 seconds of it, peaking at 332 mmHg, which is not a blood pressure but a pressurised flush of the tubing. Auto-ranged, those 13 readings take the axis to 72–332 and leave the window’s other 278 occupying 5.8% of the frame: a flat line along the bottom and one spike, from the same rule that turned a 2 mmHg window into a mountain range. Same default, opposite failure, and nothing in the picture distinguishes the two cases. The same mechanism makes the history unstable, which is the part with no analogue anywhere else in the taxonomy. Because the range follows the window, one fresh extreme rescales everything already drawn: scroll a live panel forward and the past changes shape behind you, so the same minute is a different height in every screenshot and two pictures of one series are not comparable at all. This is where a dashboard is worse than a printed chart rather than better. A printed chart is wrong in one way, permanently, and a reader who learns the trick can correct for it; a panel that refits itself is wrong in a new way every time it is looked at, and the change a reader remembers seeing between Monday and Friday may be nothing but the axis moving. The tools are not hiding any of this — Grafana’s own manual says that “by default, Grafana sets the range for the y-axis automatically based on the dataset”, and describes the cure by naming the disease: soft min and soft max, it says, “can prevent small variations in the data from being magnified when it’s mostly flat”. The setting exists because the failure is well known to the people who built the thing. It is simply not the default, and a default is what almost every chart is drawn with. (Each screenshot here is fitted exactly to its own minimum and maximum; a real panel pads a little or rounds the bounds outward, which softens the effect without changing it. VitalDB publishes the readings; the tile, the windows, the arithmetic and both drawings are ours, and the desk, the read-out and the recommendation on the slide are invented.)
03The fix
Choose the range once, and choose it from the question rather than from the data. The test is whether you could state the bounds before seeing the series: a mean arterial pressure is read in a band a clinician can name, a latency budget has a target, a share has a hundred in it. The redraw takes 40 to 130 mmHg, says so in the subtitle, and then does nothing cleverer than hold it — on the full record, on the four looks, and on the flush panel alike. On that one ruler the case reads immediately. The pressure dips into the fifties and sixties around minute 20, settles into the seventies by minute 26, sits in a band between roughly 70 and 90 for three hours, and contains two excursions rather than the four the slide advertised or the none its numbers implied: 69 to 126 and back to 70, beginning three minutes after the surgeon started at minute 44.7, and 76 to 131 over the closing ten minutes as the case ends. The four screenshots caught one of the two, and reported all four as events. Drawn side by side on the same range, the looks print their own movement above them — 57 mmHg, then 6, then 2, then 4 — and three of them are visibly straight lines sitting at slightly different heights, which is where the 8.53 mmHg step between the second and the third becomes a thing you can see rather than a thing you have to be told. Nothing is dropped to achieve any of that: 389 of the 7,856 readings fall outside the chosen range and every one is ticked at the frame and counted in the caption, including the 354 at the start where the arterial line is not yet transducing. Hold the range across panels, across window widths and across weeks, for the reason a small multiple holds one axis: the whole value of a comparison is that the ruler did not move between the two things being compared. Where the range genuinely has to follow the data — and often it must, because nobody can pin an axis on a metric they have never seen — make it follow slowly and visibly. Soft limits are the first line: they widen for a real excursion but will not zoom into a flat stretch, which is the single setting that would have stopped this slide. Compute the bounds from a robust range rather than from the minimum and the maximum, so one glitch cannot take the scale with it, and tick the excluded points at the frame with a count beside them rather than dropping them. Put something in the plot that does not move — a target line, a normal band, a reference series — so there is a fixed thing for the wiggle to be measured against; the redraw uses a plain 65 mmHg rule for that, and for nothing else. And print the span as a number beside the panel: “range 2 mmHg” is one line, it survives being screenshotted into a slide deck, and it is the line no charting default will ever write for you. The habit that generalises is smaller than any of that: before reading a shape off a chart, read what the height of the frame is worth. On an auto-ranged axis the shape is guaranteed to be there whatever the data does, which is exactly what makes it useless as evidence — and the places this matters most are the places with no axis at all, where the guarantee is invisible: sparklines, dashboard tiles, thumbnails, phone widgets, the little graph on a fitness app, the sparkline in a spreadsheet cell. A picture that redraws itself every time it is looked at was fitted to the looking, not to the question.