My neighbour cancelled a birthday party on a Saturday in June because the forecast said thirty percent chance of rain, and then it did not rain, and he told me the forecast had been wrong. Three weeks later the forecast said thirty percent, he held the party, it rained on the cake, and he told me the forecast had been wrong.
He is not a foolish man. He runs a business and reads two newspapers. He simply has no procedure for judging a probabilistic statement, and neither, as far as I can tell, does anyone, and I have come to think that absence is the single largest obstacle to communicating anything scientific to the public — larger than innumeracy, larger than distrust, larger than the state of science education.
Start from what the statement means, because it is more interesting than it looks. When a forecaster says thirty percent, they are not describing Saturday. Saturday will either be rained on or it will not; there is no thirty percent version of Saturday. They are describing themselves. They are saying: among all the days on which my information looked like this, about three in ten produced rain. It is a claim about the reliability of a process, and the crucial consequence is that no single day can confirm or refute it. Not one. A day is not a sample of a rate.
What can test it is a tally. Collect every day the forecaster said thirty percent, count what fraction rained, and if it is close to thirty percent the forecaster is calibrated, which is the actual virtue on offer. Do the same at every other number. A perfectly calibrated forecaster is not one who is right; they are one whose confidence means what it says. And here is the part people find genuinely counterintuitive when I explain it: a calibrated forecaster must be wrong thirty percent of the time when they say thirty percent, or they were lying about the thirty.
Weather is the easy case, because it repeats. You get a fresh Saturday every week and the tally builds itself. Nearly everything else we now ask the public to reason about does not repeat, or does not repeat fast enough for anyone to keep score, and that is where this stops being a curiosity.
Take a hurricane cone, which our county emergency manager puts on a screen every August. The cone is a confidence region: historically the centre of the storm stays inside it about two thirds of the time. It is not the storm's width. It is not the area that gets damaged. People read it as both — I have watched a room of adults read it as both — and evacuate or fail to evacuate accordingly. Nobody in that room is stupid. The picture is answering a question about the forecaster's track record, and every viewer is asking a question about their house.
The objection I take most seriously is that this is the scientist's problem, not the public's, and the fix is on our side: stop publishing probabilities to people who cannot use them and publish a recommendation instead. Say evacuate or do not evacuate. Say hold the party or cancel it. Certainly say fewer numbers.
I have real sympathy for that and it fails on a specific point. The recommendation depends on the cost of being wrong in each direction, and those costs are not the forecaster's. They are the neighbour's. Thirty percent means one thing if the party is in a garden with a garage twenty feet away and another if two hundred people have flown in, and no forecaster knows which situation you are in. Handing over a probability rather than an instruction is not evasion; it is the recognition that the decision requires two inputs and the institution only has one of them. What went wrong is that we handed over the input without ever teaching the arithmetic that combines it, and then complained that people used it as an instruction.
So the missing thing is not better numbers. It is a public habit of keeping score over time — the thing sports fans do without effort and voters do not do at all. We have no cultural apparatus for it. There is no widely read scorecard telling you which forecasters, which agencies, which commentators were calibrated last year, so a confident wrong prediction and a hedged right one look identical in the record, and the incentive points exactly where you would expect.
My neighbour, to his enormous credit, has started a spreadsheet. He has been logging the morning forecast and the outcome since August. He told me last week, with some indignation, that at thirty percent it has rained eleven times out of thirty-four, and asked me whether that meant they were any good.
I said yes. He said it did not feel like it, and then he went and looked at the seventies, which is where the interesting failures usually are.

The hurricane cone example is the one I use and it never fully lands, because the picture is answering a question about the forecaster and the viewer is asking about their house. You put that better than I have managed.