myweatherrock

The Weather Journal / Explainers

Why Forecasts Go Wrong

By the rock. August 25, 2026. Grounded in National Weather Service definitions; suitable for classrooms and other curious humans.

The short version

Forecasts go wrong because the atmosphere amplifies small errors. No set of observations is perfect or complete, and the equations of weather compound every imperfection, doubling small differences until two nearly identical starting points produce two different weekends. That is chaos, demonstrated by Edward Lorenz in 1963, and it is why forecasters now run ensembles: the same model dozens of times with slightly wobbled starts, reading confidence from the spread. You can watch the decay in this site's own graded record: as of the receipts page's August 13 update, our wet-or-dry calls verified 92 percent of the time one hour out and 77 percent at 48 hours. Same rock, same math, longer lead.

The rerun that would not repeat

In 1961, an MIT meteorologist named Edward Lorenz was running a small weather model, twelve equations, on a Royal McBee LGP-30 computer. Wanting to re-examine one stretch, he restarted the run from the middle, typing in the numbers from a printout. The printout rounded to three decimal places: he entered 0.506 where the machine had been carrying 0.506127. A difference of about one part in ten thousand, far smaller than any real weather instrument can even measure.

The rerun tracked the original for a while, then drifted, then bore no resemblance to it at all. Lorenz had discovered, by accident and then confirmed with great care, that the equations of the atmosphere amplify tiny differences instead of forgetting them. He published the result in 1963 as Deterministic Nonperiodic Flow, in the Journal of the Atmospheric Sciences, and it became one of the founding papers of chaos theory. The practical sentence inside the mathematics: even a perfect model of the atmosphere, fed almost perfect observations, produces a forecast that degrades with time, because almost is a load-bearing word.

Error grows on compound interest

Every forecast starts from a snapshot of the atmosphere built from surface stations, balloons, aircraft, and satellites, and the snapshot has holes: instruments have tolerances, oceans are sparsely sampled, and the air between observations is estimated. Those small initial errors do not sit still. The atmosphere's own physics multiplies them, the errors infect neighboring regions, small-scale mistakes grow into large-scale ones, and past a week or two the compounding swamps the signal for day-to-day details. This is not a hardware problem or a laziness problem. Better models and better observations push the horizon outward, and famously have: a widely cited 2015 review in Nature called it the quiet revolution, forecast skill gaining roughly one day of usable lead time per decade. But the horizon moves; it does not go away. Chaos sets the terms.

today, nearly identical starts day 7, different weekends day 0
An ensemble, schematically: one model, many almost identical starting points. Where the lines hug each other, trust the forecast. Where they fan out, the honest answer is the fan.

Ensembles: wobble it and run it again

Since the starting snapshot cannot be perfect, modern forecasting stopped pretending it was. Forecast centers run the same model dozens of times, each run beginning from a slightly different, equally plausible version of right now. The bundle is called an ensemble, and it turns chaos from an enemy into an instrument. When the runs agree five days out, the atmosphere is in a forgiving mood and confidence is high. When they scatter, the spread itself is the forecast: it says the answer is genuinely not knowable yet, and says it honestly. Every chance of rain you have ever read is downstream of this machinery; that page covers what the percent does and does not promise.

Our own miss record, in public

Most weather sites tell you they are accurate. This one keeps a receipts page, where every forecast the rock publishes is written down, waited on, and graded against what actually happened, and the numbers print whether they flatter the rock or not. It is the primary source for this lesson, and as of its August 13, 2026 update, with 2,827 graded calls on the ledger, it shows error growth in the wild. Wet or dry, our calls verified 92 percent one hour ahead, 91 percent at six hours, 88 percent at 24, and 77 percent at 48. Temperature misses, meanwhile, barely budged: off by 2.4 degrees at one hour and 2.6 at 48. Lead time punishes the yes-or-no questions first.

Two smaller lessons hide in the same table. First: the 12-hour row reads 96 percent, better than the one-hour row, which should smell wrong, and mildly is. That row holds 175 calls; a sample that size wobbles, which is why the receipts page flags thin samples instead of celebrating them. Second: on the day-ahead rain question, the three professional models we grade did not tie. Over roughly 600 days each, one verified 71 percent, the others around 60. Different models, same atmosphere, different errors: that disagreement is exactly the spread an ensemble is built to measure.

Where forecasts actually break

Forecasts rarely fail everywhere at once; they fail by category. Temperature usually misses small, a degree or three, and mostly from local effects. Timing misses medium: the arrives, just four hours late, and your dry morning was real while your dry afternoon was not. Precipitation placement misses biggest, especially on scattered-storm days when the atmosphere itself has not decided which county wins. Knowing WHICH kind of miss to expect is most of the skill of using a forecast.

Grade one week yourself

Tonight, write down tomorrow's forecast for your spot: high temperature, sky, whether it rains, and when. Tomorrow night, grade it against what happened, in three columns: temperature, timing, precipitation. Repeat for seven days. One week is a thin sample, as the 96 percent row just taught you, but the pattern still shows: the temperature column will mostly hold, and the misses will pool in timing and precipitation. By Sunday you will have built a miniature receipts page of your own, and you will read every forecast afterward the way forecasters do: as odds, with a known weak spot.

Words in this lesson

Every marked word above is here in full, because a definition that only appears on hover is no use to somebody reading on a phone, printing the page, or listening to it.

Front
noun say it: FRUNT
The boundary where two air masses of different temperature and moisture meet. Air gets lifted along a front, which is why most of a map's clouds and rain queue up on one.

Definitions in this explainer follow the National Weather Service. The rock checked twice. The rock always checks twice.

Try it yourself

Slide the lead time. These are not invented numbers; they are this site's own graded wet-or-dry accuracy from the receipts page, as published in its August 13, 2026 update.

1 hour

Wet or dry, we were right: --

Did it stick? Five questions.

Take it to class

This explainer is free for classrooms: print it, copy it, staple it. The rock asks only that the weather stay accurate.

Download the lesson (PDF)  ·  Download the worksheet (PDF)

Questions people actually ask

If chaos limits forecasts, why have they gotten so much better?

Better observations, better models, and better use of ensembles keep shrinking the initial errors and squeezing more out of the physics. A 2015 review in Nature put the pace at roughly one extra day of usable lead time per decade. Chaos does not forbid improvement; it sets a horizon that improvement approaches expensively.

Is a forecast that said 20 percent chance of rain wrong when it rains?

No. A 20 percent day rains one time in five, and the forecast priced that in. What the percent means, and how to hold it accountable fairly, is the chance of rain lesson; how often OUR percents kept their promises is a table on the receipts page.

Why do different weather apps disagree about the same day?

They are usually reading different models, or blending them differently, and the models genuinely disagree; on our own day-ahead grading the three we track ran from 71 percent down to 59. When apps disagree, you are looking at ensemble spread with branding.

What is the hardest thing to forecast?

Among everyday weather: exactly where and when scattered storms fire. The setup is predictable; the specific county is often not, because storm-scale details grow from below what observations resolve. That is chaos operating at its fastest scale.

Does this site really grade its own forecasts?

Every one it publishes, against independent observations of what actually happened, with the misses printed alongside the hits. The receipts page updates as the ledgers grow, so its numbers will drift from the snapshot quoted in this lesson; the drift is the point.

Back to the journal.