#4788: How Weather Prediction Became a Science

From hand calculations to chaos theory to machine learning — the full arc of forecasting.

Featuring
Listen
0:00
0:00
Episode Details
Episode ID
MWP-4967
Published
Duration
28:02
Audio
Direct link
Pipeline
V5
TTS Engine
chatterbox-regular
Script Writing Agent
deepseek-v4-pro

AI-Generated Content: This podcast is created using AI personas. Please verify any important information independently.

Weather prediction began as a pure physics problem. In 1904, Vilhelm Bjerknes laid out the equations governing the atmosphere — seven partial differential equations that describe how air moves, heats, and holds moisture. He knew they couldn't be solved analytically; you'd have to compute them numerically, step by step, across a grid covering the globe. The means didn't exist yet.

Lewis Fry Richardson tried anyway. During and after World War I, he spent months computing a single six-hour forecast by hand. His result was off by two orders of magnitude — a predicted pressure change of 145 millibars when the real change was near zero. The numerical method was unstable, amplifying tiny errors in the initial wind field. But Richardson's vision was extraordinary: a forecast factory with 64,000 human computers, each calculating a small patch of the globe, coordinated by a conductor with coloured lights. He anticipated parallel computing decades before it existed.

The first working numerical forecast ran on ENIAC in 1950, barely breaking even — 24 hours of computation for a 24-hour forecast. Operational centres followed: the US Weather Bureau in 1954, ECMWF in 1975. Models grew from single-layer barotropic approximations to full global spectral models with dozens of vertical levels. Then Edward Lorenz discovered chaos in 1961 — a deterministic system with no randomness could still be unpredictable beyond two weeks because tiny differences in initial conditions grow exponentially. Ensemble forecasting emerged as the honest response, quantifying uncertainty instead of pretending to know the one true forecast.

The economic stakes are enormous. Airlines route around forecast winds to save millions in fuel. Container ships avoid storms to shave days off crossings. Farmers time planting, irrigation, and harvest to precipitation and temperature forecasts. Grid operators balance renewables against demand using wind and solar forecasts. And now machine learning is entering the field — learned global forecasters trained on reanalysis data from the physics models they may eventually replace. But weather and climate remain fundamentally different: weather is an initial-value problem, climate is a boundary-value problem. Different mathematics, different timescales, different tools.

Downloads

Episode Audio

Download the full episode as an MP3 file

Download MP3
Transcript (TXT)

Plain text transcript file

Transcript (PDF)

Formatted PDF with styling

#4788: How Weather Prediction Became a Science

Corn
Daniel's been thinking about weather prediction — not the app on your phone, but the whole intellectual arc of trying to calculate what the atmosphere will do next. He wants the history, from the moment someone first thought of weather as a physics problem you could solve with equations, through Richardson's mad attempt to compute a forecast by hand, the first computer models after the war, the big operational centres, and the discovery of chaos that put a hard ceiling on the whole enterprise. Then he wants the uses — not in the abstract, but where a forecast actually turns into a decision with money or lives on the line. Aviation, shipping, flood warning, military planning, the works.
Herman
That's a lot of ground.
Corn
And then he wants the AI story. What machine learning is doing in this space, whether it's replacing physics models or just bolted onto them, and a sceptical read on the new learned global forecasters — what they're good at, what they can't do, and the awkward fact that they're all trained on reanalysis data from the physics models they're supposedly going to replace. Also, keep weather prediction and climate projection separate. They're different problems on different timescales, people mix them up constantly, and we shouldn't.
Herman
That last one. Yes. Weather is an initial-value problem — you take the state of the atmosphere right now and run it forward. Climate is a boundary-value problem — you change the composition of the atmosphere and ask what the new equilibrium looks like. Completely different mathematical beasts. People use the same word, model, for both and assume it's the same thing. It's not.
Corn
So where does the idea of calculating the weather even start?
Herman
Vilhelm Bjerknes, nineteen-oh-four. He was a Norwegian physicist and he published a paper that basically said: the atmosphere is a fluid obeying the laws of physics, so if you know its current state and you have the equations that govern it, you can compute its future state. He laid out the problem as a set of seven partial differential equations — conservation of momentum, conservation of mass, the ideal gas law, the first law of thermodynamics, and the continuity equation for water vapour. And he said, look, we know the equations, we just need enough observations and enough computing power.
Corn
Which in nineteen-oh-four was... people with pencils.
Herman
People with pencils, exactly. Bjerknes knew the equations were unsolvable analytically — you can't just integrate them and get a nice formula. You'd have to do it numerically, step by step, across a grid covering the whole atmosphere. He understood the vision perfectly. He just didn't have the means.
Corn
Enter Lewis Fry Richardson.
Herman
Richardson was... I love this man. He was a Quaker, a pacifist, he drove an ambulance in the First World War, and in his spare time between the trenches he was working on the first numerical weather forecast. He took Bjerknes's equations and actually tried to compute a forecast by hand.
Corn
By hand.
Herman
He divided Europe into a grid, collected observations from a single day — May twentieth, nineteen-ten — and spent months doing arithmetic. The forecast was for a six-hour period over central Europe, and it predicted a pressure change of one hundred and forty-five millibars in six hours.
Corn
Which is...
Herman
Physically impossible. The real change was essentially zero. His calculation was off by two orders of magnitude. Richardson knew it was wrong, and he spent years figuring out why. The core problem was that his numerical method was unstable — tiny errors in the initial wind field would amplify catastrophically with each time step. He wrote a book about it in nineteen twenty-two called Weather Prediction by Numerical Process, and in it he described this extraordinary vision of a forecast factory.
Corn
The forecast factory.
Herman
A giant circular hall, like a theatre, with a map of the globe painted on the walls. In the gallery, thousands of human computers — he imagined sixty-four thousand of them — each one calculating the equations for a small patch of the globe. A conductor at the centre, shining coloured lights to coordinate the calculations, sending runners with numbers between the sections. He imagined computing the weather faster than the weather itself happens, so you'd actually get a forecast before the weather arrived.
Corn
Sixty-four thousand people with slide rules, coordinated by a man with a spotlight. It's the most beautiful thing that never worked.
Herman
It's completely wonderful. And he was right about almost everything except the instability problem. He even anticipated parallel computing — the idea that you'd divide the globe into subdomains and compute them simultaneously. That's exactly how modern weather models work on supercomputers. He just didn't have the mathematical tools to handle the error growth. Those came later.
Corn
So the war ends, and suddenly there are machines that can do arithmetic without human computers.
Herman
ENIAC. Nineteen-fifty. John von Neumann, who'd been working on the Manhattan Project, saw weather prediction as the perfect test case for electronic computers. He assembled a team at Princeton, and in March nineteen-fifty they ran the first successful numerical weather forecast on ENIAC. It took about twenty-four hours of computation to produce a twenty-four-hour forecast.
Corn
Barely breaking even.
Herman
Barely. But it worked. The forecast wasn't great by modern standards, but it was recognisably a weather forecast — it captured the large-scale pressure patterns. That was the proof of concept. After that, things accelerated fast. The US Weather Bureau set up a numerical prediction unit in nineteen fifty-four. By the late fifties, operational numerical weather prediction was running daily.
Corn
And the big centres — ECMWF, NOAA — when do they arrive?
Herman
NOAA's National Meteorological Center — which later became NCEP — was established in the fifties and really scaled up through the sixties and seventies. ECMWF, the European Centre for Medium-Range Weather Forecasts, was founded in nineteen seventy-five. It was a European response to the fact that the Americans were pulling ahead. The idea was: pool resources, build one world-class centre with a single massive computer and a single model, and make the best medium-range forecasts on the planet. It worked. ECMWF is still probably the premier global forecasting centre.
Corn
And the models themselves — you mentioned barotropic toy grids. What was the trajectory?
Herman
The earliest models were barotropic — they treated the atmosphere as a single layer with no vertical structure. That's fine for capturing large-scale waves in the jet stream, but it can't produce clouds, can't produce precipitation, can't handle fronts. Through the sixties and seventies, models added vertical layers, added moisture, added radiation, added parameterizations for things like turbulence and convection that happen at scales smaller than the grid cells. By the eighties you had full global spectral models with dozens of vertical levels, and by the two-thousands you had models with grids fine enough to start resolving individual storms.
Corn
And at some point in this story, Edward Lorenz walks in and ruins everything.
Herman
Nineteen sixty-one. Lorenz was running a simple weather model on a small computer at MIT — a toy model with just twelve variables, nothing like a real forecast. He wanted to re-run a section of a previous simulation, so he typed in the numbers from a printout. But the printout had rounded to three decimal places, while the computer was using six. The difference was one part in a thousand. And the re-run diverged completely.
Corn
The butterfly.
Herman
The butterfly was actually a later talk he gave — Predictability: Does the Flap of a Butterfly's Wings in Brazil Set Off a Tornado in Texas? But the discovery was that moment in sixty-one. A deterministic system, governed by exact equations, with no randomness anywhere, can produce behaviour that is unpredictable beyond a certain horizon because tiny differences in the initial state grow exponentially. That's the predictability horizon. For weather, it's about two weeks. No matter how good your model is, no matter how many observations you have, the errors in your initial conditions will eventually swamp the forecast.
Corn
And that's not a computing problem. That's a property of the atmosphere.
Herman
It's a property of the equations. It's built into the physics. You can't engineer your way around it. What you can do is run the model many times with slightly different initial conditions and see how the forecasts diverge. That's ensemble forecasting.
Corn
Which is why when you look at a hurricane track forecast, you see a cone, not a line.
Herman
The cone is the spread of the ensemble. A tight cone means the members agree — high confidence. A wide cone means the atmosphere is in a chaotic regime where small differences matter a lot. Ensemble forecasting started operationally in the nineties at ECMWF and NCEP, and it's now standard practice everywhere. It's the honest response to chaos — instead of pretending you know the one true forecast, you quantify your uncertainty.
Corn
All right. So we've got the physics, the computing, the chaos, the ensembles. Let's talk about where this actually touches the world. You mentioned aviation — what does a forecast become a decision?
Herman
Aviation is the cleanest example. Every commercial flight files a flight plan that depends on forecast winds aloft. The jet stream can add or subtract a hundred knots of ground speed. If you get the winds wrong, you burn more fuel than planned, and fuel is the single largest variable cost for an airline. Worse, you might not have enough fuel to reach your destination with legal reserves. Airlines have entire dispatch teams whose job is to integrate weather forecasts into routing decisions, and they're making choices with tens of thousands of dollars of fuel on the line, per flight.
Corn
And shipping?
Herman
Container ships route around storms. A North Atlantic storm can add two days to a crossing and burn hundreds of thousands of dollars in extra fuel. The optimal route is a calculation that balances forecast winds, waves, and currents against the cost of deviation. The shipping companies subscribe to specialised forecast services — not the public weather forecast, but bespoke marine forecasts optimised for their specific routes and vessel characteristics.
Corn
Agriculture — that's the one I think of as the oldest use case. Farmers have been trying to read the weather since there were farmers.
Herman
And now it's incredibly quantitative. Planting decisions depend on soil temperature and moisture forecasts. Irrigation scheduling depends on evapotranspiration forecasts. Harvest timing depends on precipitation forecasts — you don't want to cut hay if rain is coming in three days, because hay needs to dry in the field. And then there's the pesticide and fertiliser application: you need wind forecasts to avoid drift, and you need rain forecasts because a downpour right after application washes everything into the waterways. The economic value of a good seasonal forecast to agriculture is in the billions globally.
Corn
Energy grid — renewables scheduling, you said.
Herman
Wind and solar are the obvious ones. A wind farm operator needs to know what the wind speed will be at turbine hub height, typically eighty to a hundred metres up, for the next twenty-four to forty-eight hours, to bid into the electricity market. Solar is somewhat easier because cloud cover is the main variable and the sun's position is deterministic. But the grid operator has the harder problem: they need to balance supply and demand in real time, and renewables introduce variability on the supply side that used to be completely controllable. A bad wind forecast can mean firing up a gas peaker plant at the last minute, which is expensive and emits carbon. A good forecast saves money and emissions simultaneously.
Corn
Insurance and catastrophe modelling — that's where the money gets really large.
Herman
Catastrophe modelling is a whole industry. The big firms — AIR Worldwide, RMS, now Moody's — build models that simulate tens of thousands of synthetic hurricane tracks, or earthquake scenarios, or flood events, and estimate the probability distribution of losses for a given portfolio of properties. Weather forecasts feed into this in real time: when a hurricane is actually approaching, the models shift from climatological probabilities to conditional forecasts based on the current track and intensity predictions. That's what determines whether an insurer issues a binding restriction — a notice that no new policies can be written in certain ZIP codes until the storm passes.
Corn
Flood warning is the one where the forecast directly saves lives.
Herman
Flood forecasting is a chain. The meteorological forecast tells you how much rain will fall where. That feeds into a hydrological model that tells you how that rain will translate into river discharge. That feeds into a hydraulic model that tells you which areas will flood and to what depth. Each link in the chain has its own errors, and the errors compound. The European Flood Awareness System, EFAS, runs ensemble forecasts out to ten days and issues alerts to national authorities. The decision point is: do you evacuate? Evacuating a town costs millions and disrupts thousands of lives. Not evacuating when you should have costs lives. That's a decision made on the basis of a probabilistic forecast, and it's about as high-stakes as civilian forecasting gets.
Corn
And military planning — that's the one Daniel flagged that most people don't think about.
Herman
The military runs its own weather models. The US Air Force has the 557th Weather Wing, which operates a global forecast model separate from NOAA's. They care about things civilian forecasters don't: visibility for helicopter pilots, sea state for amphibious operations, upper-level winds for paradrops, atmospheric refractivity for radar and communications. During the Gulf War, the forecasts for wind direction mattered because of the chemical weapons threat — if the wind shifted, a Scud attack with chemical warheads could drift over your own troops. The military also has a long history of weather modification research — cloud seeding over the Ho Chi Minh Trail in Vietnam, for instance — but that's a different and much darker story.
Corn
All right. That's the physics-based world. Now the part where everything gets disrupted. How did machine learning get into weather forecasting?
Herman
It crept in through the side doors first. The core of a weather model is the dynamical core — the fluid dynamics solver that moves the atmosphere forward in time. That's pure physics, and for decades it was untouchable. But around the dynamical core there are dozens of parameterizations — sub-grid processes that can't be resolved directly, like cloud formation, turbulence, convection, radiative transfer. Those parameterizations are semi-empirical. They're tuned by hand. And about ten years ago, people started asking: what if we replaced some of these with learned models trained on high-resolution simulations or observations?
Corn
So initially, ML wasn't replacing the model. It was improving the bits the model couldn't do well.
Herman
Post-processing too — taking the raw model output and correcting its systematic biases using a learned mapping from past forecasts to past observations. That's been standard practice for years and nobody calls it AI, they just call it model output statistics. Data assimilation — the process of blending observations with a short-term forecast to produce the initial conditions for the next forecast — has also started incorporating machine learning, though that's harder because assimilation is a Bayesian inference problem with very strong physical constraints.
Corn
And then the emulators arrived.
Herman
The emulators. The idea here is: you have a physics-based model that's very expensive to run. You run it many times, generate a training dataset of inputs and outputs, and then train a neural network to approximate the mapping. The neural network runs thousands of times faster. This has been done successfully for specific tasks — emulating the radiation scheme, for instance, which is one of the most expensive parts of a climate model. But it's still a copy of the physics model. It can't outperform it.
Corn
And then the big shift — the learned global forecast models.
Herman
This is the thing everyone's talking about. Starting around twenty twenty-two, a series of papers showed that you could train a deep neural network directly on reanalysis data — ERA5 from ECMWF is the standard — and produce a global forecast that is competitive with the operational physics-based models. The big names: GraphCast from Google DeepMind, FourCastNet from NVIDIA, Pangu-Weather from Huawei, and now AIFS from ECMWF itself.
Corn
And reanalysis data is...
Herman
It's the output of a physics-based model that has been run retrospectively, assimilating all available historical observations to produce the best possible estimate of what the atmosphere was doing at every point in time. ERA5 goes back to nineteen-forty. It's the gold standard. And it's produced by ECMWF's integrated forecasting system — the very physics model these AI systems are being compared against.
Corn
So they're trained on the output of the thing they're supposed to replace.
Herman
Yes. That's the awkward fact Daniel mentioned. These models learn to imitate the physics model. If the physics model has systematic biases — and it does — the AI model will inherit them. It might even amplify them. And if you ask the AI model to forecast something that's outside the distribution of its training data — an unprecedented heatwave, say, or a storm of an intensity that never appeared in ERA5 — there's no guarantee it will behave physically.
Corn
But they're fast.
Herman
They're absurdly fast. GraphCast produces a ten-day global forecast in under a minute on a single TPU. The ECMWF operational model takes hours on a supercomputer with thousands of nodes. The energy savings alone are enormous. And on standard skill scores — root mean square error of five-hundred-hectopascal geopotential height, that sort of thing — GraphCast matches or beats the operational model on most variables out to about ten days.
Corn
So what's the sceptical read?
Herman
Several things. First, the skill scores are averages over a test set. They don't tell you about performance on the most extreme events, which are the ones you actually care about. Second, these models are deterministic — they produce a single forecast. You can make an ensemble by adding noise to the initial conditions or by training multiple models, but the ensemble spread may not be well-calibrated. Third, they don't conserve mass or energy. A physics model enforces conservation laws by construction. A neural network doesn't — it learns to approximate them from data, and it can drift. Over a ten-day forecast the drift is small, but if you tried to run one of these models out to a season or a year, it would probably produce nonsense.
Corn
And that's where the weather versus climate distinction bites hardest.
Herman
Climate models need to conserve energy to within a fraction of a watt per square metre over centuries. If they drift, the whole simulation becomes unphysical. AI weather models don't have that constraint, and it's not clear they can acquire it from data alone. For weather forecasting out to ten or fifteen days, the speed and skill of the learned models is impressive and they will almost certainly become operational. For climate projection — what happens to the monsoon in a world with six hundred parts per million of carbon dioxide — you still need a physics-based model that respects conservation laws. The two problems are not interchangeable.
Corn
And the other thing — these models predict what the atmosphere will do, but they don't tell you why.
Herman
That's the interpretability problem. A physics model has a diabatic heating rate you can inspect. You can trace a forecast feature back to a specific physical process. With a neural network, you get a number. You can do attribution studies — ablate an input, see how the forecast changes — but it's not the same as having a causal physical understanding. For operational forecasters, that matters. When the model predicts something unusual, they want to know whether it's seeing a real physical signal or hallucinating.
Corn
And yet. A minute on a single chip versus hours on a supercomputer. That's not a small difference.
Herman
It's not. And I think the most likely near-term future is hybrid. You use the fast AI model to run a massive ensemble — hundreds of members, exploring the full probability space — and then you use the physics model to verify the most concerning scenarios. Or you use the AI model for the first seven days and then hand off to the physics model for the extended range. Or you use the AI model as a proposal generator for the data assimilation system. There are a lot of architectures being explored, and the field is moving extremely fast.
Corn
The other thing that strikes me — all of this is downstream of the observation network. If the satellites go dark, or the weather balloon programme gets cut, or the aircraft reports stop coming, the models have nothing to assimilate. AI or physics, doesn't matter.
Herman
That's the unglamorous infrastructure layer that makes everything else possible. The global observing system is a patchwork of satellites, radiosondes, surface stations, aircraft, buoys, and radar, coordinated through the World Meteorological Organization. It's expensive and it's fragile. During COVID, the loss of aircraft observations — because commercial flights were grounded — measurably degraded forecast skill. The aircraft provide temperature and wind profiles at cruise altitude that nothing else can match. The system is more robust than it looks, but it's not infinitely robust.
Corn
And the satellites are the backbone now.
Herman
Polar-orbiting satellites give you global temperature and moisture soundings twice a day. Geostationary satellites give you continuous imagery and atmospheric motion vectors — you track cloud features from one image to the next and infer the wind. The big leap in forecast skill in the Southern Hemisphere over the last forty years is almost entirely due to satellite data. Before satellites, the Southern Ocean was a data desert. Now it's well-observed, and the forecasts are nearly as good as the Northern Hemisphere.
Herman
I'm going to stop there because I want to hear what Hilbert thinks of all this. There's no way he doesn't have an angle on weather forecasting.

Hilbert: The Met Office, Bracknell. Nineteen eighty-seven. I was in the mail room.
Herman
Of course you were.

Hilbert: October fifteenth, nineteen eighty-seven. The great storm. Michael Fish went on the BBC the night before and told everyone not to worry, a woman had called in about a hurricane and he said there wasn't one coming. Then the storm hit, twenty-two people died, fifteen million trees down. I was in the mail room the next morning and nobody was talking. The place was silent.
Corn
The forecast missed it.

Hilbert: The model said the storm would track up the Channel. It didn't. It went across southern England. The French model got it right. The French model.
Herman
That storm is a famous case study. The observing network missed a key feature in the initial conditions — a small low-pressure centre that deepened explosively.

Hilbert: After that, they bought more computing power. The Treasury opened the purse. A bad forecast that kills people is the best thing that ever happens to a forecasting centre's budget. Nobody wants to hear that, but it's true.
Corn
There's a grim logic to it.

Hilbert: The other thing nobody talks about is the fax machines. In nineteen eighty-seven, the forecasts went out by fax. If the fax machine jammed, the forecast didn't arrive. We had a whole room of fax machines and a man whose job was to unjam them. His name was Keith. He had a stick.
Herman
A stick.

Hilbert: He'd walk up and down the rows, and if a machine started making the noise — the grinding noise — he'd whack it on the side with the stick. Worked about sixty percent of the time.
Corn
What happened the other forty percent?

Hilbert: We'd get a call from an RAF base asking where their forecast was, and we'd say it's in the machine, and they'd say get it out of the machine, and we'd say we're trying. Keith retired in ninety-four. They gave him a plaque. The plaque said "For services to facsimile transmission." I'm not joking.
Herman
The whole edifice of numerical weather prediction — Bjerknes, Richardson, von Neumann, the supercomputers, the satellites, the ensembles — and it all depends on a man with a stick.

Hilbert: It did for a while. The stick was a broom handle with gaffer tape round the end. I still have it.
Corn
Of course you do.

Hilbert: Keith gave it to me when he left. Said I'd need it. I did.
Herman
The thing that strikes me about that story is that it's not really about the model. The model might be excellent. But the chain from the model output to the decision — that's where the failure happens. The fax machine, the interpretation, the communication. The forecast can be right and still be useless if it doesn't reach the person who needs it in a form they can act on.

Hilbert: The forecast for the eighty-seven storm was wrong. But the French one wasn't. And nobody in England saw the French forecast because there was no system for sharing them. The Met Office got the French output on a telex, but the duty forecaster didn't look at it because he was busy with the UK model. That's not a physics problem. That's a human problem.
Corn
One forecaster, one telex, one decision not to look.

Hilbert: That's the whole thing. Everyone thinks weather forecasting is about the equations. It's not. It's about whether the person on shift trusts the output of the foreign model more than their own, at two in the morning, with a storm coming. And most of the time they don't.
Corn
The AI models change that calculus, though. If the French model runs on a laptop in thirty seconds, there's no reason not to look at it.

Hilbert: There's always a reason not to look at it. The reason is you've got seventeen things to look at and you're tired. The technology changes. The human doesn't.
Herman
I think that's about right. We've been talking about the arc of prediction — Bjerknes to GraphCast — but Hilbert's point is that the last metre, the bit where the forecast meets the decision, hasn't fundamentally changed. It's still a person, looking at information, under time pressure, making a call. And that's where the failures happen, and where the successes happen too.
Corn
The stick is optional.

Hilbert: The stick is never optional.
Corn
This has been My Weird Prompts. Thanks to our producer, Hilbert Flumingtop, who apparently owns a piece of meteorological history and a broom handle with gaffer tape.
Herman
If you want more episodes, we're at my weird prompts dot com. New episodes every day. We'll be back tomorrow.

This episode was generated with AI assistance. Hosts Herman and Corn are AI personalities.