Last episode we traced weather forecasting from Lewis Fry Richardson's beautiful, doomed forecast factory through to satellite-fed supercomputers, and Daniel came back with the obvious next question. If the models are that good now, what's left for the humans? He's been watching the Israel Meteorological Service — their public-facing setup looks like something out of a space program — and he wants to know what a forecaster there actually does in a shift in 2026. Three specific things he's asking: walk us through the daily workflow, tell us how much of the job is knowing where the models fail from hard-won experience, and what the shift looked like this year at a place like the IMS.
And the timing's right, because the IMS is a perfect case study. About two hundred staff, maybe sixty forecasters, and they're dealing with some of the most complicated local weather on the planet — Mediterranean fronts colliding with desert heat, the Jordan Rift Valley screwing up every wind model ever written. If you want to see where the human still matters, that's the place to look.
So we're going to shadow a forecaster through a 2026 shift — not in theory, in the concrete workflow — and figure out where the human is indispensable and where they're starting to look like a bottleneck. Let's get into it.
Here's the tension that makes this worth a whole episode. The ECMWF's high-resolution model runs at nine kilometers globally. Google's GraphCast and the other AI models are matching or beating physics-based systems on a lot of standard metrics. You'd look at that and think — why is any national service still employing dozens of forecasters? The answer isn't that the models are bad. It's that models are general, and weather is local, consequential, and absolutely full of edge cases. A model will give you a forty percent chance of thunderstorms. That's a number. A forecaster has to decide whether that number means cancel the outdoor wedding in Netanya or just keep an eye on it.
The model outputs a probability density function. The forecaster outputs a decision.
No, wait, I just agreed with a setup. Let me start that again. The forecaster's job is to collapse the probability into a specific action for a specific audience. And the IMS has to do that for audiences that couldn't be more different. Ben Gurion Airport needs to know about wind shear and visibility for the next three hours. The Water Authority needs to know about precipitation over the Kinneret watershed for the next three days. The Emergency Management Authority needs to know right now whether a flash flood warning should go out to communities in the Arava. Same model output, completely different translations required.
And Daniel mentioned that the IMS looks impressively high-tech from the outside. It is. They've got dual-polarization C-band radars, they're pulling satellite feeds from Himawari-9, GOES-18, and the European MTG-I1, and they run their own high-resolution WRF model at two-point-five kilometers over Israel. But the monitors full of radar displays aren't just for show. Someone has to sit in front of them and know what they're looking at.
Let's walk through the shift. It starts at six in the morning with a handover from the night forecaster. The night shift is skeleton crew — usually one or two people — and they've been watching the overnight model runs come in. The zero-zero UTC runs from ECMWF, GFS, ICON, and the IMS's own WRF. The night forecaster briefs the incoming team on what developed while everyone was asleep. Maybe a convective system popped up over the Negev that wasn't in the previous day's forecast. Maybe there's a dust storm forming over Egypt that's going to reach the coast by noon.
So the first thing the morning forecaster does is sit down and compare the zero-zero Z runs across all the models. And this isn't just reading the pretty output maps. They're overlaying the satellite imagery, the lightning detection network data, and the radar composites. They're looking for where the models disagree with each other and where they disagree with what's actually happening right now.
That's the critical skill that takes years to develop. You're looking at four or five different versions of the future and you have to decide which one to believe for which parameter. Here's a concrete example. The GFS model has a known systematic bias in the Jordan Valley. It overpredicts wind speeds by fifteen to twenty percent because its representation of the topography is too coarse — the valley is narrow and deep, and a nine-kilometer grid just smears it out. Every forecaster at the IMS knows this. It's documented in their model bias climatology, but it also lives in the heads of the senior forecasters. You don't just read the GFS wind output for the Jordan Valley — you mentally knock fifteen percent off it before you even start.
Which raises the question Daniel was getting at. How much of this job is a database lookup and how much is a craft?
It's both, and the balance is shifting. The IMS maintains a formal bias database that tracks how each operational model performs for different parameters, regions, seasons, and synoptic patterns. So a junior forecaster can look up "GFS, spring, Jordan Valley, wind speed" and see the correction factor. But the tacit knowledge — the senior forecaster who knows that the WRF model develops a dry bias in spring because it doesn't handle the Sharav low properly, the hot dry wind from the east — that's harder to transfer. It takes two to three years for a new forecaster to become fully operational, and most of that isn't learning the software. It's building the mental model of how the actual weather in this specific complicated place relates to what the models predict.
So by eight in the morning, they've ingested all of this, they've formed a picture of what the atmosphere is doing and what it's likely to do, and now they have to produce forecasts for audiences that care about completely different things. Walk me through that.
Three main products first thing. The public forecast — a three-day outlook that goes on the website and the app. That's what you check to see if you need a jacket. Then the aviation forecast — TAFs for Ben Gurion, Eilat, and Haifa. A TAF is a terminal aerodrome forecast, and it's incredibly specific. Wind direction and speed, visibility, cloud ceiling, any significant weather within eight kilometers of the airport. It's valid for a specific time window and it gets updated every few hours. If you get it wrong, planes can't land.
And the thresholds are completely different from the public forecast. The public wants to know if it'll rain. The pilot wants to know if there's going to be wind shear on approach to runway two-six.
Right. Then the marine forecast for the Mediterranean — wave height, swell direction, wind speed at the surface. The shipping lanes off Haifa and Ashdod depend on that. And then the forecaster checks the red flag criteria. The IMS has automated alerting for parameters that approach dangerous thresholds — flash flood potential in the wadis, heat stress, hazardous air quality from dust events. The system flags them, but the forecaster makes the final call on whether to escalate to the Emergency Management Authority. That decision cannot be automated, because the cost of a false alarm is real — you evacuate a community for nothing, and next time they might not listen.
That's the moment where the model says "seventy percent chance of heavy precipitation in this catchment" and the forecaster has to decide whether seventy percent is high enough to pull the trigger. And that decision is weighted with everything they know about the specific wadi, the time of year, the soil saturation from last week's rain, and whether the model has been overpredicting precipitation in that region all season.
Mid-morning, the forecaster joins a coordination call with the Israel Electric Corporation and the Water Authority. Both are major users of tailored forecasts and both are making operational decisions worth millions of shekels based on what the forecaster tells them. The electric company needs wind and solar generation forecasts for the next forty-eight hours so they can decide whether to spin up gas turbines or draw from storage. The Water Authority needs precipitation forecasts for reservoir management — if the Kinneret is going to get a major inflow, they might release water preemptively to avoid flooding downstream.
And the forecaster isn't reading them a number off a screen. They're interpreting the ensemble spread. The ECMWF ensemble has fifty members, each with slightly different initial conditions, and the forecaster is saying — look, most members show between five and fifteen millimeters in the watershed, but there's a cluster of about eight members showing thirty-plus, and that cluster is the one that handles Mediterranean lows best. So I'd put the probability of significant inflow at about twenty percent, but I want you to know that twenty percent is there.
That's not a technical skill. That's judgment, communication, and — honestly — a relationship. The person on the other end of that call has been talking to this forecaster for years. They know when the forecaster is worried. They can hear it in their voice. That's not something you can replace with a dashboard.
So that's the routine day. But the real value of the forecaster shows up when things go wrong — when the models disagree and someone has to make a call.
And here's where we get to the question about institutional knowledge. The IMS has a case study that's become a kind of legend internally. November 2024, flash floods in the Arava Valley. Multiple models showed moderate precipitation — nothing that would trigger an automatic warning. But a senior forecaster recognized a specific pattern. A cut-off low with a moisture trajectory coming up from the Red Sea, feeding into a convergence zone over the eastern Negev. She'd seen it before, in 2018, and it had produced catastrophic flooding then. The models didn't flag it because the absolute precipitation numbers didn't look extreme — but the combination of the synoptic pattern and the moisture source was almost identical to the 2018 event.
And the AI post-processing system didn't catch it either.
No, because it hadn't seen enough examples of that specific configuration in its training data. It's a one-in-ten-year event. The AI is great at the ninety-fifth percentile of weather. It's useless at the ninety-ninth. So the forecaster escalated the warning to extreme six hours before the event. Communities were evacuated. The floods came exactly as predicted. Without that human in the loop, the warning wouldn't have gone out until the water was already rising.
That's the case for the defense. But there's a 2026 twist here, and I want to sit with it. The IMS now runs a machine learning post-processing system — similar to what ECMWF is doing with their AIFS, but trained on local data — that learns model biases automatically and produces corrected output. It's designed to do exactly what the senior forecaster did: notice systematic errors and correct for them. And for the routine ninety-five percent of forecasts, it works. It's reducing the need for human bias correction on the day-to-day.
The division of labor is shifting. AI handles the routine forecasts, freeing forecasters to focus on the high-impact, low-probability events, on tailoring outputs for specific users, and on communication with emergency managers. The job title isn't changing, but the job is. It's moving from predictor to interpreter and risk communicator.
Which raises an uncomfortable question. As the AI models get better, and as the training datasets expand to include more rare events — either through synthetic data generation or just longer observational records — does the forecaster's edge erode? If the AI eventually sees enough cut-off lows with Red Sea moisture trajectories, does it catch the next Arava flood on its own?
Some services are already betting yes. The Finnish Meteorological Institute has been experimenting with what they call forecaster-on-the-loop. The AI produces the forecast, and the human only intervenes when the system's confidence is below a threshold. The forecaster isn't building the forecast — they're monitoring it, like a pilot watching the autopilot. If everything looks normal, they let it run. If something looks off, they step in.
And the IMS?
More conservative. They require human sign-off on all warnings. Every alert that goes to the Emergency Management Authority has a forecaster's name on it. That's partly regulatory, but it's also cultural. The IMS has a institutional memory of events like the Arava floods, and they're not ready to hand the keys to a system that's never been tested on a genuine outlier.
So we've got a spectrum. Finland at one end, running an experiment in human-supervised automation. Israel at the other, still requiring a human to pull the trigger on every warning. And most national services somewhere in between.
And here's what I think Daniel was really getting at with his question about institutional knowledge. It's not just about knowing that the GFS overpredicts wind in the Jordan Valley. It's about the mental model the forecaster builds over years of watching the atmosphere behave in this specific place. They know that a certain upper-level trough configuration reliably produces fog in the Hula Valley but not in the Coastal Plain. They know that the WRF model has a dry bias in spring because of the Sharav. They know that when the pressure gradient looks like this, the sea breeze front will stall over Tel Aviv and produce thunderstorms right at rush hour.
Some of that is documented. The IMS has a formal bias climatology. But a lot of it isn't. It's tacit knowledge that lives in the heads of the senior forecasters and gets passed down through the shift handovers, through the coordination calls, through the years of working alongside people who've seen things the models haven't.
The question is whether that tacit knowledge is a permanent feature of the job or a transitional phase. As the AI post-processing gets better, as the models get higher resolution, as the training datasets get larger — does the institutional knowledge become obsolete? Or does it just shift to a higher level? Instead of knowing the GFS wind bias, maybe the forecaster of 2036 will know the AI's failure modes. Which synoptic patterns the AI handles well and which ones it doesn't.
That's the optimistic case. The pessimistic case is that the AI eventually knows its own failure pattern better than the human does, and the forecaster becomes a liability rather than an asset — someone who overrides the model and makes things worse.
There's some evidence for that concern already. Studies from the UK Met Office have shown that forecaster overrides of model guidance are wrong about as often as they're right on routine forecasts. The human edge really only shows up in the high-impact, low-probability events — the ones where the model hasn't seen enough examples. If the models eventually see enough examples of everything, the human edge might disappear entirely.
But we're not there yet. And the Arava flood case suggests we're not close. The AI post-processing system didn't flag it in 2024, and that's with years of training data. The atmosphere has a lot of ways to surprise us.
Which brings me to the other dimension of this that I think is underappreciated. Climate change is shifting the baseline. The thirty-year normals that forecasters used to rely on are becoming less reliable. Events that used to be one-in-a-hundred-year are happening every decade. The models are constantly playing catch-up with a climate that's moving under their feet. In that environment, the forecaster's role as a sanity check — someone who can look at the model output and say "that doesn't look right, given what we've been seeing lately" — might actually become more important, not less.
The job isn't static. It's evolving in response to both better AI and a more chaotic climate. And the forecaster of 2026 is doing something qualitatively different from what the forecaster of 2006 was doing, even if the job title is the same.
Let's talk about what the afternoon shift looks like, because it's different from the morning. The morning is about production — getting the forecasts out, briefing the key users. The afternoon is more about monitoring and updating. The forecaster is watching the radar, watching the satellite, watching whether the convection that was supposed to fire up over the Judean Hills at two PM actually fired up at noon. If it did, the aviation forecast for Ben Gurion might need to be amended. If the sea breeze front is moving faster than expected, the marine forecast might need updating.
They're also doing the shift handover documentation. Writing up what happened during their shift, what they changed and why, what the incoming forecaster should watch for. That documentation feeds back into the institutional knowledge base. Over time, it becomes part of the record that junior forecasters study to understand how the atmosphere behaves in this place.
Then there's the training component. Senior forecasters at the IMS spend part of their time working with the new hires, walking them through cases, explaining why they made the calls they made. That apprenticeship model is still the primary way the tacit knowledge gets transferred. You can't learn it from a manual. You have to sit next to someone who's been doing it for twenty years and absorb it.
Okay. So we've walked through the shift, we've talked about institutional knowledge, and we've talked about how AI is changing the role. I want to pull on one more thread before we bring Hilbert in. Daniel mentioned the API — the IMS has a public API that lets you pull forecast data programmatically. That's part of a broader trend toward open data in national weather services. And it's changing who the forecaster's audience is.
Right. Ten years ago, the forecaster produced forecasts for a handful of defined audiences — the public, aviation, marine, agriculture. Now the data is out there, and anyone can build an application on top of it. A startup can take the IMS forecast and combine it with soil moisture data to tell farmers exactly when to irrigate. A logistics company can use it to route trucks around incoming storms. The forecaster isn't just serving the traditional users anymore — they're serving an ecosystem of downstream applications, and they have to think about how their forecast will be interpreted by machines as well as humans.
That's a new kind of pressure. If your forecast is being ingested by an automated irrigation system that's going to water or not water a thousand hectares based on your precipitation probability, the stakes are different. You can't rely on the human at the other end to apply their own judgment. The machine is going to take your number and act on it.
Which circles back to the communication point. The forecaster's job is increasingly about understanding how different audiences will use the forecast and tailoring the output accordingly. It's not just about getting the weather right. It's about presenting it in a way that leads to the right decision by the right person at the right time.
Hilbert: I spent six months in 1994 as a junior forecaster at the UK Met Office. Before they had the Unified Model running operationally. My job was drawing isobars on paper charts and phoning them to the BBC. I quit because I was bored. I thought computers would make the job obsolete within a decade.
Here you are, thirty-two years later, listening to us talk about sixty forecasters at a mid-sized national service.
Hilbert: I was spectacularly wrong, obviously. But what strikes me about this picture you're painting is how much the job has become about managing relationships. In 1994 I handed a forecast to a switchboard operator who read it to the public. Now the forecaster is in a meeting with the head of the Water Authority explaining why the ensemble spread is wider than usual. That's not a technical skill. That's diplomacy. And AI is terrible at diplomacy.
That's a different way of framing the human edge. It's not about beating the model on accuracy. It's about being the person the emergency manager trusts when you tell them to evacuate.
Hilbert: The Met Office had a forecaster named Margaret. She'd been there since the sixties. She could look at a synoptic chart and tell you what the weather would do in the Channel in six hours without touching a computer. But that's not why people called her. They called her because she'd been right in front of them a hundred times. When Margaret said it was going to be bad, you cancelled the ferry. The model could say the same thing and you'd still call Margaret to check.
Trust built over years of being right, and occasionally wrong, together. That's not something you can automate with a better loss function.
Hilbert: The IMS forecaster on that coordination call with the electric company — they're not just transmitting a number. They're maintaining a relationship that's been built through a hundred previous calls, some of which went badly. The electric company remembers the time the forecaster said thirty percent chance of high winds and it turned into a dust storm that knocked out transmission lines. The forecaster remembers it too. Next time they say thirty percent, both sides know what that number actually means in practice.
The forecast number is almost a shorthand for a shared history. The number means something different to those two people than it would to an outsider reading it on an app.
Hilbert: That's the part I don't think the automation people understand. They think the forecast is the product. It's not. The forecast is a prompt for a conversation. And the conversation is where the actual work happens.
That's a sharper version of what I was trying to say earlier about the forecaster as interpreter and risk communicator. The forecast is the starting point, not the endpoint.
Hilbert: I still have my Met Office training manual. It's in a box somewhere. The section on drawing isobars is about forty pages long. There's a whole chapter on the correct angle to hold the pencil. None of it is relevant anymore. But the chapter on briefing external users is about four pages and it's mostly just "be clear and don't panic." I think we've got the ratio backwards.
The manual was teaching you to do the thing the computers now do better, and barely touching on the thing that's become the whole job.
Hilbert: That's about right.
Where does that leave us? The forecaster's job in 2026 is a hybrid of three things. One, monitoring and correcting the models, which is shrinking as the AI post-processing gets better. Two, making judgment calls on high-impact events, which is still a human domain but might not be forever. Three, managing relationships and communicating risk, which is expanding and might be the hardest thing to automate.
The open question is whether that third thing is enough to sustain a profession. If the first two shrink to near zero, does the communication role alone justify sixty forecasters at the IMS? Or does it become a much smaller job, a handful of senior people who mostly talk to other humans while the machines do the rest?
I don't think we know yet. But I think the Finland experiment is going to be instructive. If forecaster-on-the-loop works — if the AI produces the forecast and the human only intervenes when confidence is low — then the communication role might be the only thing left. And I'm not sure that's a sixty-person job.
Unless the communication demands grow faster than the prediction demands shrink. If every municipality wants a tailored briefing, if every major event needs a forecaster embedded with the organizers, if the climate keeps throwing unprecedented events at us — maybe the communication role expands to fill the space the prediction role left behind.
Daniel's question was about what forecasters actually do. And the answer, I think, is that they're doing something different from what they were doing ten years ago, and they'll be doing something different again in ten years. The job isn't dying. It's mutating. And the mutation is toward the human things — judgment, trust, communication — and away from the computational things.
The cutting-room floor detail I wanted to mention: the IMS WRF model at two-point-five kilometers is actually not their highest resolution. They run an experimental nest at seven hundred fifty meters over the Tel Aviv metropolitan area for urban heat island studies. It's not operational yet, but it's a glimpse of where things are going — hyperlocal forecasting for specific cities, which creates a whole new set of interpretation challenges.
That's a good place to leave it. Thanks to our producer Hilbert Flumingtop for keeping us honest, as always.
This has been My Weird Prompts. If you want to dig into the IMS API yourself, links are in the show notes. And if you have a weird prompt about another invisible profession that's quietly transforming, send it to show at my weird prompts dot com.
We'll be back soon.