Briefly: On August 6, 2026, Google DeepMind and Google Research published the WeatherNext Cyclones results in Nature and released the code and weights for WeatherNext 2. In evaluations on 2023–2025 cyclones, the model delivered an average lead-time advantage of one day or more in forecast accuracy for track, intensity, and wind structure. For businesses, the news matters not because it promises “perfect weather,” but because it makes it possible to see several scenarios earlier—and connect them in advance to decisions that can be validated.
This article is intended for chief operating officers, logistics leaders, retail, energy, agriculture, and data teams. We examine global AI weather forecasting and its role in business processes; we do not assess accuracy for any specific location in Russia and do not replace official meteorological warnings.
Updated: August 9, 2026. The facts about the model and the research are current as of this date.
Contents
- What happened on August 6, 2026
- What WeatherNext 2 and WeatherNext Cyclones are
- What the result proves—and what it does not
- Where AI weather forecasting changes business decisions
- The PDCA framework: from probability to action
- How to test a scenario in shadow mode
- How to choose a weather data source
- Metrics that matter more than overall accuracy
- Frequently asked questions
- How AI dawn helps embed the weather signal
- Conclusion
What happened on August 6, 2026
In the article Operational Tropical Cyclone Forecasting with AI Nature describes WeatherNext Cyclones as a global AI ensemble forecasting model. An ensemble is not one answer, but many possible weather scenarios that make it possible to estimate the range and probability of outcomes.
The authors tested the model on tropical cyclones from 2023–2025. According to Nature, forecasts of track, intensity, and wind radii gained an average lead-time advantage of a day or more versus leading operational models. Lead time here means the amount of time over which a forecast remains comparably useful: in average accuracy terms, the new system’s three-day forecast was comparable to the previous two-day level.
Google DeepMind says the model was trained on nearly 20 TB of global atmospheric data and a database of nearly 5,000 historical storms. The system produces forecasts up to 15 days out, and the ensemble scales to 1,000 possible scenarios. According to the developers, a single 15-day forecast is generated in under a minute on TPU.
| Fact | Verified context | What cannot be transferred automatically |
|---|---|---|
| a day or more of lead-time advantage | 2023–2025 cyclones, compared with leading operational models | accuracy for any city and any type of weather |
| up to 1,000 scenarios | WeatherNext Cyclones ensemble | 1,000 ready-made business decisions |
| forecast up to 15 days | global medium-range model | reliable 15-day inventory ordering without local validation |
| nearly 20 TB and nearly 5,000 storms | training described by Google DeepMind | the presence of a specific company’s data |
| code and weights are open | WeatherNext repository | free production infrastructure and a ready-made SLA |
The main change on August 6 was not the appearance of WeatherNext 2 from scratch: Google had introduced the base model earlier. The news is the published cyclone evaluation, collaboration with specialized meteorologists, and the release of the code and weights for WeatherNext 2 and WeatherNext Cyclones.
What WeatherNext 2 and WeatherNext Cyclones are
WeatherNext 2 is a family of global medium-range weather models from Google DeepMind and Google Research. Google Developers documentation indicates global coverage, forecast steps up to 15 days, and variables such as temperature, humidity, wind, pressure, and precipitation. Google recommends WeatherNext 2 for new projects in the WeatherNext family.
WeatherNext Cyclones is a configuration and research setup focused on tropical cyclones. The model simultaneously forecasts the global atmosphere and cyclone characteristics: track, intensity, and wind structure. This is the specific task for which Nature publishes the latest comparative evaluation.
The difference is fundamental. If a retail chain wants to forecast ice cream sales, it should not take the cyclone track metric and treat it as proof of demand forecast quality. It needs weather variables for specific locations, sales history, promotions, the calendar, inventory, and its own loss function.
Data can be obtained in several ways. The official repository lists Google Cloud, Weather Lab, and Open-Meteo; running it yourself requires ML infrastructure and expertise. This is a model layer, not a complete logistics or inventory management system.
What the result proves—and what it does not
The study proves measurable improvement in the stated sample and on the stated meteorological metrics. The authors compared forecasts on cyclones from 2023–2025, described an ensemble methodology, and involved specialists from the National Hurricane Center, CIRA, and the UK Met Office. That is substantially stronger than a marketing demo.
But even a high-quality forecast is not the same as a high-quality decision. Four layers still remain between them:
- Localization: whether the global resolution is sufficient for the specific location, road, warehouse, or asset.
- Mapping: how the weather variable relates to sales, incidents, ETA, or output in your own history.
- Action rule: which threshold changes the plan and who is authorized to change it.
- Cost of error: which is more expensive—extra caution or a missed event.
Google separately notes that official forecasts and warnings should be obtained from the local meteorological service. So WeatherNext can be viewed as an additional analytical source or product component, but not as a standalone emergency alert channel.
There is also a scientific limitation. As of publication, Nature shows an unedited manuscript: the paper has been accepted but is still undergoing editorial processing. The result is published, but wording and formatting may change before the final version.
Where AI weather forecasting changes business decisions
AI forecasting is useful where an event is observable in advance, a decision can be changed before it happens, and the impact of that decision is recorded. If a team cannot act or does not store outcomes, extra accuracy will remain a nice map.
| Process | Weather signal | Possible decision | What counts as an error |
|---|---|---|---|
| Long-haul logistics | probability of strong winds, precipitation, or extreme temperatures along the corridor | change the time window, route, or backup transport | delays, extra mileage, unnecessary reroutes |
| Last Mile | short-term precipitation and temperature by zone | adjust shifts and the promised ETA | courier underload/overload, SLA breach |
| Retail | temperature, precipitation, humidity by store | change the order, promotion, or merchandising of weather-sensitive SKUs | out-of-stock, write-offs, extra discounting |
| Energy | wind, cloud cover, and temperature | refine output, load, and reserve capacity | imbalance, unavailable reserve, incorrect maintenance |
| Agriculture | precipitation, wind, humidity, and extreme-event risk | reschedule treatment, irrigation, or harvest | lost window, overspend, crop damage |
| Facilities Operations | heat, frost, wind, snow load | increase on-call coverage or reschedule work | downtime, false dispatch, missed incident |
These solutions are not ready-made recommendations: thresholds depend on the asset, contract, region, and cost of error. For the Russian market, it is useful to compare the global layer with a local commercial source. For example, Yandex Weather for Business offers an API, dashboard, historical data, and custom triggers for retail, energy, agriculture, and logistics.
PSDK Framework: From Probability to Action
AI Dawn offers the PSDK framework: Threshold → Scenarios → Action → Control. This is an analytics framework for process design, not a feature of WeatherNext.
1. Threshold
Start by defining the event that truly changes operations. Not “bad weather,” but, for example, the probability of wind above the allowable limit for high-rise work between 12:00 and 18:00, or the probability of precipitation after which delivery time historically increases.
You should not take a threshold from someone else’s case study. It is calibrated on historical data to account for the cost of two errors: a false alarm and a miss. If stopping an operation is expensive, the threshold may be higher; if a miss creates risk for people or infrastructure, the priority shifts toward caution and mandatory human decision-making.
2. Scenarios
One forecast hides uncertainty. An ensemble shows the range: what is likely, what is possible, and how far the options diverge. For the process, it is useful to store not only the mean, but also quantiles or probabilities of crossing a critical threshold.
Example: the average wind speed does not exceed the limit, but 20% of scenarios cross the danger line. This is not an automatic cancellation of work; it is a prompt to apply a pre-agreed rule — request confirmation, check a second source, or reduce the time window.
3. Action
Match each risk range with an allowed action:
- low risk — no change to the plan;
- medium — notify the owner and verify with a second source;
- high — prepare a backup plan;
- critical — decision by the responsible specialist under the approved procedure.
First automate reversible actions: recalculate ETA, suggest a backup route, prepare a draft order. Canceling a trip, stopping an asset, or changing pricing requires separate policy and authority.
4. Control
After the event, save the forecast, the decision made, the actual weather, and the business outcome. Without this, the system does not learn from mistakes and cannot answer whether the earlier signal helped.
weather scenarios → process threshold → action option → confirmation → actual outcome → error review
The point of PSDK is that the model does not directly control the process. It supplies scenarios; rules and people determine the permitted action; control feeds the facts into the next cycle.
How to test a scenario in shadow mode
Shadow mode is a mode in which a new system generates forecasts and recommendations but does not change the production process. The team compares its decisions with the current workflow and the actual outcome. This reduces risk and creates a fair baseline.
A practical four-week protocol is an estimated implementation plan, not a universal timeline:
- Week 1 — task and history. Choose one process, 20–100 relevant episodes, and the current metric: SLA, write-offs, downtime, output, or reserve cost. Record the decisions made at the time before reviewing the new forecast.
- Week 2 — sources. Connect at least two weather sources. Normalize time, coordinates, units, horizons, and missing values to a single schema. Separate forecast issue time and valid time, or you will get future leakage.
- Week 3 — rules. Run PSDK on historical data. For each event, calculate what action the system would have suggested and what actually happened. Do not change the threshold after every failed case without a separate validation set.
- Week 4 — shadow production. Run recommendations in parallel with the real process. The owner explains why they accepted or rejected the advice. Automatic intervention stays off.
Before rollout, define stop conditions: a sharp rise in false alarms, a missed high-risk event, source degradation, time sync issues, or no accountable owner. For more on choosing the process and baseline, see the guide how to start implementing AI into business processes.
How to choose a weather data source
The choice is not limited to “the smartest model.” Compare how ready the layer is for your operation.
| Approach | When it makes sense | Advantages | Limitations |
|---|---|---|---|
| Ready-made local API/platform | you need a fast launch in Russia and supported business parameters | localization, SLA/support, dashboard, and triggers | vendor lock-in, pricing, limited control over the model |
| WeatherNext data feed | you have a data team and need a global probabilistic layer | scenarios, global coverage, fast access to output | you need your own feature/decision layer |
| Self-host WeatherNext | research, a specialized model, or control over compute | code and weights, reproducibility, customization | ML infrastructure, updates, initial conditions, monitoring |
| Combined Circuit | the cost of a mistake is high and a cross-check is needed | comparison of global and local sources | more complex data, conflict resolution rules, and cost |
For Russia, WeatherNext’s global coverage means a forecast grid is available, but does not prove local accuracy for a specific road, field, or site. Check spatial resolution, elevation, terrain, observation availability, and historical bias. Where block- or street-level scale matters, the global model usually needs to be supplemented with local radar, stations, or a commercial nowcast.
Also review the licenses and terms for each component separately. Open source does not mean free compute, a free data feed, guaranteed SLA, or an automatic right to send production data to any service.
Metrics that matter more than overall accuracy
Overall accuracy is often useless: a rare dangerous event can be “accurately” not predicted almost all the time. You need forecast and decision metrics.
| Metric | Formula / meaning | Management question |
|---|---|---|
| Lead time | time between the signal and the event at acceptable quality | do we have time to take action |
| False alarm rate | false alarms / all alarms | how much unnecessary rework does the system create |
| Miss rate | missed events / all actual events | what risk remains unseen |
| Decision coverage | cases with an applicable rule / all cases | where the forecast is actually tied to the process |
| Override rate | human-rejected recommendations / all recommendations | are the rules understandable and trusted by the owner |
| Decision loss | the cost of action and error according to the chosen loss function | did the process get better than the baseline in money or SLA terms |
Example: source A is better at predicting rain, while source B detects rare heavy precipitation earlier. For delivery, A may be better for average ETA, while B is better at preventing disruptions on critical days. The winner depends on the loss function, not on a single ranking.
Compare four options: the current process without the new signal, each source individually, and the combined rule. If improvement appears only after manual intervention by an experienced dispatcher, that is not a failure: the right product may be a decision assistant, not an autonomous agent. A similar two-loop principle — generation and verification — is covered in the article on AI business tasks.
Frequently Asked Questions
What is WeatherNext 2?
WeatherNext 2 is a family of Google DeepMind and Google Research global medium-range weather models. The documentation indicates global coverage, a horizon of up to 15 days, and an ensemble forecast for WeatherNext 2. A recent Nature paper separately evaluates WeatherNext Cyclones on tropical cyclones from 2023–2025.
Is WeatherNext 2 free?
The code and weights are open in Google DeepMind’s repository, and ready-to-use forecast streams are available through several platforms. But compute, storage, cloud queries, commercial APIs, and production support may cost money. Check the current licenses and pricing for the channel you choose.
Does WeatherNext replace official warnings?
No. Google explicitly recommends using your local meteorological service or national weather service for official forecasts and warnings. For critical decisions, the AI model should be an additional source within an approved procedure.
Is WeatherNext 2 suitable for Russia?
The model is global, so it covers Russian territory at the global grid level. That does not confirm its quality for a specific city, road, or site. You need a backtest on actual outcomes and a comparison with a local source.
What should be measured in a weather forecast pilot?
Measure lead time, false alarm rate, miss rate, decision coverage, override rate, and decision loss versus the current baseline. Separately track data quality, missing values, update latency, and cases where sources conflict.
Where should you start with a weather signal implementation?
Choose one process, its current metric, the source of truth, constraints, and the acceptance criterion. Then run a shadow-mode backtest without automatic action. Only after measuring errors should you allow reversible actions.
How AI Dawn helps embed a weather signal
AI Dawn connects the forecast not to a model demo, but to a specific process and decision criterion. For a weather-dependent task, the team can take four actions:
- Describe the process, baseline, loss function, and acceptable thresholds.
- Prepare and synchronize weather, operational, and actual outcome data.
- Build an MVP forecast loop and run historical and shadow-mode testing.
- Integrate the validated rules with ERP, 1C, an analytics platform, or an internal system, and add a decision log.
Safe first step: choose one process, its current baseline, data sources, constraints, and acceptance criterion. That is enough to determine whether you need a ready-made API, an additional WeatherNext layer, or your own model.
Conclusion
WeatherNext Cyclones is important news: on August 6, 2026, Nature published an evaluation showing a lead-time advantage of a day or more for cyclone forecasting, and Google released the code and weights for WeatherNext 2. This expands access to probabilistic forecasting, but it does not turn the model into a ready-made dispatcher.
For business, value appears only in the full chain: weather scenario → process threshold → allowed action → confirmation → actual result. Use the PSDK loop, start in shadow mode, and compare sources by the cost of errors, not by overall accuracy.
The first practical step is to take one weather-dependent process and 20–100 historical episodes. If the new signal provides more time and reduces decision loss versus the baseline, expand the loop. If not, you have a useful negative result before expensive integration.