Short- to long-range climate forecasts with deep learning
Copernicus Publications (2026)
Abstract:
Uncertainty in projections of future regional climate change remains large, driven by structural differences among Earth System Models and the influence of internal climate variability. Existing uncertainty-reduction approaches, including emergent constraints and Bayesian variants, primarily focus on forced climate responses derived from simple aggregate metrics, thereby requiring strong assumptions and exploiting only low-dimensional climate information. Here we propose a data-driven deep-learning framework that directly forecasts spatially and monthly resolved decadal mean climatologies of surface temperature anomalies from the 2030s to the 2090s, using only recent monthly trajectories spanning 1980-2025. The training ensemble contains 265 historical+SSP2-4.5 simulations, distributed across 40 ESMs from 25 different families (i.e., modelling centers) over which the cross validation is performed. The architecture couples pluri-annual to multi-decadal temporal convolutions with a spatial U-Net encoder-decoder and is evaluated on CMIP6 simulations using a leave-one-model-family-out cross-validation (LOMFO-CV) design to ensure generalisation across separately developed ESMs. Predictive uncertainty is quantified via LOMFO-CV errors, yielding conservative and reliable ranges that incorporate irreducible internal variability and systematic model shifts.To further evaluate the predictive capacity beyond the CMIP6 distribution, we evaluated the network on historical+SSP2-4.5 simulations from a recent HadGEM3-GC5 model hierarchy developed within the European Eddy-Rich ESMs (EERIE) project, the European contribution to HighResMIP2 for CMIP7. In particular, the eddy-rich GC5-HH configuration explicitly simulates mesoscale ocean dynamics that are absent in CMIP6-type models, providing a rigorous test of generalisation to richer and more realistic physical representations. Despite these substantial differences, the network successfully reproduces warming trajectories and future climate patterns for all three model configurations (GC5-LL, GC5-MM, GC5-HH), with forecast errors largely contained within empirically calibrated uncertainty bounds from the LOMFO-CV, both globally and locally. These results, notably for GC5-HH and its more realistic physics, strengthens confidence in the applicability of the framework to real-world data.When applied to observations, the extracted end-of-century global-mean surface temperature and its uncertainty range are consistent with prior estimates from Bayesian frameworks. At local scales, the network reduces uncertainty by 40% (2030s) to 30% (2090s) on average, and by up to 75% in some regions for all future decades. Importantly, these uncertainty estimates account not only for uncertainty in the forced response (as emergent constraint methods do), but also for errors associated with predicting different realisations of internal variability, providing a physically meaningful reduction of local and global climate uncertainty.Spatial Generalization Tests for Machine Learning-based Weather Models as a Requirement for Climate Predictions
Copernicus Publications (2026)
Abstract:
Machine learning-based weather prediction is revolutionizing weather forecasting by learning from present-day climate. However, generalization to other climates remains a major challenge. With melting sea ice, land-use change and increasing ocean temperatures, boundary conditions are changing. Therefore, generalization in time will likely only be possible if generalization in space is also given. The physics of the atmosphere is invariant in space, and as such, a model should demonstrate the same to accurately represent the real world.Here, we present three test cases to evaluate whether machine learning-based weather and climate models generalize spatially and apply them to multiple AI weather models. The tests consist of reversing the entirety of the input data and boundary conditions in latitude (Test 1), reversing them in longitude (Test 2), as well as rotating them by 180˚ in longitude (Test 3), while keeping all aspects of the simulation physically consistent. For a deterministic model that generalizes in space, each of these test cases yields the same predictions as the baseline case, only subject to a rounding error. With these test cases, we investigate whether data-driven models hardcode representations of spatial relationships in the training data into their latent space. We show that currently, both fully data-driven and hybrid general circulation models do not pass these tests, instead performing poorly with unphysical results. This implies that they have likely not learned underlying atmospheric physics principles, but instead local spatial relationships statistically dependent on geographical location. This calls into question the ability of such models to simulate a changing regional climate. As such, we propose that machine learning-based climate models be evaluated using our spatial tests during model development to reduce overfitting on present-day regional climate.Data-Driven Stochastic Parameterization of MCS Latent Heating in the Grey Zone
Copernicus Publications (2025)
Abstract:
Mesoscale Convective Systems (MCSs), with length scales of 100 to 1000 km or more, fall into the "grey zone" of global models with grid spacings of 10s of km. Their under-resolved nature leads to model deficiencies in representing MCS latent heating, whose vertical structure critically shapes large-scale circulations. To address this challenge, we use analysis increments—the corrections applied by Data Assimilation (DA) to the model's prior state—from a 10 km Met Office operational forecast model to inform the development of a stochastic parameterization for MCS latent heating. To focus on errors in MCS feedback rather than errors due to a missing MCS, we select analysis increments from 1037 MCS tracks that the model successfully captures at the start of the DA cycle.A Machine Learning–based Gaussian Mixture Model reveals that the vertical structure of temperature analysis increments is probabilistically linked to the atmospheric environment. Bottom-heavy heating increments tend to occur in low Total Column Water Vapor (TCWV) conditions, suggesting that the model underestimates low-level convective heating in relatively dry environments. In contrast, top-heavy heating increments are linked to a moist layer overturning structure—characterized by high TCWV and strong vertical wind shear—indicating model underestimation of upper-level condensate detrainment in such environments. This probabilistic relationship is implemented in the Met Office operational forecast model as part of the MCS: PRIME stochastic scheme, which corrects MCS-related uncertainties during model integration. By enhancing top-heavy heating, the scheme backscatters kinetic energy from the mesoscale to larger scales, improving predictions of Indian seasonal rainfall and the Madden–Julian Oscillation (MJO). Future work will assess its impact on forecast busts and its potential to extend predictability.Precipitation rate, convective diagnostics and spin-up compared across physics suites in the model uncertainty model intercomparison project (MUMIP)
Copernicus Publications (2025)
Abstract:
A parameterisation suite is the combination of all parameterisation schemes that is used by a numerical model of the atmosphere. These parameterisation (or “physics”) suites are widely seen as the most uncertain components of atmospheric models. In MUMIP we compare deterministic parameterisation suites from across different modelling centres under common prescribed large-scale dynamics. In the first MUMIP experiment, these dynamical tendencies have been derived by coarse-graining the convection-permitting ICON DYAMOND simulation to 0.2 degree resolution. We use these realistic spatiotemporal dynamical patterns to drive millions of single column model simulations over the tropical Indian Ocean with prescribed SSTs. We use this data to estimate the uncertainty from their physics across four models, each using their default convection-parametrised physics suites. The models are: IFS, GFS, RAP and ARPEGE. The distributions of precipitation rate, convective available potential energy (CAPE), convective inhibition (CIN) and level of neutral buoyancy are analysed, as well as individual model tendencies and rate of change of CAPE and CIN as a function of lead time and, for instance, the diurnal cycle . We find notable differences across the physics suites and even more strongly between convection-parameterised physics suites and the convection-permitting ICON DYAMOND benchmark. Furthermore, we relate these diagnostics to biases in temperature and specific humidity. We also develop a framework for the detection of statistical relations among diagnostics and/or their change. The framework may for instance be used to quantify the impact of spin-up compared to persistence ("memory") and randomness within a dataset and to identify similarity in the physics across modelling centres. In this contribution some of the early results of the international MUMIP project will be presented and we hope to encourage other researchers to use and/or complement the data of MUMIP. Please refer to https://mumip.web.ox.ac.uk for details of how to get involved.Characterizing uncertainty in deep convection triggering using explainable machine learning
Journal of the Atmospheric Sciences American Meteorological Society 82:6 (2025) 1093-1111