SearcharxivSearch

arXiv subjects

Hannah M. Christensen

Publications and source records attributed to Hannah M. Christensen.

16 recordsLinked to original sources

How Do AI Climate Models Respond to Warming Across Climate Zones?

Regional climate zones are expected to shift under global warming. Whether AI climate models have learned to generalize climate-zone distributions under warming in a physically meaningful way affects their suitability for climate projection. We address this question by applying a Köppen-Geiger climate-zone decomposition to AIMIP Phase 1 models under prescribed +4K SST forcing and comparing their responses to physics-based AMIP models. Using this diagnostic, we compare baseline classification skill, per-zone responses in temperature, precipitation, and near-surface specific humidity, and the spatial structure of departures from physics-based models. All AI models considered reproduce the 1979-2014 ERA5 climatology within the physics-based models' range, but only the hybrid physics-AI model NeuralGCM-HRD reorganizes zones in agreement with established thermodynamic and hydrological scaling relations. The remaining emulators have distinct failure modes traceable to their architectural treatment of land cells. A physically consistent climate-zone response is therefore necessary for AI models intended for climate projection.

physics.ao-ph

Spatial Generalization Tests for Machine Learning-based Weather Models to Assess Physical Consistency

Machine learning-based weather prediction is revolutionizing weather forecasting by learning from weather data in present-day climate. However, generalization to other climates remains a major challenge. With melting sea ice, land-use change, and increasing ocean temperatures, boundary conditions are changing. Therefore, generalization in time depends on generalization in space. Here, we present three test cases to evaluate whether machine learning-based weather and climate models generalize in space and apply them to GraphCast and NeuralGCM. We reverse or rotate the planet in longitude or latitude under the model's coordinate system and adapt all boundary conditions and forcings accordingly. Physics-based general circulation models simulate a rotated/reversed planet with only rounding errors, but GraphCast and NeuralGCM fail these tests. The analyses furthermore revealed unphysical variable mappings based on correlation rather than causation. We argue that machine learning-based climate models should be designed to pass generalization tests to prevent overfitting on present-day regional climate.

physics.ao-ph

An observationally constrained probabilistic trigger for organized deep convection in an NWP ensemble

A novel stochastic parametrization scheme representing organized convection is described. The effects of mesoscale convective systems (MCSs) are represented in an observationally constrained manner, by probabilistically triggering an MCS scheme in regions of enhanced environmental total column water vapour. In combination with the probabilistic trigger, patterns with given spatiotemporal scales determine where and when the scheme is active. Our scheme builds on the multiscale coherent structure parametrization (MCSP), which represents the top-heavy heating structure associated with MCSs. The original and new MCSP schemes are tested in a numerical weather prediction (NWP) ensemble. Both MCSP schemes improve the spatiotemporal scales of tropical precipitation compared to a control. When the spread-error relationship of tropical precipitation is analysed, the new scheme successfully boosts spread compared to original MCSP, improving the underdispersion of the ensemble seen with the original MCSP.

physics.ao-ph

ACE2-NEMO: Coupling an ML atmospheric emulator to a full-depth dynamical ocean model

Understanding how fast atmospheric variability shapes slow climate variability and sensitivity remains a central challenge in Earth-system science. Recent advances in machine-learned (ML) atmospheric models have demonstrated remarkable skill on weather timescales, but their emergent behaviour in a fully coupled climate system remains largely unexplored. We present early results from a new hybrid modelling framework, in which the ACE2 ML atmospheric emulator is interactively coupled to the NEMO ocean model. We report on a set of 70-year coupled simulations (1950-2020 historical forcing and fixed-1950s control). These experiments represent, to our knowledge, the first multi-decadal integrations of a machine-learned atmosphere interacting with a full-depth dynamical ocean. Several historical and fixed-1950s control simulations from the fully dynamic global coupled climate model EC-Earth, which has the same ocean component used in ACE2-NEMO, are also considered for comparison. We assess the behaviour of the coupled system, with particular focus on low-frequency tropical variability and the climate response to greenhouse-gas forcing. Analysis of potentially emergent El Niño-like variability reveals realistic fast timescale air-sea coupling in the tropical Pacific, but the temporal variability is unrealistic, with very low amplitude oscillations; this appears to be due to weak atmospheric feedback in the tropical Pacific. The response to CO2 forcing shows initial agreement with EC-Earth3P, but deviates due to reduced downward short-wave radiation in ACE2. These results provide a unique test of physical realism for atmospheric emulators, and evaluate the possible role of entirely machine-learned components in next-generation Earth system models.

physics.ao-ph

No Epoch Like the Present: Robust Climate Emulation Requires Out-of-Distribution Generalisation

Climate emulation is an out-of-distribution (OOD) projection task. This is precisely the challenge where modern Machine Learning (ML) methods are most prone to failure. Consequently, while current ML emulators trained on present climate achieve high in-distribution performance, their future reliability under the inevitable distribution shifts of a changing climate remains a critical, poorly understood blind spot. Addressing this challenge requires a fundamental shift in how we understand, evaluate, and design climate emulators. In this work, we first confirm that climate change drives a statistically significant and progressively growing shift in atmospheric state distributions, rendering standard evaluation protocols insufficient. We empirically establish that seasonal variation serves as an effective proxy for these long-term climate shifts, providing access to $\textit{real-world}$ distribution shifts without recourse to heuristics like synthetic perturbations. Motivated by this link, we introduce a novel evaluation framework that leverages seasonal shifts as a rigorous, zero-overhead testbed for emulator robustness. Our systematic characterisation confirms that current state-of-the-art hybrid-ML emulators degrade significantly under these realistic shifts. Finally, we chart a path forward by identifying compositional generalisation, the ability to form novel combinations from observed elementary components, as a principled route towards robust climate emulation. We demonstrate that physically motivated decompositions substantially improve OOD performance with only modest trade-offs against in-distribution performance, providing an avenue towards ML-driven climate emulators robust to an unknown future.

cs.LG

Role of the ocean for fast atmospheric evolution revealed by machine learning

There have recently been many efforts to create machine learnt atmospheric emulators designed to replace physical models. So far these have mainly focused on medium-range weather forecasting, where these `Machine Learnt Weather Prediction' (MLWP) models can outperform leading operational forecasting centres. However, because of this focus on shorter timescales, many of these emulators ignore the effects of the ocean, and take no ocean variables as inputs. We hypothesise that such MLWP models have learnt a best-guess of the evolution of the atmosphere, by implicitly inferring ocean conditions from atmospheric states, with no access to ocean data. Turning this limitation into a strength, we use it as a means to study the role of the oceans on the evolution of the atmosphere. By exploring how model forecast errors relate to properties of the air-sea interface, we infer what ocean information these atmospheric emulators are able to derive from atmospheric data alone, and what they cannot. This highlights the regions and processes through which the ocean independently influences the atmosphere on fast timescales. We perform this analysis for GraphCast, finding clear relationships between air-sea properties and the forecast errors over the ocean, including clear seasonal effects. We then explore what this reveals about GraphCast's internal representation of the ocean. In addition to understanding real-world ocean-atmosphere interactions, this analysis provides guidance for improving forecast skill and physical realism in MLWP models, and for informing how future machine learning models should use ocean information on short timescales.

physics.ao-ph

Error in ERA5 2m Temperature identified using GraphCast

Reanalyses such as ERA5 have long been foundational for weather and climate science. They have also found a new use case, as training and verification data for machine-learnt weather prediction (MLWP) models. Here we compare short-lead time (6h) forecasts from the MLWP model GraphCast against ERA5. In doing so, we identify a recurrent, spatially coherent error in 2m Temperature centred on the Ethiopian Highlands, that occurs predominantly at 0600 UTC. We show that these error events are not an error in the forecast from GraphCast, but are in fact an error in ERA5, and are also present in the ECMWF operational analysis. They arise from the 2D optimal interpolation procedure, when surface reports are assimilated that are temporally displaced compared to the background forecast. This produces spuriously warm analysis increments over Ethiopia on approximately 7\% of dates at 0600 UTC across the reanalysis record. The spread from the ensemble of data assimilation partially flags these cases but is underdispersive. We assess the impact on GraphCast, which was trained on ERA5. While GraphCast can largely ignore these unphysical error events, a small systematic degradation in forecast skill over the region is observed. We discuss implications for using reanalysis as truth in machine learning training and verification, and recommend simple changes to reduce such artefacts in future analyses.

physics.ao-ph

Epistemic and Aleatoric Uncertainty Quantification in Weather and Climate Models

Representing and quantifying uncertainty in physical parameterisations is a central challenge in weather and climate modelling, and approaches are often developed separately for different timescales. Here, we introduce a unified framework for analysing uncertainty in parameterisations across weather and climate regimes. Using the Lorenz 1996 system as a testbed for simplified chaotic dynamics, we quantify uncertainties in a subgrid-scale parameterisation using a Bayesian Neural Network (BNN). This allows us to disentangle aleatoric uncertainty, arising from internal variability in the training data, and epistemic uncertainties, arising from poorly constrained parameters during training. At runtime, we sample uncertainties in line with stochastic approaches in weather models and perturbed-parameter methods in climate models. On weather timescales, aleatoric uncertainty dominates, underscoring the value of stochastic parameterisations. On longer, climate timescales and under changing forcings, accounting for both types of uncertainty is necessary for well-calibrated ensembles, with epistemic uncertainty widening the range of explored climate states, and aleatoric uncertainty promoting transitions between them. Constraining parameter uncertainty with short simulations reduces epistemic uncertainty and improves long-term model behaviour under perturbed forcings. This framework links concepts from machine learning with traditional uncertainty quantification in Earth system modelling, offering a pathway toward seamless treatment of uncertainty in weather and climate prediction.

physics.ao-ph

Seasonal forecasting using the GenCast probabilistic machine learning model

Machine-learnt weather prediction (MLWP) models are now well established as being competitive with conventional numerical weather prediction (NWP) models in the medium range. However, there is still much uncertainty as to how this performance extends to longer timescales, where interactions with slower components of the earth system become important. We take GenCast, a state-of-the-art probabilistic MLWP model, and apply it to the task of seasonal forecasting with prescribed sea surface temperature (SST), by providing anomalies persisted over climatology (GenCast-Persisted) or forcing with observations (GenCast-Forced). The forecasts are compared to the European Centre for Medium-Range Weather Forecasts seasonal forecasting system, SEAS5. Our results indicate that, despite being trained at short timescales, GenCast-Persisted produces much of the correct precipitation patterns in response to El Niño and La Niña events, with several erroneous patterns in GenCast-Persisted corrected with GenCast-Forced. The uncertainty in precipitation response, as represented by the ensemble, compares favourably to SEAS5. Whilst SEAS5 achieves superior skill in the tropics for 2-metre temperature and mean sea level pressure (MSLP), GenCast-Persisted achieves significantly higher skill in some areas in higher latitudes, including mountainous areas, with notable improvements for MSLP in particular; this is reflected in a higher correlation with the observed NAO index. Reliability diagrams indicate that GenCast-Persisted is overconfident compared to SEAS5, whilst GenCast-Forced produces well-calibrated seasonal 2-metre temperature predictions. These results provide an indication of the potential of MLWP models similar to GenCast for the `full' seasonal forecasting problem, where the atmospheric model is coupled to ocean, land and cryosphere models.

physics.ao-ph

The Link between Gulf Stream Precipitation and European Blocking in General Circulation Models and the Role of Horizontal Resolution

Past studies show that coupled model biases in European blocking and North Atlantic eddy-driven jet variability decrease as one increases the horizontal resolution in the atmospheric and oceanic model components. This has commonly been argued to be related to an alleviation of sea surface temperature (SST) biases due to increased oceanic resolution in particular, with a physical pathway via changes to surface baroclinicity. On the other hand, many studies have now highlighted the key role of diabatic processes in the Gulf Stream region on blocking formation and maintenance. Here, following recent work by Schemm, we leverage a large multi-model ensemble to show that Gulf Stream precipitation variability in coupled models is tightly linked to the simulated frequency of European blocking and northern jet excursions. Furthermore, the reduced biases in blocking and jet variability are consistent with greater precipitation variability as a result of increased atmospheric horizontal resolution. By contrast, typical North Atlantic SST biases are found to share only a weak or negligible relationship with blocking and jet biases. Finally, while previous studies have used a comparison between coupled models and models run with prescribed SSTs to argue for the role of ocean resolution, we emphasise here that models run with prescribed SSTs experience greatly reduced precipitation variability due to their excessive thermal damping, making it unclear if such a comparison is meaningful. Instead, we speculate that most of the reduction in coupled model biases may actually be due to increased atmospheric resolution.

physics.ao-ph

Defining error accumulation in ML atmospheric simulators

Machine learning (ML) has recently shown significant promise in modelling atmospheric systems, such as the weather. Many of these ML models are autoregressive, and error accumulation in their forecasts is a key problem. However, there is no clear definition of what `error accumulation' actually entails. In this paper, we propose a definition and an associated metric to measure it. Our definition distinguishes between errors which are due to model deficiencies, which we may hope to fix, and those due to the intrinsic properties of atmospheric systems (chaos, unobserved variables), which are not fixable. We illustrate the usefulness of this definition by proposing a simple regularization loss penalty inspired by it. This approach shows performance improvements (according to RMSE and spread/skill) in a selection of atmospheric systems, including the real-world weather prediction task.

cs.LG

Machine Learning for Stochastic Parametrisation

Atmospheric models used for weather and climate prediction are traditionally formulated in a deterministic manner. In other words, given a particular state of the resolved scale variables, the most likely forcing from the sub-grid scale processes is estimated and used to predict the evolution of the large-scale flow. However, the lack of scale-separation in the atmosphere means that this approach is a large source of error in forecasts. Over recent years, an alternative paradigm has developed: the use of stochastic techniques to characterise uncertainty in small-scale processes. These techniques are now widely used across weather, sub-seasonal, seasonal, and climate timescales. In parallel, recent years have also seen significant progress in replacing parametrisation schemes using machine learning (ML). This has the potential to both speed up and improve our numerical models. However, the focus to date has largely been on deterministic approaches. In this position paper, we bring together these two key developments, and discuss the potential for data-driven approaches for stochastic parametrisation. We highlight early studies in this area, and draw attention to the novel challenges that remain.

cs.LG

Using Probabilistic Machine Learning to Better Model Temporal Patterns in Parameterizations: a case study with the Lorenz 96 model

The modelling of small-scale processes is a major source of error in climate models, hindering the accuracy of low-cost models which must approximate such processes through parameterization. Red noise is essential to many operational parameterization schemes, helping model temporal correlations. We show how to build on the successes of red noise by combining the known benefits of stochasticity with machine learning. This is done using a physically-informed recurrent neural network within a probabilistic framework. Our model is competitive and often superior to both a bespoke baseline and an existing probabilistic machine learning approach (GAN) when applied to the Lorenz 96 atmospheric simulation. This is due to its superior ability to model temporal patterns compared to standard first-order autoregressive schemes. It also generalises to unseen scenarios. We evaluate across a number of metrics from the literature, and also discuss the benefits of using the probabilistic metric of hold-out likelihood.

cs.LG

The Fractal Nature of Clouds in Global Storm-Resolving Models

Clouds in observations are fractals: they show self-similarity across scales ranging from one to 1000 km. This includes individual storms and large-scale cloud structures typical of organised convection. It is not known whether global storm-resolving models reproduce the observed fractal scaling laws for clouds and organised convection. We compute the fractal dimension of clouds using Himawari satellite data and compare this to global storm-resolving model simulations completed as part of the DYAMOND intercomparison project. We find cloud fields in these simulations are indeed fractal, and reproduce the observed fractal dimension to within 10\%. We find the fractal dimension is sensitive to the choice of boundary layer parametrisation scheme used in each model simulation, and not to the convection parametrisation as might have been expected. The fractal dimension is independent of cloud area distributions, providing a complementary metric to assess the multi-scale structure of convection and convective organisation in model simulations.

physics.ao-ph

Machine Learning for Stochastic Parameterization: Generative Adversarial Networks in the Lorenz '96 Model

Stochastic parameterizations account for uncertainty in the representation of unresolved sub-grid processes by sampling from the distribution of possible sub-grid forcings. Some existing stochastic parameterizations utilize data-driven approaches to characterize uncertainty, but these approaches require significant structural assumptions that can limit their scalability. Machine learning models, including neural networks, are able to represent a wide range of distributions and build optimized mappings between a large number of inputs and sub-grid forcings. Recent research on machine learning parameterizations has focused only on deterministic parameterizations. In this study, we develop a stochastic parameterization using the generative adversarial network (GAN) machine learning framework. The GAN stochastic parameterization is trained and evaluated on output from the Lorenz '96 model, which is a common baseline model for evaluating both parameterization and data assimilation techniques. We evaluate different ways of characterizing the input noise for the model and perform model runs with the GAN parameterization at weather and climate timescales. Some of the GAN configurations perform better than a baseline bespoke parameterization at both timescales, and the networks closely reproduce the spatio-temporal correlations and regimes of the Lorenz '96 system. We also find that in general those models which produce skillful forecasts are also associated with the best climate simulations.

physics.ao-ph

Constraining stochastic parametrisation schemes using high-resolution simulations

Stochastic parametrisations are used in weather and climate models to improve the representation of unpredictable unresolved processes. When compared to a deterministic model, a stochastic model represents `model uncertainty', i.e., sources of error in the forecast due to the limitations of the forecast model. We present a technique for systematically deriving new stochastic parametrisations or for constraining existing stochastic approaches. A high-resolution model simulation is coarse-grained to the desired forecast model resolution. This provides the initial conditions and forcing data needed to drive a Single Column Model (SCM). By comparing the SCM parametrised tendencies with the evolution of the high resolution model, we can estimate the error in the SCM tendencies that a stochastic parametrisation seeks to represent. We use this approach to assess the physical basis of the widely used Stochastically Perturbed Parametrisation Tendencies (SPPT) scheme. We find justification for the multiplicative nature of SPPT, and for the use of spatio-temporally correlated stochastic perturbations. We find evidence that the stochastic perturbation should be positively skewed, indicating that occasional large-magnitude positive perturbations are physically realistic. However other key assumptions of SPPT are less well justified, including coherency of the stochastic perturbations with height, coherency of the perturbations for different physical parametrisation schemes, and coherency for different prognostic variables. Relaxing these SPPT assumptions allows for an error model that explains a larger fractional variance than traditional SPPT. In particular, we suggest that independently perturbing the tendencies associated with different parametrisation schemes is justifiable, and would improve the realism of the SPPT approach.

physics.ao-ph