SearcharxivSearch

arXiv subjects

Elizabeth A. Barnes

Publications and source records attributed to Elizabeth A. Barnes.

At least 19 recordsLinked to original sources

Extremes on Rewind: Generating 1,000-Member Ensembles Initialized at a Final Condition

Scenario planning for rare, high-impact events often requires massive ensembles to stochastically sample relevant trajectories. Although autoregressive weather emulators can efficiently generate such ensembles, isolating trajectories of interest requires sifting through petabytes of data, a challenge that grows exponentially with lead time and rarity. In contrast, a non-autoregressive foundation model like Climate in a Bottle video (cBottle-video) can directly sample trajectories terminating in extremes, avoiding large-ensemble search. We use cBottle-video to generate 1000-member ensembles with start- and/or end-conditioning across three extreme events---the 2021 Pacific Northwest (PNW) heatwave, Superstorm Sandy, and Hurricane Ian. Antecedent 500 hPa geopotential height ($z_{500}$) spread at the free end of end-conditioned ensembles reaches 84--89\% of the final-state spread of start-conditioned ensembles, revealing substantial diversity consistent with each extreme event. For the 2021 PNW heatwave, end-conditioned ensemble members begin uniformly warmer than reanalysis and stay warm, replacing the observed rapid intensification with persistent antecedent heat. For Superstorm Sandy, the leading modes of $z_{500}$ at the antecedent end of the end-conditioned ensemble explain 44\% of the variance in track latitude, and roughly 10\% of ensemble members begin as stronger hurricanes than Sandy. For Hurricane Ian, variation in the first landfall location among end-conditioned trajectories underscores the importance of accounting for intermediate hazard exposure in risk planning.

physics.ao-ph

Do AI Forecast Ensembles Sample the Correct Conditional Distribution?

Ensemble forecasting aims to sample the conditional distribution of outcomes; whether AI forecast ensembles do this correctly in a joint sense remains largely untested. We train a diffusion model for probabilistic subseasonal coastal sea level forecasts at eight US East Coast tide gauge stations, with sea level derived from reanalysis, and find that marginal and joint forecast quality decouple: positive skill at every station and lead time marginally, while joint spatial structure is worse than climatological draws. A shuffle-based permutation decomposition reveals this failure is invisible to the energy score but detected by the variogram score. Lorenz-96 experiments across 0.7-170 equivalent years show the gap persists regardless of training volume and is reproduced by a linear baseline, indicating structural inadequacy of the learned distribution. A dynamical ensemble does not replicate the failure while a deterministic emulator does, suggesting it is specific to learned emulators rather than ensemble forecasting generally.

physics.ao-ph

Multi-Year-to-Decadal Temperature Prediction using a Machine Learning Model-Analog Framework

Multi-year-to-decadal climate predictions are a key tool in understanding the range of potential regional climate futures. Here, we present a framework that combines machine learning and analog forecasting for regional predictions on these timescales. A neural network is used to learn a mask of weights that highlights important global precursors to the evolution of a specific prediction target (region, variable, and lead time). A library of mask-weighted model states, or potential analogs, are then compared to a mask-weighted observational state. The known future of the best matching potential analogs serve as the prediction for the future of the observational state. We predict 2-meter temperature using the Berkeley Earth Surface Temperature dataset for observations, with a multi-model potential analog library of CMIP6 simulations. Using a 30-year climatology reference, we compare our analog method to two other analog prediction methods and the CMIP6 library using the continuous ranked probability score to assess the quality of predicted distributions, and mean squared error to assess the quality of predicted ensemble means. For nearly all cases explored, our analog method produces skillful predictions. We find higher distribution skill over the CMIP6 library in all cases, which is due to improved prediction of the distribution spread. We find overall higher skill than other analog methods in the predicted distributions and ensemble means. Finally, we find broadly similar skill to an ensemble of bias-corrected initialized Earth system models. Benefits of our analog method include low computational cost, ensemble size flexibility, and interpretability.

physics.ao-ph

Precipitation diffusion downscaling and application to out-of-distribution simulations with and without stratospheric aerosol injection

Stratospheric aerosol injection (SAI), a possible climate engineering strategy where reflective particles are injected into the stratosphere, has been explored to mitigate global warming and its associated risks, such as the intensification of extreme precipitation events. However, current Earth system models (ESMs) often used to simulate SAI and other climate change scenarios are too coarse to properly assess such risks. Traditional statistical downscaling methods, used to project higher resolution impacts, may be biased and unrealistic. To address this, we train a deep learning diffusion downscaler to generate 0.25° contiguous United States (CONUS) daily precipitation using historical and future climate simulations from the Mesoscale Atmosphere-Ocean Interaction in Seasonal-to-Decadal Climate Prediction (MESACLIP) project, then apply the diffusion downscaler to out-of-distribution CESM2 simulations with and without SAI. The diffusion model generates realistic downscaled precipitation using either MESACLIP or CESM2 inputs. It also faithfully recreates the climate change projections of extreme precipitation in MESACLIP. Diffusion-downscaled projections of the future CESM2 SAI scenarios suggest that SAI could nearly cut in half the CONUS-average increase in yearly max precipitation, compared to the non-SAI scenario. However, there is considerable regional variation and internal variability, with SAI modeled to only slightly reduce increases in extreme precipitation frequency in the Mid Atlantic and the Pacific Northwest, but mitigating most intensification in other regions. Future application of diffusion downscaling to a wider variety of SAI scenarios would provide valuable insight into how proposed SAI strategies may affect precipitation variability on fine spatial scales for regional impact assessments.

physics.ao-ph

AI-informed model-analogs for understanding subseasonal-to-seasonal jet stream and North American temperature predictability

Subseasonal-to-seasonal forecasting is crucial for public health, disaster preparedness, and agriculture, and yet it remains a particularly challenging timescale to predict. We explore the use of an interpretable AI-informed model analog forecasting approach, previously employed on longer timescales, to improve S2S predictions. Using an artificial neural network, we learn a mask of weights to optimize analog selection and showcase its versatility across three varied prediction tasks: 1) classification of Week 3-4 Southern California summer temperatures; 2) regional regression of Month 1 midwestern U.S. summer temperatures; and 3) classification of Month 1-2 North Atlantic wintertime upper atmospheric winds. The AI-informed analogs outperform traditional analog forecasting approaches, as well as climatology and persistence baselines, for deterministic and probabilistic skill metrics on both climate model and reanalysis data. We find the analog ensembles built using the AI-informed approach also produce better predictions of temperature extremes and improve representation of forecast uncertainty. Finally, by using an interpretable-AI framework, we analyze the learned masks of weights to better understand S2S sources of predictability.

physics.ao-ph

Am I Confused or Is This Confusing?: Deep Ensembles for ENSO Uncertainty Quantification

Faithful uncertainty quantification (UQ) is paramount in high stakes climate prediction. Deep ensembles, or ensembles of probabilistic neural networks, are state of the art for UQ in machine learning (ML) and are growing increasingly popular for weather and climate prediction. However, detailed analyses of the mechanisms, strengths, and limitations of ensembles in these complex problem settings are lacking. We take a step towards filling this gap by deploying deep ensembles for predictability analysis of the El-Niño Southern Oscillation (ENSO) in the Community Earth System Model 2 Large Ensemble (CESM2-LE). Principally, we show that epistemic uncertainty, modeled by ensemble disagreement, robustly signals predictive error growth associated with shifts in the distributions of monthly sea-surface temperature (SST), ocean heat content (OHC), and zonal surface wind stress ($τ_x$) anomalies under a climate change scenario. Conversely, we find that aleatoric uncertainty, which remains a popular measure of model confidence, becomes less reliable and behaves counterintuitively under climate-change-induced distributional shift. We highlight that, because ensemble performance improvement relative to the expected single model scales with epistemic uncertainty, ensemble improvement increases with distributional shift from climate change. This work demonstrates the utility of deep ensembles for modeling aleatoric and epistemic uncertainty in ML climate prediction, as well as the growing importance of robustly quantifying these two forms of uncertainty under anthropogenic warming.

physics.ao-ph

Watch an AI Weather Model Learn (and Unlearn) Tropical Cyclones

In a changing climate, artificial intelligence (AI) weather models have the potential to provide cheaper, faster, and more accurate forecasts of high-impact weather events. To realize this potential and gauge trustworthiness, there is a need for more research on how models learn extreme events and how that learning might be improved. Here, we investigate how a Spherical Fourier Neural Operator (SFNO) learns tropical cyclones (TCs) by saving every checkpoint from training and analyzing storm specific metrics. We find evidence that for some storms the SFNO learns information about TC intensity that it loses later in training. This unlearning pattern is associated with anomalously moist environments and may be due to the model unlearning the relationship between moisture and TC intensity. This work provides a first example of leveraging task-specific training dynamics to further our understanding of how AI weather models learn extreme events.

physics.ao-ph

Forecasting the Future with Yesterday's Climate: Temperature Bias in AI Weather and Climate Models

AI-based climate and weather models have rapidly gained popularity, providing faster forecasts with skill that can match or even surpass that of traditional dynamical models. Despite this success, these models face a key challenge: predicting future climates while being trained only with historical data. In this study, we investigate this issue by analyzing boreal winter land temperature biases in AI weather and climate models. We examine two weather models, FourCastNet V2 Small (FourCastNet) and Pangu Weather (Pangu), evaluating their predictions for 2020-2025 and Ai2 Climate Emulator version 2 (ACE2) for 1996-2010. These time periods lie outside of the respective models' training sets and are significantly more recent than the bulk of their training data, allowing us to assess how well the models generalize to new, i.e. more modern, conditions. We find that all three models produce cold-biased mean temperatures, resembling climates from 15-20 years earlier than the period they are predicting. In some regions, like the Eastern U.S., the predictions resemble climates from as much as 20-30 years earlier. Further analysis shows that FourCastNet's and Pangu's cold bias is strongest in the hottest predicted temperatures, indicating limited training exposure to modern extreme heat events. In contrast, ACE2's bias is more evenly distributed but largest in regions, seasons, and parts of the temperature distribution where climate change has been most pronounced. These findings underscore the challenge of training AI models exclusively on historical data and highlight the need to account for such biases when applying them to future climate prediction.

physics.ao-ph

Digestible Pieces: comparing three options for partitioning the Northeast Pacific Coast for S2S sea surface height prediction

We discuss the utility of applying clustering as a preprocessing step for identifying subseasonal to seasonal forecasts of opportunity of coastal sea level using convolutional neural networks (CNNs). Clustering leverages potential covariance among points along the same coastline or in the same ocean basin. To evaluate the utility of clustering for reliably identifying forecasts of opportunity, we compare CNNs trained to predict sea level probability distributions in three ways: over the whole Northeast Pacific Coast simultaneously, over predetermined clusters within this coastline, and at individual gridpoints near tide gauges. All CNN prediction tasks (Whole Coast, Cluster, Point), outperform climatology by a similar margin at Week 3 when the entire test set is used to evaluate CNN skill. However, when comparing the skill of each tasks' 20% most confident predictions, we find the skill of the Cluster and Point tasks to be on par with each other and substantially more skillful than the Whole Coast task. Of the Cluster and Point task, the Cluster task represents all gridpoints in the Northeast Pacific Coast with minimal tunable parameters. Throughout this exercise we learned that clustering gridpoints as a pre-processing step is the preferred approach between the three for making S2S predictions of coastal sea level.

physics.ao-ph

How does an AI Weather Model Learn to Forecast Extreme Weather?

In a warming climate with more frequent severe weather, artificial intelligence (AI) weather models have the potential to provide cheaper, faster, and more accurate forecasts of high-impact weather events. To realize this potential, there is a need for more research on how models learn extreme events and how that learning might be improved. We investigate how a spherical Fourier neural operator model (SFNO) learns extreme weather by saving every checkpoint throughout training and analyzing a collection of 9 extreme weather events including heatwaves, atmospheric rivers, and tropical cyclones. The SFNO learns heatwaves similarly to other weather days, but we find evidence that the model learns information about atmospheric river and tropical cyclone forecasts that it loses later in training. We propose a possible training strategy to improve the forecasting of extreme events by retaining information from earlier training checkpoints, and provide initial evidence of its utility.

physics.ao-ph

Uncertainty-permitting machine learning reveals sources of dynamic sea level predictability across daily-to-seasonal timescales

Reliable dynamic sea level forecasts are hindered by numerous sources of uncertainty on daily-to-seasonal timescales (1-180 days) due to atmospheric boundary conditions and internal ocean variability. Studies have demonstrated that certain initial states can extend predictability horizons; thus, identifying these initial conditions may help improve forecast skill. Here, we identify sources of dynamic sea level predictability on daily-to-seasonal timescales using neural networks trained on CESM2 large ensemble data to forecast dynamic sea level. The forecasts yield not only a point estimate for sea level but also a standard deviation to quantify forecast uncertainty based on the initial conditions. Forecasted uncertainties can be leveraged to identify state-dependent sources of predictability at most locations and forecast leads. Network forecasts, particularly in the low-latitude Indo-Pacific, exhibit skillful deterministic predictions and skillfully forecast exceedance probabilities relative to local linear baselines. For networks trained at Guam and in the western Indian Ocean, the transfer of sources of predictability from local sources to remote sources is presented by the deteriorating utility of initial condition information for predicting exceedance events. Propagating Rossby waves are identified as a potential source of predictability for dynamic sea level at Guam. In the Indian Ocean, persistence of thermosteric sea level anomalies from the Indian Ocean Dipole may be a source of predictability on subseasonal timescales, but El Niño drives predictability on seasonal timescales. This work shows how uncertainty-quantifying machine learning can help identify changes in sources of state-dependent predictability over a range of forecast leads.

physics.ao-ph

Reanalysis-based Global Radiative Response to Sea Surface Temperature Patterns: Evaluating the Ai2 Climate Emulator

The sensitivity of the radiative flux at the top of the atmosphere to surface temperature perturbations cannot be directly observed. The relationship between sea surface temperature (SST) and top-of-atmosphere radiation can be estimated with Green's function simulations by locally perturbing the sea surface temperature boundary conditions in atmospheric climate models. We perform such simulations with the Ai2 Climate Emulator (ACE), a machine learning-based emulator trained on ERA5 reanalysis data (ACE2-ERA5). This produces a sensitivity map of the top-of-atmosphere radiative response to surface warming that aligns with our physical understanding of radiative feedbacks. However, ACE2-ERA5 likely underestimates the radiative response to historical warming. We compare to two additional versions of ACE and traditional climate models. We argue that Green's function experiments can be used to evaluate the performance and limitations of machine learning-based climate emulators by examining if causal physical relationships are correctly represented and testing their capability for out-of-distribution predictions.

physics.ao-ph

Turning Up the Heat: Assessing 2-m Temperature Forecast Errors in AI Weather Prediction Models During Heat Waves

Extreme heat is the deadliest weather-related hazard in the United States. Furthermore, it is increasing in intensity, frequency, and duration, making skillful forecasts vital to protecting life and property. Traditional numerical weather prediction (NWP) models struggle with extreme heat for medium-range and subseasonal-to-seasonal (S2S) timescales. Meanwhile, artificial intelligence-based weather prediction (AIWP) models are progressing rapidly. However, it is largely unknown how well AIWP models forecast extremes, especially for medium-range and S2S timescales. This study investigates 2-m temperature forecasts for 60 heat waves across the four boreal seasons and over four CONUS regions at lead times up to 20 days, using two AIWP models (Google GraphCast and Pangu-Weather) and one traditional NWP model (NOAA United Forecast System Global Ensemble Forecast System (UFS GEFS)). First, case study analyses show that both AIWP models and the UFS GEFS exhibit consistent cold biases on regional scales in the 5-10 days of lead time before heat wave onset. GraphCast is the more skillful AIWP model, outperforming UFS GEFS and Pangu-Weather in most locations. Next, the two AIWP models are isolated and analyzed across all heat waves and seasons, with events split among the model's testing (2018-2023) and training (1979-2017) periods. There are cold biases before and during the heat waves in both models and all seasons, except Pangu-Weather in winter, which exhibits a mean warm bias before heat wave onset. Overall, results offer encouragement that AIWP models may be useful for medium-range and S2S predictability of extreme heat.

physics.ao-ph

Predicting Tropical Cyclone Track Forecast Errors using a Probabilistic Neural Network

A new method for estimating tropical cyclone track uncertainty is presented and tested. This method uses a neural network to predict a bivariate normal distribution, which serves as an estimate for track uncertainty. We train the network and make predictions on forecasts from the National Hurricane Center (NHC), which currently uses static error distributions based on forecasts from the past five years for most applications. The neural network-based method produces uncertainty estimates that are dynamic and probabilistic. Further, the neural network-based method allows for probabilistic statements about tropical cyclone trajectories, including landfall probability, which we highlight. We show that our predictions are well calibrated using multiple metrics, that our method produces better uncertainty estimates than current NHC approaches, and that our method achieves similar performance to the Global Ensemble Forecast System. Once trained, the computational cost of predictions using this method is negligible, making it a strong candidate to improve the NHC's operational estimations of tropical cyclone track uncertainty.

physics.ao-ph

Recommendations for Comprehensive and Independent Evaluation of Machine Learning-Based Earth System Models

Machine learning (ML) is a revolutionary technology with demonstrable applications across multiple disciplines. Within the Earth science community, ML has been most visible for weather forecasting, producing forecasts that rival modern physics-based models. Given the importance of deepening our understanding and improving predictions of the Earth system on all time scales, efforts are now underway to develop forecasting models into Earth-system models (ESMs), capable of representing all components of the coupled Earth system (or their aggregated behavior) and their response to external changes. Modeling the Earth system is a much more difficult problem than weather forecasting, not least because the model must represent the alternate (e.g., future) coupled states of the system for which there are no historical observations. Given that the physical principles that enable predictions about the response of the Earth system are often not explicitly coded in these ML-based models, demonstrating the credibility of ML-based ESMs thus requires us to build evidence of their consistency with the physical system. To this end, this paper puts forward five recommendations to enhance comprehensive, standardized, and independent evaluation of ML-based ESMs to strengthen their credibility and promote their wider use.

cs.LG

Using Neural Networks to Learn the Jet Stream Forced Response from Natural Variability

Two distinct features of anthropogenic climate change, warming in the tropical upper troposphere and warming at the Arctic surface, have competing effects on the mid-latitude jet stream's latitudinal position, often referred to as a "tug-of-war". Studies that investigate the jet's response to these thermal forcings show that it is sensitive to model type, season, initial atmospheric conditions, and the shape and magnitude of the forcing. Much of this past work focuses on studying a simulation's response to external manipulation. In contrast, we explore the potential to train a convolutional neural network (CNN) on internal variability alone and then use it to examine possible nonlinear responses of the jet to tropospheric thermal forcing that more closely resemble anthropogenic climate change. Our approach leverages the idea behind the fluctuation-dissipation theorem, which relates the internal variability of a system to its forced response but so far has been only used to quantify linear responses. We train a CNN on data from a long control run of the CESM dry dynamical core and show that it is able to skillfully predict the nonlinear response of the jet to sustained external forcing. The trained CNN provides a quick method for exploring the jet stream sensitivity to a wide range of tropospheric temperature tendencies and, considering that this method can likely be applied to any model with a long control run, could lend itself useful for early stage experiment design.

physics.ao-ph

Investigating the fidelity of explainable artificial intelligence methods for applications of convolutional neural networks in geoscience

Convolutional neural networks (CNNs) have recently attracted great attention in geoscience due to their ability to capture non-linear system behavior and extract predictive spatiotemporal patterns. Given their black-box nature however, and the importance of prediction explainability, methods of explainable artificial intelligence (XAI) are gaining popularity as a means to explain the CNN decision-making strategy. Here, we establish an intercomparison of some of the most popular XAI methods and investigate their fidelity in explaining CNN decisions for geoscientific applications. Our goal is to raise awareness of the theoretical limitations of these methods and gain insight into the relative strengths and weaknesses to help guide best practices. The considered XAI methods are first applied to an idealized attribution benchmark, where the ground truth of explanation of the network is known a priori, to help objectively assess their performance. Secondly, we apply XAI to a climate-related prediction setting, namely to explain a CNN that is trained to predict the number of atmospheric rivers in daily snapshots of climate simulations. Our results highlight several important issues of XAI methods (e.g., gradient shattering, inability to distinguish the sign of attribution, ignorance to zero input) that have previously been overlooked in our field and, if not considered cautiously, may lead to a distorted picture of the CNN decision-making strategy. We envision that our analysis will motivate further investigation into XAI fidelity and will help towards a cautious implementation of XAI in geoscience, which can lead to further exploitation of CNNs and deep learning for prediction problems.

physics.geo-ph

Carefully choose the baseline: Lessons learned from applying XAI attribution methods for regression tasks in geoscience

Methods of eXplainable Artificial Intelligence (XAI) are used in geoscientific applications to gain insights into the decision-making strategy of Neural Networks (NNs) highlighting which features in the input contribute the most to a NN prediction. Here, we discuss our lesson learned that the task of attributing a prediction to the input does not have a single solution. Instead, the attribution results and their interpretation depend greatly on the considered baseline (sometimes referred to as reference point) that the XAI method utilizes; a fact that has been overlooked so far in the literature. This baseline can be chosen by the user or it is set by construction in the method s algorithm, often without the user being aware of that choice. We highlight that different baselines can lead to different insights for different science questions and, thus, should be chosen accordingly. To illustrate the impact of the baseline, we use a large ensemble of historical and future climate simulations forced with the SSP3-7.0 scenario and train a fully connected NN to predict the ensemble- and global-mean temperature (i.e., the forced global warming signal) given an annual temperature map from an individual ensemble member. We then use various XAI methods and different baselines to attribute the network predictions to the input. We show that attributions differ substantially when considering different baselines, as they correspond to answering different science questions. We conclude by discussing some important implications and considerations about the use of baselines in XAI research.

physics.geo-ph