SearcharxivSearch

arXiv subjects

Tom Beucler

Publications and source records attributed to Tom Beucler.

At least 19 recordsLinked to original sources

Stress-Testing Dynamical and Generative Downscaling Using Subseasonal Extreme Precipitation Forecasts

Coarse spatial resolution limits the ability of subseasonal prediction models to resolve extreme precipitation. Downscaling with either dynamical or deep generative models can overcome this issue, but the comparative performance of these models for extremes across different atmospheric regimes remains poorly understood. In this work, we evaluate the Weather Research and Forecasting (WRF) model against a diffusion-based generative model by downscaling two physically distinct, extreme precipitation events up to lead times of 3 weeks. For a fair comparison with WRF, which can downscale boundary conditions from different driving models without model-specific training, the diffusion model is trained in an unpaired fashion. Both approaches improve upon the raw European Centre for Medium-Range Weather Forecasts forecasts, in comparison to fused rain gauge-radar observations in Switzerland (CombiPrecip), but exhibit regime-dependent strengths. WRF achieves the highest probabilistic skill for a multicell, non-stationary event. Conversely, the diffusion model is more consistent across different performance metrics for the two events, outperforming WRF in a more stationary supercell event. These results demonstrate that explicit dynamical modeling can add value for specific precipitation events for subseasonal lead times, and that generative downscaling adds value more broadly in different situations.

physics.comp-ph

Maximum updraft velocity beyond CAPE: the role of boundary layer dynamics and pressure perturbations

Deep convective updraft velocities play a key role in the Earth's climate system, influencing precipitation extremes, lightning, and the planetary energy budget. While Convective Available Potential Energy (CAPE) is widely used to explain maximum updraft velocity ($w_{\max}$), CAPE is an imperfect predictor as updrafts are also influenced by entrainment, boundary layer dynamics, pressure perturbations, and condensate loading. However, the relative importance of these processes and how they interact to set $w_{\max}$ in individual clouds remains unclear. Here, we use equation learning to identify compact, physically interpretable relationships linking environmental and in-cloud conditions to $w_{\max}$ in individual tracked clouds across idealized radiative-convective equilibrium regimes spanning a range of sea surface temperatures and radiative cooling rates. For pre-storm prediction, CAPE and local mean boundary layer vertical velocity ($\overline{w_{\mathrm{bl}}}$) together explain nearly half the variance in $w_{\max}$ across regimes ($R^2=0.47$). While CAPE captures regime-mean differences, it has little predictive value within a single simulation. $\overline{w_{\mathrm{bl}}}$ is essential for capturing cloud-to-cloud variability, including the suppression of $w_{\max}$ even at high CAPE values. At the time of peak intensity, a simple approximate Bernoulli-like invariant combining maximum pressure perturbation and maximum cloud condensate explains 89\% of the variance ($R^2=0.89$). The tight link between $w_{\max}$ and pressure perturbation supports the sticky thermals hypothesis and highlights the importance of dynamic pressure effects, often neglected in updraft theories. These results highlight $\overline{w_{\mathrm{bl}}}$ as an important regulator of convective intensity alongside CAPE, and demonstrate that dynamic pressure plays an important role within individual updrafts.

physics.ao-ph

PRecover 1.0: Process Rate Recovery with Machine Learning

Comprehensive information on cloud microphysical process rates from numerical simulations allows for better understanding of precipitation formation pathways and aerosol-cloud interactions. However, resource limitations often make it impractical to include all microphysical process rates in the model output, limiting in-depth analyses. To address this shortcoming, we introduce PRecover, a data-driven post-processing approach to recover microphysical process rates that are not stored during runtime from standard output of a numerical weather prediction model. In particular, we train random forests, gradient boosting models, and feed-forward neural networks to recover microphysical process rates from a two-moment bulk microphysics scheme in the ICOsahedral Nonhydrostatic (ICON) model. We use cloud variables as input, obtained from high-resolution simulations in a limited-area setup over Europe. Warm-rain and ice microphysical process rates are recovered with a two-step classification-regression approach for both instantaneous and accumulated process rates. As a physics-based baseline, we assess whether process rates can be directly recalculated from stored ICON output variables. Accurate recalculation is possible for process rates such as accretion and self-collection but not for the autoconversion, rain melting or heterogeneous ice nucleation rate. Using PRecover, we successfully recover most of the process rates that are accumulated over output time steps of 10 minutes or less, but the values are increasingly difficult to recover for rates accumulated over longer accumulation intervals. To quantify predictive uncertainty, we provide calibrated prediction intervals through conformalized quantile regression. We demonstrate spatial transferability of the models with two case studies over different regional domains and simulation settings unseen during training.

physics.ao-ph

SwAIther-Precip: Lead-Time-Aware Bias Correction Enables Kilometer-Scale Downscaling of Global AI Precipitation Forecasts over Switzerland

Skillful medium-range precipitation forecasting at kilometer scale remains challenging over complex terrain because precipitation arises from multiscale nonlinear processes that global models cannot explicitly resolve at affordable cost. Global AI weather models can produce skillful medium-range forecasts, but their native 0.25 degrees resolution limits direct use for local hazard applications. Statistical downscaling can help bridge this gap, yet existing approaches often struggle with state-dependent, and especially lead-time-dependent, biases in global forecasts. We introduce SwAIther-Precip, a lead-time-aware downscaling framework that converts coarse-resolution AIFS forecasts into probabilistic km-scale precipitation fields over Switzerland. First, a U-Net conditioned on lead time via feature-wise linear modulation deterministically corrects systematic biases at coarse resolution. This targeted correction enables a cheaper super-resolution stage conditioned only on corrected precipitation, allowing direct training on observations rather than on the full atmospheric state. A diffusion-based model then generates fine-scale spatial variability independently of lead time. Using AIFS forecasts and CombiPrecip radar-gauge observations, SwAIther-Precip reduces CRPS by 48% relative to raw AIFS. The generated fields reproduce observed spatial variability with spectral fidelity above 0.85 at large scales and 0.88 at small scales, corresponding to an effective resolution of approximately 4 km on a 1 km grid for lead times up to 5 days. Training across lead times further improves long-range performance, yielding a 13% CRPS reduction at 6 days relative to lead-time-specific models. These results show that explicitly correcting lead-time-dependent biases before generative super-resolution is key to efficient km-scale probabilistic downscaling of global AI precipitation forecasts.

physics.ao-ph

A Scale-Adaptive Framework for Joint Spatiotemporal Super-Resolution with Diffusion Models

Deep-learning video super-resolution has progressed rapidly, but climate applications typically super-resolve (increase resolution) either space or time, and joint spatiotemporal models are often designed for a single pair of super-resolution (SR) factors (upscaling spatial and temporal ratio between the low-resolution sequence and the high-resolution sequence), limiting transfer across spatial resolutions and temporal cadences (frame rates). We present a scale-adaptive framework that reuses the same architecture across factors by decomposing spatiotemporal SR into a deterministic prediction of the conditional mean, with attention, and a residual conditional diffusion model, with an optional mass-conservation (same precipitation amount in inputs and outputs) transform to preserve aggregated totals. Assuming that larger SR factors primarily increase underdetermination (hence required context and residual uncertainty) rather than changing the conditional-mean structure, scale adaptivity is achieved by retuning three factor-dependent hyperparameters before retraining: the diffusion noise schedule amplitude beta (larger for larger factors to increase diversity), the temporal context length L (set to maintain comparable attention horizons across cadences) and optionally a third, the mass-conservation function f (tapered to limit the amplification of extremes for large factors). Demonstrated on reanalysis precipitation over France (Comephore), the same architecture spans super-resolution factors from 1 to 25 in space and 1 to 6 in time, yielding a reusable architecture and tuning recipe for joint spatiotemporal super-resolution across scales.

cs.LG

Emulating Non-Differentiable Metrics via Knowledge-Guided Learning: Introducing the Minkowski Image Loss

The ``differentiability gap'' presents a primary bottleneck in Earth system deep learning: since models cannot be trained directly on non-differentiable scientific metrics and must rely on smooth proxies (e.g., MSE), they often fail to capture high-frequency details, yielding ``blurry'' outputs. We develop a framework that bridges this gap using two different methods to deal with non-differentiable functions: the first is to analytically approximate the original non-differentiable function into a differentiable equivalent one; the second is to learn differentiable surrogates for scientific functionals. We formulate the analytical approximation by relaxing discrete topological operations using temperature-controlled sigmoids and continuous logical operators. Conversely, our neural emulator uses Lipschitz-convolutional neural networks to stabilize gradient learning via: (1) spectral normalization to bound the Lipschitz constant; and (2) hard architectural constraints enforcing geometric principles. We demonstrate this framework's utility by developing the Minkowski image loss, a differentiable equivalent for the integral-geometric measures of surface precipitation fields (area, perimeter, connected components). Validated on the EUMETNET OPERA dataset, our constrained neural surrogate achieves high emulation accuracy, completely eliminating the geometric violations observed in unconstrained baselines. However, applying these differentiable surrogates to a deterministic super-resolution task reveals a fundamental trade-off: while strict Lipschitz regularization ensures optimization stability, it inherently over-smooths gradient signals, restricting the recovery of highly localized convective textures. This work highlights the necessity of coupling such topological constraints with stochastic generative architectures to achieve full morphological realism.

cs.LG

Dissipating the correlation smokescreen: Causal decomposition of the radiative effects of biomass burning aerosols over the South-East Atlantic

Biomass burning aerosols (BBAs) from Southern Africa seasonally overlie the semi-permanent South-East Atlantic (SEA) stratocumulus deck, impacting the region's energy budget through complex aerosol-cloud-radiation-meteorology interactions. Climate model intercomparison initiatives, like the Aerosol Comparisons between Observations and Models (AeroCom), have highlighted the large inter-model variability for BBA radiative effects, especially over the SEA, due to parameterization of emission modeling and smoke properties. Observational constraints are needed to reduce these uncertainties, but correlative observational studies are typically affected by confounding meteorological influences. We propose a physically informed statistical approach, based on causal graphs applied to satellite observations, to disentangle BBA influences on shortwave radiation over the SEA and identify the main sources of statistical biases plaguing observational studies. We find that, during the fire season, BBAs cause a regional shortwave cooling of -2.5 W m$^{-2}$, which can be decomposed into equal contributions from three physical pathways: aerosol-radiation interactions (ARI), adjustments to ARI, and aerosol-cloud interactions (ACI). We also perform ablation experiments with graph variants to investigate the main sources of confounding - like large-scale winds, humidity-biased retrievals or spatial aggregation of data - and show that they result in biased radiative effect estimates (between -50 $\%$ and +15 $\%$). Once free of such biases, our derived causal estimates of smoke radiative effects can be used as observational constraints to improve climate models.

physics.ao-ph

Calibrated Conformal Prediction Intervals for Microphysical Process Rates

Conformal prediction can yield statistically valid prediction intervals for any regression model, with no model modifications and small computational costs. To assess its practical value, we apply conformal methods to quantify uncertainty in machine learning emulators of six microphysical process rates. Microphysical process rates describe small-scale processes in atmospheric clouds such as precipitation formation and aerosol-cloud interactions, and help understand weather and climate. The emulators are trained on simulation output from the ICOsahedral Nonhydrostatic (ICON) model in a limited-area numerical weather prediction configuration. We compare split conformal prediction for deterministic emulators with conformalized quantile regression for quantile regression emulators. Both conformal prediction methods yield well-calibrated and sharp prediction intervals on average, but conformalized quantile regression provides more consistent intervals across several orders of magnitude, making it preferable for the uncertainty quantification of climate variables.

physics.ao-ph

Data-Driven Integration Kernels for Interpretable Nonlocal Operator Learning

Machine learning models can represent climate processes that are nonlocal in horizontal space, height, and time, often by combining information across these dimensions in highly nonlinear ways. While this can improve predictive skill, it makes learned relationships difficult to interpret and prone to overfitting as the extent of nonlocal information grows. We address this challenge by introducing data-driven integration kernels, a framework that adds structure to nonlocal operator learning by explicitly separating nonlocal information aggregation from local nonlinear prediction. Each spatiotemporal predictor field is first integrated using learnable kernels (defined as continuous weighting functions over horizontal space, height, and/or time), after which a local nonlinear mapping is applied only to the resulting kernel-integrated features and optional local inputs. This design confines nonlinear interactions to a small set of integrated features and makes each kernel directly interpretable as a weighting pattern that reveals which horizontal locations, vertical levels, and past timesteps contribute most to the prediction. We demonstrate the framework for South Asian monsoon precipitation using a hierarchy of neural network models with increasing structure, including baseline, nonparametric kernel, and parametric kernel models. Across this hierarchy, kernel models achieve near-baseline performance with far fewer trainable parameters, indicating that much of the relevant nonlocal information can be captured through a small set of interpretable integrations when appropriate structural constraints are imposed.

cs.LG

Machine Learning of Vertical Fluxes by Unresolved Midlatitude Mesoscale Processes

Machine learning (ML) can represent processes unresolved in coarse-resolution Earth system models (ESMs) by learning from high-resolution climate data. Such ML parameterization approaches have been primarily tested in idealized setups where they have focused on deep convection. It remains largely unexplored whether these approaches could be used in a more targeted fashion to learn vertical fluxes resulting from midlatitude mesoscale processes, such as slantwise convection and frontal dynamics in extratropical cyclones, which are not well represented in ESMs. To address this, we employ a variable-resolution CESM2 simulation with a refined area over the North Atlantic (14-km grid refinement) that resolves such midlatitude mesoscale processes. We train an artificial neural network to predict vertical profiles of mesoscale moisture, heat, and momentum fluxes from the perspective of a coarse-resolution (111-km grid) model. Our results show that a large number of features are required to achieve reasonable model performance when data come from the midlatitudes of real-geography atmospheric simulations, especially when coarse-grained vertical velocities, which we show are not representative of vertical velocities in a coarse-resolution model, are excluded as inputs. Feature importance analysis reveals the importance of vertically non-local information in temperature, moisture, and the meridional wind. We suggest that these non-local relationships capture the influence of cold air outbreaks and fronts on mesoscale fluxes. Our results demonstrate the importance of vertically non-local processes, clarify the regime-dependent predictability of mesoscale fluxes, and identify variables most informative for their parameterization, providing guidance for improving ESMs with ML and advancing our understanding of multi-scale interactions in the midlatitudes.

physics.ao-ph

TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale

TCBench is a benchmark for evaluating global, short to medium-range (1-5 days) forecasts of tropical cyclone (TC) track and intensity. To allow a fair and model-agnostic comparison, TCBench builds on the IBTrACS observational dataset and formulates TC forecasting as predicting the time evolution of an existing tropical system conditioned on its initial position and intensity. TCBench includes state-of-the-art physics-based (TIGGE) and Artificial Intelligence Weather Prediction (AIWP) models (AIFS, Pangu-Weather, FourCastNet v2, GenCast, FNV3). If not readily available (e.g., from the NOAA website as is done with TIGGE), TC tracks are consistently derived from model outputs using the TempestExtremes library. TCBench provides deterministic and probabilistic storm-following metrics. On 2023 test cases, AIWP models skillfully forecast TC tracks, while skillful intensity forecasts require additional steps such as post-processing or task-specific training. Designed for accessibility, TCBench helps AI practitioners tackle domain-relevant TC challenges and equips tropical meteorologists with data-driven tools and workflows to improve prediction and TC process understanding. By lowering barriers to reproducible, process-aware evaluation of extreme events, TCBench aims to democratize data-driven TC forecasting.

cs.CE

Crowdsourcing the Frontier: Advancing Hybrid Physics-ML Climate Simulation via a $50,000 Kaggle Competition

Subgrid machine-learning (ML) parameterizations have the potential to introduce a new generation of climate models that incorporate the effects of higher-resolution physics without incurring the prohibitive computational cost associated with more explicit physics-based simulations. However, important issues, ranging from online instability to inconsistent online performance, have limited their operational use for long-term climate projections. To more rapidly drive progress in solving these issues, domain scientists and machine learning researchers opened up the offline aspect of this problem to the broader machine learning and data science community with the release of ClimSim, a NeurIPS Datasets and Benchmarks publication, and an associated Kaggle competition. This paper reports on the downstream results of the Kaggle competition by coupling emulators inspired by the winning teams' architectures to an interactive climate model (including full cloud microphysics, a regime historically prone to online instability) and systematically evaluating their online performance. Our results demonstrate that online stability in the low-resolution, real-geography setting is reproducible across multiple diverse architectures, which we consider a key milestone. All tested architectures exhibit strikingly similar offline and online biases, though their responses to architecture-agnostic design choices (e.g., expanding the list of input variables) can differ significantly. Multiple Kaggle-inspired architectures achieve state-of-the-art (SOTA) results on certain metrics such as zonal mean bias patterns and global RMSE, indicating that crowdsourcing the essence of the offline problem is one path to improving online performance in hybrid physics-AI climate simulation.

physics.ao-ph

Beyond the Training Data: Confidence-Guided Mixing of Parameterizations in a Hybrid AI-Climate Model

Persistent systematic errors in Earth system models (ESMs) arise from difficulties in representing the full diversity of subgrid, multiscale atmospheric convection and turbulence. Machine learning (ML) parameterizations trained on short high-resolution simulations show strong potential to reduce these errors. However, stable long-term atmospheric simulations with hybrid (physics + ML) ESMs remain difficult, as neural networks (NNs) trained offline often destabilize online runs. Training convection parameterizations directly on coarse-grained data is challenging, notably because scales cannot be cleanly separated. This issue is mitigated using data from superparameterized simulations, which provide clearer scale separation. Yet, transferring a parameterization from one ESM to another remains difficult due to distribution shifts that induce large inference errors. Here, we present a proof-of-concept where a ClimSim-trained, physics-informed NN convection parameterization is successfully transferred to ICON-A. The scheme is (a) trained on adjusted ClimSim data with subtracted radiative tendencies, and (b) integrated into ICON-A. The NN parameterization predicts its own error, enabling mixing with a conventional convection scheme when confidence is low, thus making the hybrid AI-physics model tunable with respect to observations and reanalysis through mixing parameters. This improves process understanding by constraining convective tendencies across column water vapor, lower-tropospheric stability, and geographical conditions, yielding interpretable regime behavior. In AMIP-style setups, several hybrid configurations outperform the default convection scheme (e.g., improved precipitation statistics). With additive input noise during training, both hybrid and pure-ML schemes lead to stable simulations and remain physically consistent for at least 20 years.

physics.ao-ph

Multidata Causal Discovery for Statistical Hurricane Intensity Forecasting

Improving statistical forecasts of tropical cyclone (TC) intensity is limited by complex nonlinear interactions and difficulty in identifying relevant predictors. Conventional methods prioritize correlation or fit, often overlooking confounding variables and limiting generalizability to unseen TCs. To address this, we leverage a multidata causal discovery framework with a replicated dataset based on Statistical Hurricane Intensity Prediction Scheme (SHIPS) using ERA5 meteorological reanalysis. We conduct experiments to identify and select predictors causally linked to TC intensity changes. We then train multiple linear regression models to compare causal feature selection with correlation, random forest feature importance, and no feature selection, across five forecast lead times from 1 to 5 days (24 to 120 hours). Causal feature selection consistently outperforms on unseen test cases, especially for lead times shorter than 3 days. Top causal features include vertical shear, mid-tropospheric potential vorticity and surface moisture conditions, which are physically significant yet often underutilized in TC intensity predictions. We build an extended predictor set (SHIPS+) by adding selected features to the standard SHIPS predictors. SHIPS+ yields increased short-term predictive skill at lead times of 24, 48, and 72 hours. Adding nonlinearity using a multilayer perceptron further extends skill to longer lead times, despite our framework being purely regional and not requiring global forecast data. Operational SHIPS tests confirm that three of the six added causally discovered predictors improve forecast skill, with the largest gains at longer lead times. Our results demonstrate that causal discovery improves TC intensity prediction and pave the way toward more empirical forecasts.

stat.AP

Global Forecasting of Tropical Cyclone Intensity Using Neural Weather Models

Numerical Weather Prediction (NWP) models that integrate coupled physical equations forward in time are the traditional tools for simulating atmospheric processes and forecasting weather. With recent advancements in deep learning, AI-based Weather Prediction models that rely on neural network architectures$\unicode{x2013}$Neural Weather Models (NeWMs)$\unicode{x2013}$have emerged as competent medium-range NWP emulators, with performances that compare favorably to state-of-the-art NWP models. However, they are commonly trained on reanalyses with limited spatial resolution (e.g., 0.25{\deg} horizontal grid spacing), which smooths out key features of weather systems. For example, tropical cyclones (TCs)$\unicode{x2013}$among the most impactful weather events due to their devastating effects on human activities$\unicode{x2013}$are challenging to forecast, as extrema are smoothed in deterministic forecasts at 0.25{\deg} resolution. To address this, we use our best observational estimates of wind gusts and minimum sea level pressure to train a hierarchy of post-processing models on NeWM outputs. Applied to Pangu-Weather and FourCastNet v2, the post-processing models produce accurate and reliable forecasts of TC intensity up to five days ahead. Our post-processing algorithm is tracking-independent, preventing full misses, and we demonstrate that even linear models extract predictive information from NeWM outputs beyond what is encoded in their initial conditions. While spatial masking improves probabilistic forecast consistency, we do not find clear advantages of convolutional architectures over simple multilayer perceptrons for our NeWM post-processing purposes. Overall, by combining the efficiency of NeWMs with a lightweight, tracking-independent postprocessing framework, our approach improves the accessibility of global TC intensity forecasts, marking a step toward their democratization.

physics.comp-ph

Setting the Standard: Recommended Practices for Data Preprocessing in Data-Driven Climate Prediction

Artificial intelligence (AI) - and specifically machine learning (ML) - applications for climate prediction across timescales are proliferating quickly. The emergence of these methods prompts a revisit to the impact of data preprocessing, a topic familiar to the climate community, as more traditional statistical models work with relatively small sample sizes. Indeed, the skill and confidence in the forecasts produced by data-driven models are directly influenced by the quality of the datasets and how they are treated during model development, thus yielding the colloquialism, "garbage in, garbage out." As such, this article establishes protocols for the proper preprocessing of input data for AI/ML models designed for climate prediction (i.e., subseasonal to decadal and longer). The three aims are to: (1) educate researchers, developers, and end users on the effects that preprocessing has on climate predictions; (2) provide recommended practices for data preprocessing for such applications; and (3) empower end users to decipher whether the models they are using are properly designed for their objectives. Specific topics covered in this article include the creation of (standardized) anomalies, dealing with non-stationarity and the spatiotemporally correlated nature of climate data, and handling of extreme values and variables with potentially complex distributions. Case studies will illustrate how using different preprocessing techniques can produce different predictions from the same model, which can create confusion and decrease confidence in the overall process. Ultimately, implementing the recommended practices set forth in this article will enhance the robustness and transparency of AI/ML in climate prediction studies.

physics.data-an

Investigating the Robustness of Extreme Precipitation Super-Resolution Across Climates

The coarse spatial resolution of gridded climate models, such as general circulation models, limits their direct use in projecting socially relevant variables like extreme precipitation. Most downscaling methods estimate the conditional distributions of extremes by generating large ensembles, complicating the assessment of robustness under distributional transformations, such as those induced by climate change. To better understand and potentially improve robustness, we propose super-resolving the parameters of the target variable's probability distribution directly using analytically tractable mappings. Within a perfect-model framework over Switzerland, we demonstrate that vector generalized linear and additive models can super-resolve the generalized extreme value distribution of summer hourly precipitation extremes from coarse precipitation fields and topography. We introduce the notion of a "robustness gap", defined as the difference in predictive error between present-trained and future-trained models, and use it to diagnose how model structure affects the generalization of each quantile to a pseudo-global warming scenario. By evaluating multiple model configurations, we also identify an upper limit on the super-resolution factor based on the spatial auto- and cross-correlation of precipitation and elevation, beyond which coarse precipitation loses predictive value. Our framework is broadly applicable to variables governed by parametric distributions and offers a model-agnostic diagnostic for understanding when and why empirical downscaling generalizes to climate change and extremes.

physics.ao-ph

Reduced Cloud Cover Errors in a Hybrid AI-Climate Model Through Equation Discovery And Automatic Tuning

Cloud-related parameterizations remain a leading source of uncertainty in climate projections. Although machine learning holds promise for Earth system models (ESMs), many data-driven parameterizations lack interpretability, physical consistency, and smooth integration into ESMs. Here, a two-step method is presented to improve a climate model with data-driven parameterizations. First, we incorporate a physically consistent cloud cover parameterization -- derived from storm-resolving simulations via symbolic regression, preserving interpretability while enhancing accuracy -- into the ICON global atmospheric model. Second, we apply the gradient-free Nelder-Mead optimizer to automatically recalibrate the hybrid model against Earth observations, tuning in nested stages (2-, 7-, 30- and 365-day runs) to ensure stability and tractability. The tuned hybrid model substantially reduces long-standing biases in cloud cover -- particularly over the Southern Ocean (by 75%) and subtropical stratocumulus regions (by 44%) -- and remains robust under +4K surface warming. These results demonstrate that interpretable machine-learned parameterizations, paired with practical tuning, can efficiently and transparently strengthen ESM fidelity.

physics.ao-ph