SearcharxivSearch

arXiv subjects

Florian Pappenberger

Publications and source records attributed to Florian Pappenberger.

15 recordsLinked to original sources

Improved rainfall forecasts in daily use over East Africa

Ensemble forecasting has proven to be a vital tool for predicting extreme, life-threatening or only partially predictable weather events. As well as providing probabilistic products, individual ensemble members provide context and realisations of possible extreme weather. However, many National Meteorological Services in East Africa do not have the computing resources to enable them to run their local area models in ensemble mode over the full period of the two-week medium range. In this paper we test the performance of a forecast system, comprising the global ECMWF ensemble forecast, post-processed using cGAN, a neural network model, and forecasts calibrated using IDR, a recent statistical method, against the probabilistic climatology. cGAN provides comparable levels of probabilistic skill as IDR applied to the ECMWF ensemble or IDR applied to the machine-learning models FuXi and GraphCast. These methods improve the raw ECMWF ensemble forecast substantially, which is itself an improvement on deterministic forecasts. The ability of cGAN to produce individual realisations of future rainfall was important when assessed in an operational context by the National Meteorological and Hydrological Services of Kenya and Ethiopia. Moreover, in common with IDR, cGAN is cheap to train/run and requires no additional post-processing. It is run on laptops to generate many thousands of ensemble members, making it suitable for Meteorological Services with limited computational facilities.

physics.ao-ph

AIFL: A Global Daily Streamflow Forecasting Model Using a Deterministic LSTM Pre-trained on ERA5-Land and Fine-tuned on IFS

Reliable global streamflow forecasting is essential for flood preparedness and water resource management, yet data-driven models often suffer from a performance gap when transitioning from historical reanalysis to operational forecast products. This paper introduces AIFL (Artificial Intelligence for Floods), a deterministic LSTM-based model designed for global daily streamflow forecasting. Trained on 18,588 basins curated from the Caravan dataset, AIFL utilises a two-stage transfer-learning strategy to bridge the reanalysis-to-forecast domain shift. The model is first pre-trained on 40 years of ERA5-Land reanalysis (1980-2019) to capture robust hydrological processes, then fine-tuned on operational Integrated Forecasting System (IFS) forecasts (2016-2019) to adapt to the specific error structures and biases of operational numerical weather prediction. Ablation experiments confirm that this two-stage approach outperforms both a naive IFS-only baseline and a mixed-forcing single-stage alternative. To our knowledge, this is the first global model trained end-to-end within the Caravan ecosystem. On an independent temporal test set (2021-2024), AIFL achieves high predictive skill with a median modified Kling-Gupta Efficiency (KGE') of 0.66 and a median Nash-Sutcliffe Efficiency (NSE) of 0.53. Benchmarking results show that AIFL achieves comparable accuracy to current state-of-the-art global systems. The model provides a streamlined and operationally robust baseline for the global hydrological community.

cs.LG

Joint distribution of upstream runoff governs downstream river-discharge prediction uncertainty in distributed ML models

Uncertainty quantification of hydrological predictions is necessary to inform operational decisions. Recent generative machine-learning methods have advanced probabilistic streamflow prediction, but have remained confined to lumped models that predict a basin outlet directly. At the same time, deterministic LSTM runoff models are increasingly applied at grid or catchment scale and routed through river networks to produce spatially continuous, physically consistent discharge fields. This technical note argues that moving probabilistic prediction from lumped to distributed models introduces a specific new requirement: the joint distribution of upstream runoff generation must be sampled jointly. In lumped inference, the model predicts the outlet distribution directly and can modulate spread from basin attributes. In distributed inference, downstream discharge is obtained by routing many upstream runoff predictions, so independent local sampling averages uncertainty away. Using Japan as a case study, we train two probabilistic basin-scale runoff LSTMs and route their runoff through a Hayami routing scheme. Randomly matching upstream ensemble members produces severely under-dispersed downstream ensembles, whereas a simple quantile matching strategy restores much of the spread of the direct basin-scale reference. The shift from lumped to distributed probabilistic hydrology therefore requires explicit attention to the spatial joint structure of runoff uncertainty.

cs.LG

From Licensing to Open Access: Designing a Sustainable Transition in Operational Weather Data

This translational article documents the European Centre for Medium-Range Weather Forecasts (ECMWF) transition from a restricted data licensing model to open access under CC BY 4.0, completed in October 2025. The policy context included EU open data requirements and alignment with international data exchange frameworks. The transition was implemented through a tiered service model that kept core forecast data open while offering operationally supported delivery as a cost-recovered service. Between 2020 and 2025, ECMWF executed an iterative planning cycle: setting an annual target for revenue reduction, specifying additions to the open tier under that target, provisioning infrastructure, and assessing outcomes to update assumptions. Drawing on internal administrative records (2014 - 2025), we describe design choices, operational constraints, and early outcomes. In the six months following the end of the transition, more than 93% of previously paying organisations retained a Service Agreement, while open endpoint download volumes increased substantially. We discuss trade-offs in defining the open tier (resolution, parameters, schedule), the reduction of compliance overheads formerly associated with redistribution restrictions, and the scalability implications of global distribution. We note an emerging sustainability question as AI-based forecast products become freely available. The early evidence is consistent with the view that a tiered service model can be designed to reconcile open-access obligations with operational sustainability, subject to monitoring over longer contract renewal cycles (typically annual).

physics.ao-ph

MSWEP V3: Machine Learning-Powered Global Precipitation Estimates at 0.1$^\circ$ Hourly Resolution (1979-Present)

We introduce Version 3 (V3) of the gridded near real-time Multi-Source Weighted-Ensemble Precipitation (MSWEP) product -- the first fully global, historical machine learning powered precipitation (P) dataset, developed to meet the growing demand for timely and accurate P estimates amid escalating climate challenges. MSWEP V3 provides hourly data at 0.1$^\circ$ resolution from 1979 to the present, continuously updated with a latency of approximately two hours. Development follows a two-stage process. First, baseline P fields are generated using machine learning model stacks that integrate satellite- and (re)analysis-based P and air-temperature products, along with static variables. The models are trained using hourly and daily observations from 15,959 P gauges worldwide. Second, these baseline P fields are corrected using daily and monthly gauge observations from 57,666 and 86,000 stations globally. To assess MSWEP V3's baseline performance, we evaluated 19 (quasi-) global gridded P products -- including both uncorrected and gauge-based products -- using observations from an independent set of 15,958 gauges excluded from the first training stage. The MSWEP V3 baseline achieved a median daily Kling-Gupta Efficiency (KGE) of 0.69, outperforming all evaluated products. Other uncorrected products achieved median daily KGE values of 0.61 (ERA5), 0.46 (IMERG-L V7), 0.38 (GSMaP V8), and 0.31 (CHIRP). Using leave-one-out cross-validation, the daily gauge correction was found to improve the median daily correlation by 0.09, constrained by the already strong baseline performance. We anticipate that MSWEP V3 -- accessible at www.gloh2o.org/mswep -- will enable more reliable monitoring, forecasting, and management of water-related risks in a variable and changing climate.

physics.ao-ph

The ecological forecast limit revisited: Potential, actual and relative system predictability

Ecological forecasts are model-based statements about currently unknown ecosystem states in time or space. For a model forecast to be useful to inform decision makers, model validation and verification determine adequateness. The measure of forecast goodness that can be translated into a limit up to which a forecast is acceptable is known as the 'forecast limit'. While verification in weather forecasting follows strict criteria with established metrics and forecast limits, assessments of ecological forecasting models still remain experiment-specific, and forecast limits are rarely reported. As such, users of ecological forecasts remain uninformed of how far into the future statements can be trusted. In this work, we synthesise existing approaches to define empirical forecast limits in a unified framework for assessing ecological predictability and offer recipes for their computation. We distinguish the model's potential and absolute forecast limit, and show how a benchmark model can help determine its relative forecast limit. The approaches are demonstrated with three case studies from population, ecosystem, and Earth system research. We found that forecast limits can be computed with three requirements: A verification reference, a scoring function, and a predictive error tolerance. Within our framework, forecast limits are defined for practically any ecological forecast and support research on ecological predictability analysis.

stat.AP

Hydra-LSTM: A semi-shared Machine Learning architecture for prediction across Watersheds

Long Short Term Memory networks (LSTMs) are used to build single models that predict river discharge across many catchments. These models offer greater accuracy than models trained on each catchment independently if using the same data. However, the same data is rarely available for all catchments. This prevents the use of variables available only in some catchments, such as historic river discharge or upstream discharge. The only existing method that allows for optional variables requires all variables to be considered in the initial training of the model, limiting its transferability to new catchments. To address this limitation, we develop the Hydra-LSTM. The Hydra-LSTM processes variables used across all catchments and variables used in only some catchments separately to allow general training and use of catchment-specific data in individual catchments. The bulk of the model can be shared across catchments, maintaining the benefits of multi-catchment models to generalise, while also benefitting from the advantages of using bespoke data. We apply this methodology to 1 day-ahead river discharge prediction in the Western US, as next-day river discharge prediction is the first step towards prediction across longer time scales. We obtain state-of-the-art performance, generating more accurate median and quantile predictions than Multi-Catchment and Single-Catchment LSTMs while allowing local forecasters to easily introduce and remove variables from their prediction set. We test the ability of the Hydra-LSTM to incorporate catchment-specific data by introducing historical river discharge as a catchment-specific input, outperforming state-of-the-art models without needing to train an entirely new model.

cs.LG

Robustness of AI-based weather forecasts in a changing climate

Data-driven machine learning models for weather forecasting have made transformational progress in the last 1-2 years, with state-of-the-art ones now outperforming the best physics-based models for a wide range of skill scores. Given the strong links between weather and climate modelling, this raises the question whether machine learning models could also revolutionize climate science, for example by informing mitigation and adaptation to climate change or to generate larger ensembles for more robust uncertainty estimates. Here, we show that current state-of-the-art machine learning models trained for weather forecasting in present-day climate produce skillful forecasts across different climate states corresponding to pre-industrial, present-day, and future 2.9K warmer climates. This indicates that the dynamics shaping the weather on short timescales may not differ fundamentally in a changing climate. It also demonstrates out-of-distribution generalization capabilities of the machine learning models that are a critical prerequisite for climate applications. Nonetheless, two of the models show a global-mean cold bias in the forecasts for the future warmer climate state, i.e. they drift towards the colder present-day climate they have been trained for. A similar result is obtained for the pre-industrial case where two out of three models show a warming. We discuss possible remedies for these biases and analyze their spatial distribution, revealing complex warming and cooling patterns that are partly related to missing ocean-sea ice and land surface information in the training data. Despite these current limitations, our results suggest that data-driven machine learning models will provide powerful tools for climate science and transform established approaches by complementing conventional physics-based models.

physics.ao-ph

AIFS -- ECMWF's data-driven forecasting system

Machine learning-based weather forecasting models have quickly emerged as a promising methodology for accurate medium-range global weather forecasting. Here, we introduce the Artificial Intelligence Forecasting System (AIFS), a data driven forecast model developed by the European Centre for Medium-Range Weather Forecasts (ECMWF). AIFS is based on a graph neural network (GNN) encoder and decoder, and a sliding window transformer processor, and is trained on ECMWF's ERA5 re-analysis and ECMWF's operational numerical weather prediction (NWP) analyses. It has a flexible and modular design and supports several levels of parallelism to enable training on high-resolution input data. AIFS forecast skill is assessed by comparing its forecasts to NWP analyses and direct observational data. We show that AIFS produces highly skilled forecasts for upper-air variables, surface weather parameters and tropical cyclone tracks. AIFS is run four times daily alongside ECMWF's physics-based NWP model and forecasts are available to the public under ECMWF's open data policy.

physics.ao-ph

Advances in Land Surface Model-based Forecasting: A comparative study of LSTM, Gradient Boosting, and Feedforward Neural Network Models as prognostic state emulators

Most useful weather prediction for the public is near the surface. The processes that are most relevant for near-surface weather prediction are also those that are most interactive and exhibit positive feedback or have key role in energy partitioning. Land surface models (LSMs) consider these processes together with surface heterogeneity and forecast water, carbon and energy fluxes, and coupled with an atmospheric model provide boundary and initial conditions. This numerical parametrization of atmospheric boundaries being computationally expensive, statistical surrogate models are increasingly used to accelerated progress in experimental research. We evaluated the efficiency of three surrogate models in speeding up experimental research by simulating land surface processes, which are integral to forecasting water, carbon, and energy fluxes in coupled atmospheric models. Specifically, we compared the performance of a Long-Short Term Memory (LSTM) encoder-decoder network, extreme gradient boosting, and a feed-forward neural network within a physics-informed multi-objective framework. This framework emulates key states of the ECMWF's Integrated Forecasting System (IFS) land surface scheme, ECLand, across continental and global scales. Our findings indicate that while all models on average demonstrate high accuracy over the forecast period, the LSTM network excels in continental long-range predictions when carefully tuned, the XGB scores consistently high across tasks and the MLP provides an excellent implementation-time-accuracy trade-off. The runtime reduction achieved by the emulators in comparison to the full numerical models are significant, offering a faster, yet reliable alternative for conducting numerical experiments on land surfaces.

physics.ao-ph

AI Increases Global Access to Reliable Flood Forecasts

Floods are one of the most common natural disasters, with a disproportionate impact in developing countries that often lack dense streamflow gauge networks. Accurate and timely warnings are critical for mitigating flood risks, but hydrological simulation models typically must be calibrated to long data records in each watershed. Using AI, we achieve reliability in predicting extreme riverine events in ungauged watersheds at up to a 5-day lead time that is similar to or better than the reliability of nowcasts (0-day lead time) from a current state of the art global modeling system (the Copernicus Emergency Management Service Global Flood Awareness System). Additionally, we achieve accuracies over 5-year return period events that are similar to or better than current accuracies over 1-year return period events. This means that AI can provide flood warnings earlier and over larger and more impactful events in ungauged basins. The model developed in this paper was incorporated into an operational early warning system that produces publicly available (free and open) forecasts in real time in over 80 countries. This work highlights a need for increasing the availability of hydrological data to continue to improve global access to reliable flood warnings.

cs.LG

The rise of data-driven weather forecasting

Data-driven modeling based on machine learning (ML) is showing enormous potential for weather forecasting. Rapid progress has been made with impressive results for some applications. The uptake of ML methods could be a game-changer for the incremental progress in traditional numerical weather prediction (NWP) known as the 'quiet revolution' of weather forecasting. The computational cost of running a forecast with standard NWP systems greatly hinders the improvements that can be made from increasing model resolution and ensemble sizes. An emerging new generation of ML models, developed using high-quality reanalysis datasets like ERA5 for training, allow forecasts that require much lower computational costs and that are highly-competitive in terms of accuracy. Here, we compare for the first time ML-generated forecasts with standard NWP-based forecasts in an operational-like context, initialized from the same initial conditions. Focusing on deterministic forecasts, we apply common forecast verification tools to assess to what extent a data-driven forecast produced with one of the recently developed ML models (PanguWeather) matches the quality and attributes of a forecast from one of the leading global NWP systems (the ECMWF IFS). The results are very promising, with comparable skill for both global metrics and extreme events, when verified against both the operational analysis and synoptic observations. Increasing forecast smoothness and bias drift with forecast lead time are identified as current drawbacks of ML-based forecasts. A new NWP paradigm is emerging relying on inference from ML models and state-of-the-art analysis and reanalysis datasets for forecast initialization and model training.

physics.ao-ph

What do large-scale patterns teach us about extreme precipitation over the Mediterranean at medium- and extended-range forecasts?

Extreme Precipitation Events (EPEs) can have devastating consequences such as floods and landslides, posing a great threat to society and the economy. Predicting such events long in advance can support the mitigation of negative impacts. Here, we focus on EPEs over the Mediterranean, a region that is frequently affected by such hazards. Previous work identified strong connections between localized EPEs and large-scale atmospheric flow patterns, affecting the weather over the entire Mediterranean. We analyze the predictive skill of these patterns in the ECMWF extended-range forecasts and assess if and where these patterns can be used for indirect predictions of EPEs, using the Brier Skill Score. The results show that the ECMWF model provides skillful predictions of the Mediterranean patterns up to 2 weeks in advance. Moreover, using the forecasted patterns for indirect predictability of EPEs outperforms the reference score up to about 10 days lead time for many locations. Especially for high orography locations or coastal areas, like parts of western Turkey, western Balkans, Iberian Peninsula and Morocco this limit extends from 11 to 14 days lead time. This study demonstrates that connections between localized EPEs and large-scale patterns over the Mediterranean extend the forecasting horizon of the model by over 3 days in many locations, in comparison to forecasting based on the predicted precipitation. Thus, it is beneficial to use the predicted patterns rather than the predicted precipitation at longer lead times for EPEs forecasting. The model's performance is also assessed from a user perspective, showing that the EPEs forecasting based on the patterns increases the economic benefits at medium and extended range lead times. Such information could support higher confidence in the decision-making of various users, e.g., the agricultural sector and (re)insurance companies.

physics.ao-ph

Extreme precipitation events in the Mediterranean: Spatiotemporal characteristics and connection to large-scale atmospheric flow patterns

The Mediterranean is strongly affected by Extreme Precipitation Events (EPEs), sometimes leading to negative impacts on society, economy, and the environment. Understanding such natural hazards and their drivers is essential to mitigate related risks. Here, EPEs over the Mediterranean between 1979 and 2019 are analyzed, using ERA5 dataset from ECMWF. EPEs are determined based on the 99th percentile of the daily distribution (P99). The different EPE characteristics are assessed, based on seasonality and spatiotemporal dependencies. To better understand the connection to large-scale atmospheric flow patterns, Empirical Orthogonal Function (EOF) analysis and subsequent K-means clustering are used to quantify the importance of weather regimes to EPE frequency. The analysis is performed for three different variables, depicting atmospheric variability in the lower and middle troposphere: Sea level pressure (SLP), temperature at 850 hPa (T850), and geopotential height at 500 hPa (Z500). Results show a clear spatial division in EPEs occurrence, with winter (autumn) being the season of highest EPEs frequency for the eastern (western) Mediterranean. There is a high degree of temporal dependencies with 20% of the EPEs (median value of all studied grid-cells), occurring up to 1 week after a preceding P99 event at the same location. Local orography is a key modulator of the spatiotemporal connections and substantially enhances the probability of co-occurrence of EPEs even for distant locations. The clustering clearly demonstrates the prevalence of distinct synoptic-scale atmospheric conditions during the occurrence of EPEs for different locations within the region. Results indicate that clustering based on a combination of SLP and Z500 can increase the conditional probability of EPEs by more than three (3) times (median value for all grid cells) from the nominal probability of 1% for the P99 EPEs.

physics.ao-ph

Statistical post-processing of heat index ensemble forecasts: is there a royal road?

We investigate the effect of statistical post-processing on the probabilistic skill of discomfort index (DI) and indoor wet-bulb globe temperature (WBGTid) ensemble forecasts, both calculated from the corresponding forecasts of temperature and dew point temperature. Two different methodological approaches to calibration are compared. In the first case, we start with joint post-processing of the temperature and dew point forecasts and then create calibrated samples of DI and WBGTid using samples from the obtained bivariate predictive distributions. This approach is compared with direct post-processing of the heat index ensemble forecasts. For this purpose, a novel ensemble model output statistics model based on a generalized extreme value distribution is proposed. The predictive performance of both methods is tested on the operational temperature and dew point ensemble forecasts of the European Centre for Medium-Range Weather Forecasts and the corresponding forecasts of DI and WBGTid. For short lead times (up to day 6), both approaches significantly improve the forecast skill. Among the competing post-processing methods, direct calibration of heat indices exhibits the best predictive performance, very closely followed by the more general approach based on joint calibration of temperature and dew point temperature. Additionally, a machine learning approach is tested and shows comparable performance for the case when one is interested only in forecasting heat index warning level categories.

stat.AP