SearcharxivSearch

arXiv subjects

Sebastian Lerch

Publications and source records attributed to Sebastian Lerch.

At least 19 recordsLinked to original sources

Towards Fair Comparisons of AI- and Physics-Based Weather Models for Extreme Events via the Weighted Potential CRPS

We study whether deterministic AI weather prediction (AIWP) models issue more informative forecasts for extreme weather events than deterministic numerical weather prediction (NWP) models. The deterministic model output is subjected to statistical post-processing via isotonic distributional regression (IDR), or EasyUQ, before the resulting probabilistic forecasts are assessed using weighted versions of the continuous ranked probability score (CRPS). This extends the Potential CRPS (PCRPS) measure proposed by Gneiting et al. (2026) to focus on extreme outcomes. Since IDR exhibits optimality properties with respect to weighted versions of the CRPS, the proposed approach inherits desirable properties of the PCRPS, and, in particular, facilitates fair comparisons between data-driven and physics-based models when forecasting extreme weather events. We apply this evaluation framework to forecasts in the WeatherBench 2 dataset issued by the AIWP models GraphCast, Pangu-Weather, and FuXi, with the ECMWF's high-resolution NWP model serving as a physics-based reference. The forecast models are compared when predicting mean sea level pressure, temperature, wind speed, and precipitation extremes, defined as exceedances or non-exceedances of thresholds obtained from historical observation data. We additionally study forecast performance when predicting record-breaking events, though the ordering of the different methods is largely insensitive to the thresholds on which emphasis is placed. We find that AIWP models, particularly FuXi, result in the most informative forecasts for extreme weather events across most settings, suggesting that AIWP models have the potential to outperform NWP models when forecasting extremes.

stat.AP

Uncertainty Quantification in Forecast Comparisons

Skill scores, which measure the relative improvement of a forecasting method over a benchmark via consistent scoring functions and proper scoring rules, are a standard tool in forecast evaluation, yet their sampling uncertainty is rarely rigorously quantified. With modern forecasting applications being increasingly multivariate and involving evaluations across multiple horizons, variables, spatial locations, and forecasting methods, standard tools like the pairwise Diebold-Mariano forecast accuracy test or pointwise confidence intervals fail to account for the multiple comparison problem, leading to inflated Type I error rates and invalid joint inference. To address the lack of a coherent, statistically rigorous framework for quantifying uncertainty across these multi-dimensional evaluation problems, we introduce simultaneous confidence bands for expected scores and skill scores. Our framework provides a versatile tool for joint inference that is applicable to any forecast type from mean and quantile to full distributional forecasts. We develop a bootstrap implementation and show that our bands are valid under multivariate extensions of the classical Diebold-Mariano assumptions. We demonstrate the practical utility of the approach in two case studies by quantifying the benefits of time-varying parameter models for macroeconomic forecasting, and by comparing data-driven and physics-based models in probabilistic weather forecasting.

stat.ME

Energy-Arena: A Dynamic Benchmark for Operational Energy Forecasting

Energy forecasting research faces a persistent comparability gap that makes it difficult to measure consistent progress over time. Reported accuracy gains are often not directly comparable because models are evaluated under study-specific datasets, time periods, information sets, and scoring setups, while widely used benchmarks and competition datasets are typically tied to fixed historical windows. This paper introduces the Energy-Arena, a dynamic benchmarking platform for operational energy time series forecasting that provides a continuously updated reference point as energy systems evolve. The platform operates as an open, API-based submission system and standardizes challenge definitions and submission deadlines aligned with operational constraints. Performance is reported on rolling evaluation windows via persistent leaderboards. By moving from retrospective backtesting to forward-looking benchmarking, the Energy-Arena enforces standardized ex-ante submission and ex-post evaluation, thereby improving transparency by preventing information leakage and retroactive tuning. The platform is publicly available at Energy-Arena.org.

econ.EM

Post-processing of ensemble photovoltaic power forecasts with distributional and quantile regression methods

Accurate and reliable forecasting of photovoltaic (PV) power generation is crucial for grid operations, electricity markets, and energy planning, as solar systems now contribute a significant share of the electricity supply in many countries. PV power forecasts are often generated by converting forecasts of relevant weather variables to power predictions via a model chain. The use of ensemble simulations from numerical weather prediction models results in probabilistic PV forecasts in the form of a forecast ensemble. However, weather forecasts often exhibit systematic errors that propagate through the model chain, leading to biased and/or uncalibrated PV power predictions. These deficiencies can be mitigated by statistical post-processing. Using PV production data and corresponding short-term PV power ensemble forecasts at seven utility-scale PV plants in Hungary, we systematically evaluate and compare seven state-of-the-art methods for post-processing PV power forecasts. These include both parametric and non-parametric techniques, as well as statistical and machine learning-based approaches. Our results show that compared to the raw PV power ensemble, any form of statistical post-processing significantly improves the predictive performance. Non-parametric methods outperform parametric models, with advanced nonlinear quantile regression models showing the best results. Furthermore, machine learning-based approaches surpass their traditional statistical counterparts.

stat.AP

Operational convection-permitting COSMO/ICON ensemble predictions at observation sites (CIENS)

We present the CIENS dataset, which contains ensemble weather forecasts from the operational convection-permitting numerical weather prediction model of the German Weather Service. It comprises forecasts for 55 meteorological variables mapped to the locations of synoptic stations, as well as additional spatially aggregated forecasts from surrounding grid points, available for a subset of these variables. Forecasts are available at hourly lead times from 0 to 21 hours for two daily model runs initialized at 00 and 12 UTC, covering the period from December 2010 to June 2023. Additionally, the dataset provides station observations for six key variables at 170 locations across Germany: pressure, temperature, hourly precipitation accumulation, wind speed, wind direction, and wind gusts. Since the forecast are mapped to the observed locations, the data is delivered in a convenient format for analysis. The CIENS dataset complements the growing collection of benchmark datasets for weather and climate modeling. A key distinguishing feature is its long temporal extent, which encompasses multiple updates to the underlying numerical weather prediction model and thus supports investigations into how forecasting methods can account for such changes. In addition to detailing the design and contents of the CIENS dataset, we outline potential applications in ensemble post-processing, forecast verification, and related research areas. A use case focused on ensemble post-processing illustrates the benefits of incorporating the rich set of available model predictors into machine learning-based forecasting models.

physics.ao-ph

Probabilistic measures afford fair comparisons of AIWP and NWP model output

We introduce a new measure for fair and meaningful comparisons of single-valued output from artificial intelligence based weather prediction (AIWP) and numerical weather prediction (NWP) models, called potential continuous ranked probability score (PC). In a nutshell, we subject the deterministic backbone of physics-based and data-driven models post hoc to the same statistical postprocessing technique, namely, isotonic distributional regression (IDR). Then we find PC as the mean continuous ranked probability score (CRPS) of the postprocessed probabilistic forecasts. The nonnegative PC measure quantifies potential predictive performance and is invariant under strictly increasing transformations of the model output. PC attains its most desirable value of zero if, and only if, the weather outcome Y is a fixed, non-decreasing function of the model output X. The PC measure is recorded in the unit of the outcome, has an upper bound of one half times the mean absolute difference between outcomes, and serves as a proxy for the mean CRPS of real-time, operational probabilistic products. When applied to WeatherBench 2 data, our approach demonstrates that the data-driven GraphCast model outperforms the leading, physics-based European Centre for Medium Range Weather Forecasts (ECMWF) high-resolution (HRES) model. Furthermore, the PC measure for the HRES model aligns exceptionally well with the mean CRPS of the operational ECMWF ensemble. Across application domains, our approach affords comparisons of single-valued forecasts in settings where the pre-specification of a loss function -- which is the usual, and principally superior, procedure in forecast contests, administrative, and benchmarks settings -- places competitors on unequal footings.

stat.AP

Probabilistic intraday electricity price forecasting using generative machine learning

The growing importance of intraday electricity trading in Europe calls for improved price forecasting and tailored decision-support tools. In this paper, we propose a novel generative neural network model to generate probabilistic path forecasts for intraday electricity prices and use them to construct effective trading strategies for Germany's continuous-time intraday market. Our method demonstrates competitive performance in terms of statistical evaluation metrics compared to two state-of-the-art statistical benchmark approaches. To further assess its economic value, we consider a realistic fixed-volume trading scenario and propose various strategies for placing market sell orders based on the path forecasts. Among the different trading strategies, the price paths generated by our generative model lead to higher profit gains than the benchmark methods. Our findings highlight the potential of generative machine learning tools in electricity price forecasting and underscore the importance of economic evaluation.

stat.AP

Windows of opportunity in subseasonal weather regime forecasting: A statistical-dynamical approach

MJO and SPV are prominent sources of subseasonal predictability in the Extratropics. With relevance for European weather it has been shown that the joint interaction of MJO and the SPV can modulate the preferred phase of the NAO and the occurrence of weather regimes. However, improving extended-range NWP at three-week lead times remain under-explored. This study investigates how MJO and SPV phases affect Greenland Blocking (GL) activity and integrates atmospheric state information into a neural network to enhance week-three weather regime activity forecasts. We define a weather regime activity metric using ECMWF reanalysis and reforecasts. In reanalyses we find increased GL activity following MJO phases 7,8 and 1, as well as weak SPV phases, indicating climatological windows of opportunity in line with previous studies. However, ECMWF forecast skill improves only in MJO phases 8 and 1 and weak SPV phases, identifying somewhat different model windows of opportunities. Next we explore using these findings in post-processing tools. Climatological forecasts based on MJO/SPV-NAO relationships provide a purely statistical approach to extended-range GL activity forecasting, independent of NWP models. Notably, MJO conditioned climatological forecasts show clear signals when evaluated against observed GL activity. Statistical-dynamical models, using neural networks that combine historical atmospheric state data with NWP-derived weather regime metrics improve weather regime activity forecasts across all regimes considered, achieving an absolute accuracy increase of 2.9% in forecasting the dominant weather regime compared to ECMWF. This is particularly beneficial to Blocking in the European domain, where NWP models often underperform. Atmospheric conditioned and neural network forecasts serve as valuable decision-support tools alongside NWP models, enhancing the reliability of S2S predictions.

physics.ao-ph

Learning low-dimensional representations of ensemble forecast fields using autoencoder-based methods

Large-scale numerical simulations often produce high-dimensional gridded data that is challenging to process for downstream applications. A prime example is numerical weather prediction, where atmospheric processes are modeled using discrete gridded representations of the physical variables and dynamics. Uncertainties are assessed by running the simulations multiple times, yielding ensembles of simulated fields as a high-dimensional stochastic representation of the forecast distribution. The high-dimensionality and large volume of ensemble datasets poses major computing challenges for subsequent forecasting stages. Data-driven dimensionality reduction techniques could help to reduce the data volume before further processing by learning meaningful and compact representations. However, existing dimensionality reduction methods are typically designed for deterministic and single-valued inputs, and thus cannot handle ensemble data from multiple randomized simulations. In this study, we propose novel dimensionality reduction approaches specifically tailored to the format of ensemble forecast fields. We present two alternative frameworks, which yield low-dimensional representations of ensemble forecasts while respecting their probabilistic character. The first approach derives a distribution-based representation of an input ensemble by applying standard dimensionality reduction techniques in a member-by-member fashion and merging the member representations into a joint parametric distribution model. The second approach achieves a similar representation by encoding all members jointly using a tailored variational autoencoder. We evaluate and compare both approaches in a case study using 10 years of temperature and wind speed forecasts over Europe. The approaches preserve key spatial and statistical characteristics of the ensemble and enable probabilistic reconstructions of the forecast fields.

cs.LG

Graph Neural Networks and Spatial Information Learning for Post-Processing Ensemble Weather Forecasts

Ensemble forecasts from numerical weather prediction models show systematic errors that require correction via post-processing. While there has been substantial progress in flexible neural network-based post-processing methods over the past years, most station-based approaches still treat every input data point separately which limits the capabilities for leveraging spatial structures in the forecast errors. In order to improve information sharing across locations, we propose a graph neural network architecture for ensemble post-processing, which represents the station locations as nodes on a graph and utilizes an attention mechanism to identify relevant predictive information from neighboring locations. In a case study on 2-m temperature forecasts over Europe, the graph neural network model shows substantial improvements over a highly competitive neural network-based post-processing method.

cs.LG

Improving Model Chain Approaches for Probabilistic Solar Energy Forecasting through Post-processing and Machine Learning

Weather forecasts from numerical weather prediction models play a central role in solar energy forecasting, where a cascade of physics-based models is used in a model chain approach to convert forecasts of solar irradiance to solar power production, using additional weather variables as auxiliary information. Ensemble weather forecasts aim to quantify uncertainty in the future development of the weather, and can be used to propagate this uncertainty through the model chain to generate probabilistic solar energy predictions. However, ensemble prediction systems are known to exhibit systematic errors, and thus require post-processing to obtain accurate and reliable probabilistic forecasts. The overarching aim of our study is to systematically evaluate different strategies to apply post-processing methods in model chain approaches: Not applying any post-processing at all; post-processing only the irradiance predictions before the conversion; post-processing only the solar power predictions obtained from the model chain; or applying post-processing in both steps. In a case study based on a benchmark dataset for the Jacumba solar plant in the U.S., we develop statistical and machine learning methods for post-processing ensemble predictions of global horizontal irradiance and solar power generation. Further, we propose a neural network-based model for direct solar power forecasting that bypasses the model chain. Our results indicate that post-processing substantially improves the solar power generation forecasts, in particular when post-processing is applied to the power predictions. The machine learning methods for post-processing yield slightly better probabilistic forecasts, and the direct forecasting approach performs comparable to the post-processing strategies.

stat.AP

Multivariate post-processing of probabilistic sub-seasonal weather regime forecasts

Reliable forecasts of quasi-stationary, recurrent, and persistent large-scale atmospheric circulation patterns (weather regimes) are crucial for various socio-economic sectors. Despite steady progress, probabilistic weather regime predictions still exhibit biases in the exact timing and amplitude of weather regimes. This study thus aims at advancing probabilistic weather regime predictions in the North Atlantic-European region through ensemble post-processing. Here, we focus on the representation of seven year-round weather regimes in the sub-seasonal to seasonal reforecasts of the European Centre for Medium-Range Weather Forecasts. The manifestation of each of the seven regimes can be expressed by a continuous weather regime index, representing the projection of the instantaneous 500-hPa geopotential height anomalies (Z500A) onto the respective mean regime pattern. We apply a two-step ensemble post-processing involving first univariate ensemble model output statistics and second ensemble copula coupling, which restores the multivariate dependency structure. Compared to current forecast calibration practices, which rely on correcting the Z500 field by the lead time dependent mean bias, our approach extends the forecast skill horizon for daily/instantaneous regime forecasts moderately by 1.2 days to 14.5 days. Additionally, to our knowledge our study is the first to systematically evaluate the multivariate aspects of forecast quality for weather regime forecasts. Our method outperforms current practices in the multivariate aspect, as measured by the energy and variogram score. Still our study shows, that even with advanced post-processing weather regime prediction becomes difficult beyond 14 days, which likely points towards intrinsic limits of predictability for daily/instantaneous regime forecasts. The proposed method can easily be applied to operational weather regime forecasts.

physics.ao-ph

Uncertainty quantification for data-driven weather models

Artificial intelligence (AI)-based data-driven weather forecasting models have experienced rapid progress over the last years. Recent studies, with models trained on reanalysis data, achieve impressive results and demonstrate substantial improvements over state-of-the-art physics-based numerical weather prediction models across a range of variables and evaluation metrics. Beyond improved predictions, the main advantages of data-driven weather models are their substantially lower computational costs and the faster generation of forecasts, once a model has been trained. However, most efforts in data-driven weather forecasting have been limited to deterministic, point-valued predictions, making it impossible to quantify forecast uncertainties, which is crucial in research and for optimal decision making in applications. Our overarching aim is to systematically study and compare uncertainty quantification methods to generate probabilistic weather forecasts from a state-of-the-art deterministic data-driven weather model, Pangu-Weather. Specifically, we compare approaches for quantifying forecast uncertainty based on generating ensemble forecasts via perturbations to the initial conditions, with the use of statistical and machine learning methods for post-hoc uncertainty quantification. In a case study on medium-range forecasts of selected weather variables over Europe, the probabilistic forecasts obtained by using the Pangu-Weather model in concert with uncertainty quantification methods show promising results and provide improvements over ensemble forecasts from the physics-based ensemble weather model of the European Centre for Medium-Range Weather Forecasts for lead times of up to 5 days.

physics.ao-ph

Comparison of Model Output Statistics and Neural Networks to Postprocess Wind Gusts

Wind gust prediction plays an important role in warning strategies of national meteorological services due to the high impact of its extreme values. However, forecasting wind gusts is challenging because they are influenced by small-scale processes and local characteristics. To account for the different sources of uncertainty, meteorological centers run ensembles of forecasts and derive probabilities of wind gusts exceeding a threshold. These probabilities often exhibit systematic errors and require postprocessing. Model Output Statistics (MOS) is a common operational postprocessing technique, although more modern methods such as neural network-bases approaches have shown promising results in research studies. The transition from research to operations requires an exhaustive comparison of both techniques. Taking a first step into this direction, our study presents a comparison of a postprocessing technique based on linear and logistic regression approaches with different neural network methods proposed in the literature to improve wind gust predictions, specifically distributional regression networks and Bernstein quantile networks. We further contribute to investigating optimal design choices for neural network-based postprocessing methods regarding changes of the numerical model in the training period, the use of persistence predictors, and the temporal composition of training datasets. The performance of the different techniques is compared in terms of calibration, accuracy, reliability and resolution based on case studies of wind gust forecasts from the operational weather model of the German weather service and observations from 170 weather stations.

stat.AP

Postprocessing of Ensemble Weather Forecasts Using Permutation-invariant Neural Networks

Statistical postprocessing is used to translate ensembles of raw numerical weather forecasts into reliable probabilistic forecast distributions. In this study, we examine the use of permutation-invariant neural networks for this task. In contrast to previous approaches, which often operate on ensemble summary statistics and dismiss details of the ensemble distribution, we propose networks that treat forecast ensembles as a set of unordered member forecasts and learn link functions that are by design invariant to permutations of the member ordering. We evaluate the quality of the obtained forecast distributions in terms of calibration and sharpness and compare the models against classical and neural network-based benchmark methods. In case studies addressing the postprocessing of surface temperature and wind gust forecasts, we demonstrate state-of-the-art prediction quality. To deepen the understanding of the learned inference process, we further propose a permutation-based importance analysis for ensemble-valued predictors, which highlights specific aspects of the ensemble forecast that are considered important by the trained postprocessing models. Our results suggest that most of the relevant information is contained in a few ensemble-internal degrees of freedom, which may impact the design of future ensemble forecasting and postprocessing systems.

stat.ML

Deep learning for post-processing global probabilistic forecasts on sub-seasonal time scales

Sub-seasonal weather forecasts are becoming increasingly important for a range of socio-economic activities. However, the predictive ability of physical weather models is very limited on these time scales. We propose several post-processing methods based on convolutional neural networks to improve sub-seasonal forecasts by correcting systematic errors of numerical weather prediction models. Our post-processing models operate directly on spatial input fields and are therefore able to retain spatial relationships and to generate spatially homogeneous predictions. They produce global probabilistic tercile forecasts for biweekly aggregates of temperature and precipitation for weeks 3-4 and 5-6. In a case study based on a public forecasting challenge organized by the World Meteorological Organization, our post-processing models outperform recalibrated forecasts from the European Centre for Medium-Range Weather Forecasts (ECMWF), and achieve improvements over climatological forecasts for all considered variables and lead times. We compare several model architectures and training modes and demonstrate that all approaches lead to skillful and well-calibrated probabilistic forecasts. The good calibration of the post-processed forecasts emphasizes that our post-processing models reliably quantify the forecast uncertainty based on deterministic input information in form of the ECMWF ensemble mean forecast fields only.

physics.ao-ph

Direction Augmentation in the Evaluation of Armed Conflict Predictions

In many forecasting settings, there is a specific interest in predicting the sign of an outcome variable correctly in addition to its magnitude. For instance, when forecasting armed conflicts, positive and negative log-changes in monthly fatalities represent escalation and de-escalation, respectively, and have very different implications. In the ViEWS forecasting challenge, a prediction competition on state-based violence, a novel evaluation score called targeted absolute deviation with direction augmentation (TADDA) has therefore been suggested, which accounts for both for the sign and magnitude of log-changes. While it has a straightforward intuitive motivation, the empirical results of the challenge show that a no-change model always predicting a log-change of zero outperforms all submitted forecasting models under the TADDA score. We provide a statistical explanation for this phenomenon. Analyzing the properties of TADDA, we find that in order to achieve good scores, forecasters often have an incentive to predict no or only modest log-changes. In particular, there is often an incentive to report conservative point predictions considerably closer to zero than the forecaster's actual predictive median or mean. In an empirical application, we demonstrate that a no-change model can be improved upon by tailoring predictions to the particularities of the TADDA score. We conclude by outlining some alternative scoring concepts.

stat.AP

Learning to forecast: The probabilistic time series forecasting challenge

We report on a course project in which students submit weekly probabilistic forecasts of two weather variables and one financial variable. This real-time format allows students to engage in practical forecasting, which requires a diverse set of skills in data science and applied statistics. We describe the context and aims of the course, and discuss design parameters like the selection of target variables, the forecast submission process, the evaluation of forecast performance, and the feedback provided to students. Furthermore, we describe empirical properties of students' probabilistic forecasts, as well as some lessons learned on our part.

stat.OT