SearcharxivSearch

arXiv subjects

Nicholas Loveday

Publications and source records attributed to Nicholas Loveday.

10 recordsLinked to original sources

On equitable scoring functions and optimal forecasting behaviour

Equitable scoring functions have a long history in meteorological forecast verification and have recently gained renewed prominence through the use of the Stable Equitable Error in Probability Space (SEEPS) score for evaluating precipitation forecasts from numerical and machine-learning weather prediction systems. This paper provides a systematic analysis of equitable scoring functions through the optimal forecasts they induce. We introduce the generalized diagonal score, which defines a broad class of equitable scoring functions that includes, up to equivalence, the SEEPS, Gerrity, Peirce, Barnston and diagonal scores. Within this framework, we show that optimal single-valued and categorical forecasts are characterized by crossing points between the predictive and climatological cumulative distribution functions. Moreover, we show that the generalized diagonal score, when evaluating predictive distributions, is proper but not strictly proper, and is insensitive to substantial forecast misspecification away from crossing points. The framework also yields scoring rules that are both proper and equitable for probabilistic forecasts with categorical outcomes. For single-valued and categorical forecasts, the crossing-point characterization shows that optimal forecasts can vary across climatologies even when the underlying predictive distribution is unchanged. Consequently, identical predictive distributions may lead to substantially different optimal forecasts under equitable scores, with important implications for assessing the suitability of such scores for any given application.

math.ST

Extreme Weather Bench: A framework and benchmark for evaluation of high-impact weather

Forecasting the wide variety of high-impact weather events experienced globally is a challenge for both Artificial Intelligence (AI) and Numerical Weather Prediction (NWP) models and it is critical that such models be properly verified before deployment. Although AI weather models are rapidly evolving, much of their evaluation is currently done either with a global-scale evaluation or by hand-picking a small number of case studies or a region. A widely-used open-source benchmark suite focusing on high-impact weather will help to drive the science forward for all scales of weather models, as it has for other AI fields. Here we introduce Extreme Weather Bench (EWB), a new community-driven benchmark suite that facilitates model validation and verification on a variety of high-impact hazards that matter to people around the globe. EWB provides a standard set of case studies (spanning across multiple spatial and temporal scales and different parts of the weather spectrum), observational data, impact-based metrics, and open-source code for users to evaluate their models. Verifying that a model works against a standard set of case studies, especially events that are high-impact for the general public, is a key piece of improving the trustworthiness of AI models. EWB will help to drive the science forward for all weather models, enabling true comparisons across models and evaluating models on specific high-impact phenomena through the use of case studies. EWB is a free open-source community-driven system and will continue to evolve to include additional phenomena, test cases and metrics in collaboration with the worldwide weather and forecast verification community.

cs.LG

WP-MIP: An Artificial Intelligence, Hybrid, and Physically Based Model Intercomparison Project for Weather Prediction

Rapid progress in the field of machine-learning for weather prediction has led to the emergence of algorithms whose forecasting skill can exceed that of traditional physically based models. This development represents an opportunity to improve the quality of forecasting services provided by operational centers, particularly given the speed at which machine-learning based models generate predictions. Despite the clear promise of these systems, questions remain about the ability of the current generation of machine-learning models to generate physically consistent predictions of the full suite of required forecast fields under all conditions. Answering these questions will require careful comparisons between the well-understood physically based models, current state-of-the-art machine-learning models, and the hybrid models that combine elements of these two archetypes. The Weather Prediction Model Intercomparison Project (WP-MIP) is a World Meteorological Organization-supported initiative whose initial goal is to create a centralized database of physically based, machine-learning, and hybrid model forecasts to enable a distributed assessment and evaluation effort. The first instance of WP-MIP focuses on global deterministic predictions using both center-specific and common initializations to facilitate sensitivity studies. Forecasts contributed by institutions across six continents will be used to develop AI-ready verification techniques that highlight the strengths and weaknesses of each class of prediction system, with the goal of establishing best-practice guidance to model developers and national weather centers. The broad engagement of the operational and forecast-evaluation communities in WP-MIP will ensure that the project results are highly relevant to the development and deployment of next-generation weather prediction systems.

physics.ao-ph

On the evaluation of time-to-event, survival time and first passage time forecasts

Time-to-event forecasts are essential when decisions depend on event timing. This article develops a framework for evaluating such forecasts when the event has not yet occurred or is not predicted within the forecast horizon. We introduce a theory of provisional evaluation, in which each forecast is assessed against its right-censored realization, defined as the minimum of the event time and the evaluation time. For probabilistic forecasts, we show that strictly proper scoring rules induce provisionally strictly proper scoring rules, whose expected score, computed from the right-censored realization, is optimized under truthful forecasting. Threshold-weighted versions of the continuous ranked probability score and the logarithmic score satisfy this property. We also develop a theory for scoring point (single-valued) forecasts under right-censoring. Quantile and interquartile range forecasts are shown to be provisionally elicitable, meaning that scoring functions exist for which these functionals uniquely minimize the expected score, whereas the expectation functional is not provisionally elicitable. A synthetic experiment demonstrates that the proposed scores correctly rank forecasters. Diagnostic tools, including Murphy diagrams and reliability diagrams, extend naturally. Applications to operational time-to-flood and time-to-strong-wind forecasts illustrate the approach.

math.ST

Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model

Recent advances in AI-based weather prediction have led to the development of artificial intelligence weather prediction (AIWP) models with competitive forecast skill compared to traditional NWP models, but with substantially reduced computational cost. There is a strong need for appropriate methods to evaluate their ability to predict extreme weather events, particularly when spatial coherence is important, and grid resolutions differ between models. We introduce a verification framework that combines spatial verification methods and weighted proper scoring rules. Specifically, the framework extends the High-Resolution Assessment (HiRA) approach with threshold-weighted scoring rules. It enables user-oriented evaluation consistent with how forecasts may be interpreted by operational meteorologists or used in simple post-processing systems. The method supports targeted evaluation of extreme events by allowing flexible weighting of the relative importance of different decision thresholds. We demonstrate this framework by evaluating 32 months of precipitation forecasts from an AIWP model and a high-resolution NWP model. Our results show that model rankings are sensitive to the choice of neighbourhood size. Increasing the neighbourhood size has a greater impact on scores evaluating extreme-event performance for the high-resolution NWP model than for the AIWP model. At near equivalent neighbourhood sizes, the empirical CDF of the high-resolution NWP model only outperformed the empirical CDF of the AIWP model in predicting extreme precipitation events at short lead times. We also demonstrate how this approach can be extended to evaluate discrimination ability in predicting heavy precipitation. We find that the high-resolution NWP model had superior discrimination ability at short lead times.

physics.ao-ph

Evaluation and statistical correction of area-based heat index forecasts that drive a heatwave warning service

This study evaluates the performance of the area-based, district heatwave forecasts that drive the Australian heatwave warning service. The analysis involves using a recently developed approach of scoring multicategorical forecasts using the FIxed Risk Multicategorical (FIRM) scoring framework. Additionally, we quantify the stability of the district forecasts between forecast updates. Notably, at longer lead times, a discernible overforecast bias exists that leads to issuing severe and extreme heatwave district forecasts too frequently. Consequently, at shorter lead times forecast heatwave categories are frequently downgraded with subsequent revisions. To address these issues, we demonstrate how isotonic regression can be used to conditionally bias correct the district forecasts. Finally, using synthetic experiments, we illustrate that even if an area warning is derived from a perfectly calibrated gridded forecast, the area warning will be biased in most situations. We show how these biases can also be corrected using isotonic regression which could lead to a better warning service. Importantly, the evaluation and bias correction approaches demonstrated in this paper are relevant to forecast parameters other than heat indices.

physics.ao-ph

scores: A Python package for verifying and evaluating models and predictions with xarray

`scores` is a Python package containing mathematical functions for the verification, evaluation and optimisation of forecasts, predictions or models. It supports labelled n-dimensional (multidimensional) data, which is used in many scientific fields and in machine learning. At present, `scores` primarily supports the geoscience communities; in particular, the meteorological, climatological and oceanographic communities. `scores` not only includes common scores (e.g., Mean Absolute Error), it also includes novel scores not commonly found elsewhere (e.g., FIxed Risk Multicategorical (FIRM) score, Flip-Flop Index), complex scores (e.g., threshold-weighted continuous ranked probability score), and statistical tests (such as the Diebold Mariano test). It also contains isotonic regression which is becoming an increasingly important tool in forecast verification and can be used to generate stable reliability diagrams. Additionally, it provides pre-processing tools for preparing data for scores in a variety of formats including cumulative distribution functions (CDF). At the time of writing, `scores` includes over 50 metrics, statistical techniques and data processing tools. All of the scores and statistical techniques in this package have undergone a thorough scientific and software review. Every score has a companion Jupyter Notebook tutorial that demonstrates its use in practice. `scores` supports `xarray` datatypes, allowing it to work with Earth system data in a range of formats including NetCDF4, HDF5, Zarr and GRIB among others. `scores` uses Dask for scaling and performance. Support for `pandas` is being introduced. The `scores` software repository can be found at https://github.com/nci/scores/

physics.ao-ph

The Jive Verification System and its Transformative Impact on Weather Forecasting Operations

Forecast verification is critical for continuous improvement in meteorological organizations. The Jive verification system was originally developed to assess the accuracy of public weather forecasts issued by the Australian Bureau of Meteorology. It started as a research project in 2015 and gradually evolved to be a Bureau operational verification system in 2022. The system includes daily verification dashboards for forecasters to visualize recent forecast performance and "Evidence Targeted Automation" dashboards for exploring the performance of competing forecast systems. Additionally, Jive includes a Jupyter Notebook server with the Jive Python library which supports research experiments, case studies, and the development of new verification metrics and tools. This paper describes the Jive verification system and how it helped bring verification to the forefront at the Bureau of Meteorology, leading to more accurate, streamlined forecasts. Jive has provided evidence to support forecast automation decisions and has helped to understand the evolving role of meteorologists in the forecast process. It has given operational meteorologists tools for evaluating forecast processes, including identifying when and how manual interventions lead to superior predictions. Work on Jive led to new verification science, including novel metrics that are decision-focused, including diagnostics for extreme conditions. Jive also provided the Bureau with an enterprise-wide data analysis environment and has prompted a clarification of forecast definitions. These collective impacts have resulted in more accurate forecasts, ultimately benefiting society, and building trust with forecast users. These positive outcomes highlight the importance of meteorological organizations investing in verification science and technology.

physics.ao-ph

A User-Focused Approach to Evaluating Probabilistic and Categorical Forecasts

A user-focused verification approach for evaluating probability forecasts of binary outcomes (also known as probabilistic classifiers) is demonstrated that is (i) based on proper scoring rules, (ii) focuses on user decision thresholds, and (iii) provides actionable insights. It is argued that when categorical performance diagrams and the critical success index are used to evaluate overall predictive performance, rather than the discrimination ability of probabilistic forecasts, they may produce misleading results. Instead, Murphy diagrams are shown to provide better understanding of overall predictive performance as a function of user probabilistic decision threshold. It is illustrated how to select a proper scoring rule, based on the relative importance of different user decision thresholds, and how this choice impacts scores of overall predictive performance and supporting measures of discrimination and calibration. These approaches and ideas are demonstrated using several probabilistic thunderstorm forecast systems as well as synthetic forecast data. Furthermore, a fair method for comparing the performance of probabilistic and categorical forecasts is illustrated using the FIxed Risk Multicategorical (FIRM) score, which is a proper scoring rule directly connected to values on the Murphy diagram. While the methods are illustrated using thunderstorm forecasts, they are applicable for evaluating probabilistic forecasts for any situation with binary outcomes.

stat.AP

Evaluation of point forecasts for extreme events using consistent scoring functions

We present a method for comparing point forecasts in a region of interest, such as the tails or centre of a variable's range. This method cannot be hedged, in contrast to conditionally selecting events to evaluate and then using a scoring function that would have been consistent (or proper) prior to event selection. Our method also gives decompositions of scoring functions that are consistent for the mean or a particular quantile or expectile. Each member of each decomposition is itself a consistent scoring function that emphasises performance over a selected region of the variable's range. The score of each member of the decomposition has a natural interpretation rooted in optimal decision theory. It is the weighted average of economic regret over user decision thresholds, where the weight emphasises those decision thresholds in the corresponding region of interest.

stat.AP