SearcharxivSearch

arXiv subjects

Christopher D. Roberts

Publications and source records attributed to Christopher D. Roberts.

3 recordsLinked to original sources

AIFS-SUBS: Extending Data-Driven Forecasting to Sub-Seasonal Timescales

Data-driven models now rival numerical weather prediction in the medium range, but extending them to sub-seasonal lead times raises challenges absent at shorter horizons. Errors accumulate over long autoregressive rollouts, systematic biases grow with lead time, and several years of data must be held out for independent verification, even though machine-learning models otherwise benefit from longer training records. To address these challenges, we adapt ECMWF's AIFS-CRPS medium-range model. AIFS-SUBS adopts a 24h autoregressive time step to reduce error accumulation, adds stratospheric levels and top-of-atmosphere thermal radiation as predictors, and reserves 2007--2011 as an independent verification window. We evaluate two config-durations: AIFS-SUBS, fine-tuned on operational analyses, and AIFS-SUBS-ERA5, trained on ERA5 alone. Across weeks 2--6, AIFS-SUBS matches the operational Integrated Forecasting System (IFS) in probabilistic skill while reducing systematic biases. For the convective (OLR) component of the Madden--Julian Oscillation (MJO), AIFS-SUBS extends skilful forecasts (correlation > 0.5) by eight days relative to the IFS, while matching or exceeding the IFS for the full multivariate RMM index. AIFS-SUBS also reproduces the observed MJO modulation of tropical cyclone activity comparably. Stratospheric skill is particularly strong with AIFS-SUBS reproducing sudden stratospheric warming (SSW) frequency and surface impact. In the AI Weather Quest, AIFS-SUBS-ERA5 attains a variable-averaged ranked probability skill score slightly ahead of the IFS at weeks 3 and 4. At inference, AIFS-SUBS uses about 200 times less energy than the IFS, opening the door to much larger real-time ensembles. AIFS-SUBS is ECMWF's first machine-learning model targeted at sub-seasonal time-scales.

physics.ao-ph

Ensemble size effects on conditional reliability estimates: slope attenuation bias and correction methods

The goal of ensemble forecasting is to maximise sharpness subject to reliability. Marginal reliability means that, over all cases, the ensemble is statistically consistent with reality: the ensemble mean is unbiased, the expected ensemble variance equals the expected mean-squared error of the ensemble mean, and the variance of the ensemble members matches the variance of the truth. Equivalently, forecasts that assign probability $p$ to an event verify with relative frequency $p$. However, climatological consistency is not sufficient for users acting on individual forecasts. A natural extension is to assess reliability conditional on the forecast itself, by examining whether, on average, larger ensemble means imply larger observed values, larger spreads imply larger forecast errors, or higher probabilities imply higher event frequencies. This motivates conditional reliability diagnostics such as reliability diagrams and spread-error relationships. Here we show that conditional reliability diagnostics are systematically biased for finite ensemble sizes. We present a unified framework for slope attenuation caused by finite-ensemble sampling noise, which affects conditional diagnostics for ensemble means, spreads, and probabilities. Using synthetic forecasts that are perfectly reliable by construction, we isolate finite-ensemble effects. We derive analytical expressions for the expected attenuation and propose practical estimators computable directly from ensemble data. The framework is illustrated using 2-metre temperature sub-seasonal ensemble forecasts from ECMWF, where finite-ensemble slope attenuation substantially affects the spread-error relationship and tercile-based reliability diagrams. These results demonstrate that attenuated conditional slopes should not be interpreted as evidence of forecast deficiencies unless finite-ensemble effects are explicitly taken into account.

physics.ao-ph

Unbiased calculation, evaluation, and calibration of ensemble forecast anomalies

Long-range ensemble forecasts are typically verified as anomalies with respect to a lead-time dependent climatological mean to remove the influence of systematic biases. However, common methods for calculating anomalies result in statistical inconsistencies between forecast and verification anomalies, even for a perfectly reliable ensemble. It is important to account for these systematic effects when evaluating ensemble forecast systems, particularly when tuning a model to improve the reliability of forecast anomalies or when comparing spread-error diagnostics between systems with different reforecast periods. Here, we show that unbiased variances and spread-error ratios can be recovered by deriving estimators that are consistent with the values that would be achieved when calculating anomalies relative to the true, but unknown, climatological mean. An elegant alternative is to construct forecast climatologies separately for each member, which ensures that forecast and verification anomalies are defined relative to reference climatological means with the same sampling uncertainty. This alternative approach has no impact on forecast ensemble means but systematically modifies the total variance and ensemble spread of forecast anomalies in such a way that anomaly-based spread-error ratios are unbiased without any explicit correction for climatology sample size. Furthermore, the improved statistical consistency of forecast and verification anomalies means that probabilistic forecast skill is optimised when the underlying forecast is also perfectly reliable. Alternative methods for anomaly calculation can thus impact probabilistic forecast skill, especially when anomalies are defined relative to climatologies with a small sample size. Finally, we demonstrate the equivalence of anomalies calculated using different methods after applying an unbiased statistical calibration.

physics.ao-ph