Searcharxiv⌕ Search

arXiv · 2610.09898

Learning joint probabilistic weather forecasts from station observations alone

Abstract

Assessing compound weather risks requires forecasts representing dependence between variables. CLARA (Calibrated Advection-Routing Attention) learns joint Gaussian predictive distributions of five surface variables from station observations alone, without numerical weather prediction or reanalysis; the approximately 28,000-parameter model supports CPU training and prediction. Across six multi-year folds on 96 stations, its lead-mean energy score is 4.9% lower than that of a learned comparator with matched temporal inputs (4.7% with a similar parameter count) and 11-65% lower than those of statistical baselines. Holding marginal variances fixed, removing learned correlations worsens joint negative log-likelihood by 1.0-2.8 nats per station. A covariance-scale estimator, proved consistent under stated assumptions, improves short-lead calibration but over-corrects at long leads. Synthetic interventions show an attention-bias coefficient alone does not measure forecast influence. Retrained in ten regions on six continents, CLARA outperforms persistence in all 60 multi-year region-lead comparisons and a similarly sized learned model in 57 of 60.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chaeyeon Yi, Yun Am Seo. 2026-10-07. Learning joint probabilistic weather forecasts from station observations alone. https://arxiv.org/abs/2610.09898

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

BayesJudge: Uncertainty-Aware Bayesian Meta-Evaluation of Human and LLM Judgments

AI evaluation pipelines often produce conflicting judgments rather than clean labels. In pairwise LLM evaluation, this conflict is especially visible: disagreement can arise from ambiguous items, underspecified rubrics, heterogeneous or unstable human raters, or an LLM judge whose verdict changes when the response order is swapped. We propose BayesJudge, an online Bayesian meta-evaluation layer for conflicting human-LLM judgment streams. For each comparison, BayesJudge estimates a panel-relative posterior verdict distribution over the two responses, with a tie or ambiguity state when such labels are available. At the same time, it estimates rater-specific human confusion matrices and LLM presentation-order bias. The method uses tie-open labels to keep ambiguity observable and paired order-swapped judge calls to separate response quality from presentation effects. We formulate the exact online posterior recursion and use a scalable Rao-Blackwellized assumed-density SMC approximation for streaming inference. Controlled synthetic experiments demonstrate recovery of prespecified evaluator parameters and illustrate two protocol-level identifiability mechanisms: tie-open labels expose ambiguity mass, and order-swapped paired judgments separate item preference from position bias. On a real-world SummEval dataset, BayesJudge successfully detects systematic presentation-order effects in LLM judge outputs, infers distinct expert and crowdworker behavior signatures without rater metadata, and produces posterior uncertainty estimates that correlate with human disagreement. Our code is available at https://anonymous.4open.science/r/BayesJudge-0879.

stat.AP↗

Event-Aware Spatiotemporal Precipitation Forecasting with Geographic Context and Physics-Guided Regularization

Hourly precipitation forecasting involves several distinct statistical challenges, including spatially varying predictor-precipitation relationships, a strongly imbalanced precipitation distribution, and progressive degradation of precipitation event skill with increasing lead time. We develop a spatiotemporal forecasting framework that addresses these challenges through three complementary components: explicit geographic representation, event-aware learning for imbalanced precipitation, and weak asymmetric regularization derived from the atmospheric water budget. The physical information is treated as an asymmetric constraint rather than as an additional prediction target, designed to discourage physically unsupported precipitation attenuation without replacing the data-driven forecast. Experiments using ERA5 data across regional, enlarged domain, and spatial subset settings show that explicit geographic information improves spatial field prediction, while event-aware learning provides the most consistent gains in detecting moderate and heavy precipitation events. Physical regularization has a more selective effect, mainly reducing systematic underprediction over the enlarged domain while improving longer lead precipitation event prediction in the spatial subset experiment. These results indicate that the benefit of physical guidance depends on the available data regime and becomes most apparent when data-driven precipitation information deteriorates with increasing lead time.

stat.AP↗

Opponent-Adjusted Evaluation of NFL Pass Protection and Pass Rush

Evaluating NFL pass protectors and rushers is difficult because observed win rates reflect both player performance and the difficulty of the blocking assignments each player faces. Using Hudl player coordinates and blocking annotations from the 2021 NFL regular season, we estimate blocker and rusher strength jointly from initial blocking matchups. A ridge-regularized Bradley--Terry model uses early pass rush wins and losses within 2.5 seconds of the snap, while a multinomial extension models full-play losses, wins, pressures, and sacks. To summarize full-play performance, we use the outcomes' conditional associations with offensive expected points added (EPA) to weight predicted probabilities, giving ratings of expected disruption produced by rushers or allowed by blockers. In held-out Weeks 16--18, the early and final models reduce log loss by 4.5% and 3.6%, respectively, relative to historical matchup baselines whose smoothing strengths are selected by cross-validation; 95% game-bootstrap intervals for both reductions are above zero. The resulting ratings place observed success in the context of opponents and protection help, with game-bootstrap intervals quantifying sampling uncertainty in early strength and full-play disruption.

stat.AP↗