SearcharxivSearch

arXiv subjects

Jonathan Pipping-Gamón

Publications and source records attributed to Jonathan Pipping-Gamón.

6 recordsLinked to original sources

The Blown Lead Paradox: A Pathwise Calibration Benchmark for Win Probability Forecasts

Live win probabilities have given sports collapses a common numerical language: the highest win probability attained by the team that eventually lost. Yet this familiar number is a selected pathwise extreme. The losing path is identified by the terminal outcome and then searched for its most favorable point. Under ideal sequential calibration, we derive an exact continuous-path benchmark for this statistic, a conservative bound for discretely reported paths, and a probability-integral-transform diagnostic for collections of games. Unlike fixed-time calibration summaries and proper scores, the diagnostic asks whether a published feed produces severe losing paths as often as a calibrated sequential model should. We apply the benchmark separately to public regular-season NFL and NBA feeds from 2018-2024, using a season-stratified dyadic bootstrap to account for recurring teams. We detect no global departure in the NFL and a systematic excess of extreme losing-team peaks in the NBA. In the upper 5% benchmark tail, the NBA excess is 1.3% (95% interval 0.4% to 2.3%); first crossings of a published win probability of 0.95 show the same positive gap between observed and implied loss rates.

stat.AP

Dummy RAPM: Representing Low-Minute Players in Regularized Adjusted Plus-Minus

Regularized Adjusted Plus-Minus (RAPM) uses stint-level lineup indicators to estimate player contributions to scoring margin. When low-minute player columns are removed, their stints remain in the data, but the design matrix no longer represents the complete lineup. Dummy RAPM restores this information using five indicators for the number of excluded players on each lineup side. Across 16 NBA seasons, chronological validation selects a 10-minute-per-appearance threshold and a dummy-to-player penalty ratio of 2.2. On held-out March-April games, Dummy RAPM reduces mean season game-margin RMSE from 12.897 to 12.856 and achieves lower RMSE in 13 of 16 seasons. The average reduction is 0.042 points, or 0.30%. Although the improvement in game-level predictive accuracy is small, it is consistent: RAPM performs better when it records how many excluded players are on each side.

stat.AP

Opponent-Adjusted Evaluation of NFL Pass Blocking and Pass Rushing Performance

Evaluating offensive linemen and pass rushers at the player level is difficult because observable outcomes are sparse, opponent-dependent, and strongly shaped by surrounding context. Using 2021 regular-season Hudl tracking data, we construct a blocker-rusher interaction dataset and estimate two ridge-regularized Bradley-Terry paired-comparison models: a binary win/loss model aligned with the 2.5-second pass block win-rate definition and a four-class severity model over loss, win, hit, and sack, with both models incorporating a double-team indicator. The final dataset contains 153,138 interactions across 33,283 pass plays in 266 games. On an ordered 80/20 holdout split (test n = 30,628), both models improve on global baselines and modestly outperform stronger matchup baselines under log-loss evaluation, corresponding to relative log-loss reductions of about 0.24% to 1.21%. Game-level bootstrap resampling indicates that these gains are most stable for the win model and for the severity model relative to the global baseline, while the severity-versus-matchup comparison remains directionally positive but less certain. External comparison to 2021 AP All-Pro selections provides additional face validation on the learned rankings, with the severity model showing the strongest alignment to expert recognition. Overall, ridge-regularized Bradley-Terry models provide an interpretable opponent-adjusted framework for evaluating NFL pass protection and pass rush at the interaction level.

stat.AP

A Unified Server Quality Metric for Tennis

Traditional tennis rating systems (e.g., Elo) summarize overall player strength but do not isolate the independent value of serving. Using point-by-point data from Wimbledon and the U.S.\ Open, we develop serve-specific player metrics that separate serving quality from return ability and other latent factors. For each tournament and gender, we fit logistic mixed-effects models of point outcomes using serve speed, speed variability, and placement features, with crossed server and returner random intercepts to capture unobserved player strengths. From these models we derive Server Quality Scores (SQS): partially pooled, opponent-adjusted estimates of players' serving impact. In out-of-sample evaluation, SQS aligns more strongly with serve efficiency$\unicode{x2014}$the probability of winning points within three shots$\unicode{x2014}$than weighted Elo. We further benchmark SQS against task-aligned serve-stat baselines and model ablations, quantifying the incremental value of serve features and partial pooling. Associations with overall serve win percentage are smaller and dataset-dependent, and neither SQS nor weighted Elo consistently dominates that outcome. Overall, SQS is best interpreted as a measure of serve-induced short-point advantage (serve quality plus early-point conversion), complementing holistic ratings with actionable insight for coaching, forecasting, and player evaluation.

stat.AP

Beyond Expected Goals: A Probabilistic Framework for Shot Occurrences in Soccer

Expected goals (xG) models estimate the probability that a shot results in a goal from its context (e.g., location, pressure), but they operate only on observed shots. We propose xG+, a possession-level framework that first estimates the probability that a shot occurs within the next second and its corresponding xG if it were to occur. We also introduce ways to aggregate this joint probability estimate over the course of a possession. By jointly modeling shot-taking behavior and shot quality, xG+ remedies the conditioning-on-shots limitation of standard xG. We show that this improves predictive accuracy at the team level and produces a more persistent player skill signal than standard xG models.

stat.AP

Kicking for Goal or Touch? An Expected Points Framework for Penalty Decisions in Rugby Union

Following a penalty in rugby union, teams typically choose between attempting a shot at goal or kicking to touch to pursue a try. We develop an Expected Points (EP) framework that quantifies the value of each option as a function of both field location and game context. Using phase-level data from the 2018/19 Premiership Rugby season (35,199 phases across 132 matches) and an angle-distance model of penalty kick success estimated from international records, we construct two surfaces: (i) the expected points of a possession beginning with a lineout, and (ii) the expected points of a kick at goal, taking into account the in-game consequences of made and missed kicks. We then compare these surfaces to produce decision maps that indicate where kicking for goal or kicking to touch maximizes expected return, and we analyze how the boundary shifts with game context and the expected meters gained to touch. Our results provide a unified, data-driven method for evaluating penalty decisions and can be tailored to team-specific kickers and lineout units. This study offers, to our knowledge, the first comprehensive EP-based assessment of penalty strategy in rugby union and outlines extensions to win-probability analysis and richer tracking data.

stat.AP