Searcharxiv⌕ Search

arXiv subjects

Abraham J. Wyner

Publications and source records attributed to Abraham J. Wyner.

At least 19 recordsLinked to original sources

The Blown Lead Paradox: A Pathwise Calibration Benchmark for Win Probability Forecasts

Live win probabilities have given sports collapses a common numerical language: the highest win probability attained by the team that eventually lost. Yet this familiar number is a selected pathwise extreme. The losing path is identified by the terminal outcome and then searched for its most favorable point. Under ideal sequential calibration, we derive an exact continuous-path benchmark for this statistic, a conservative bound for discretely reported paths, and a probability-integral-transform diagnostic for collections of games. Unlike fixed-time calibration summaries and proper scores, the diagnostic asks whether a published feed produces severe losing paths as often as a calibrated sequential model should. We apply the benchmark separately to public regular-season NFL and NBA feeds from 2018-2024, using a season-stratified dyadic bootstrap to account for recurring teams. We detect no global departure in the NFL and a systematic excess of extreme losing-team peaks in the NBA. In the upper 5% benchmark tail, the NBA excess is 1.3% (95% interval 0.4% to 2.3%); first crossings of a published win probability of 0.95 show the same positive gap between observed and implied loss rates.

stat.AP↗

Dummy RAPM: Representing Low-Minute Players in Regularized Adjusted Plus-Minus

Regularized Adjusted Plus-Minus (RAPM) uses stint-level lineup indicators to estimate player contributions to scoring margin. When low-minute player columns are removed, their stints remain in the data, but the design matrix no longer represents the complete lineup. Dummy RAPM restores this information using five indicators for the number of excluded players on each lineup side. Across 16 NBA seasons, chronological validation selects a 10-minute-per-appearance threshold and a dummy-to-player penalty ratio of 2.2. On held-out March-April games, Dummy RAPM reduces mean season game-margin RMSE from 12.897 to 12.856 and achieves lower RMSE in 13 of 16 seasons. The average reduction is 0.042 points, or 0.30%. Although the improvement in game-level predictive accuracy is small, it is consistent: RAPM performs better when it records how many excluded players are on each side.

stat.AP↗

Opponent-Adjusted Evaluation of NFL Pass Blocking and Pass Rushing Performance

Evaluating offensive linemen and pass rushers at the player level is difficult because observable outcomes are sparse, opponent-dependent, and strongly shaped by surrounding context. Using 2021 regular-season Hudl tracking data, we construct a blocker-rusher interaction dataset and estimate two ridge-regularized Bradley-Terry paired-comparison models: a binary win/loss model aligned with the 2.5-second pass block win-rate definition and a four-class severity model over loss, win, hit, and sack, with both models incorporating a double-team indicator. The final dataset contains 153,138 interactions across 33,283 pass plays in 266 games. On an ordered 80/20 holdout split (test n = 30,628), both models improve on global baselines and modestly outperform stronger matchup baselines under log-loss evaluation, corresponding to relative log-loss reductions of about 0.24% to 1.21%. Game-level bootstrap resampling indicates that these gains are most stable for the win model and for the severity model relative to the global baseline, while the severity-versus-matchup comparison remains directionally positive but less certain. External comparison to 2021 AP All-Pro selections provides additional face validation on the learned rankings, with the severity model showing the strongest alignment to expert recognition. Overall, ridge-regularized Bradley-Terry models provide an interpretable opponent-adjusted framework for evaluating NFL pass protection and pass rush at the interaction level.

stat.AP↗

Separating Intent from Execution: A Probabilistic Approach to Pitch Location Accuracy

Control has long been recognized as a critical component of pitcher performance, reflecting a pitcher's ability to execute pitches in alignment with his intended targets. However, accurately inferring a pitcher's intentions presents a persistent challenge. Traditional metrics typically rely on uniformity assumptions, inferring intent based on the behavior of a ``typical'' pitcher across similar situations. In this study, we propose an alternative, individualized approach to measuring control, one that eschews such assumptions in favor of personalized inference. We estimate a pitcher's intended location on a pitch-by-pitch basis, conditioning on both individual tendencies and specific game contexts. This allows us to assess control by comparing the actual pitch location to the inferred intended target, thereby aligning measurement more closely with the unique strategies of each pitcher. We introduce xCTRL, a novel metric that quantifies control as the distance between a pitch's actual location and its estimated intended location. We find that xCTRL exhibits strong stability and greater predictive power than existing control metrics. By capturing pitcher-specific intent, xCTRL enhances our understanding of control and offers a more intuitive and accurate representation of pitching performance.

stat.AP↗

Exploring the Difficulty of Estimating Win Probability: A Simulation Study

Estimating win probability is one of the classic modeling tasks of sports analytics. Many widely used win probability estimators use machine learning to fit the relationship between a binary win/loss outcome variable and certain game-state variables. To illustrate just how difficult it is to accurately fit such a model from noisy and highly correlated observational data, in this paper we conduct a simulation study. We create a simplified random walk version of football in which true win probability at each game-state is known, and we see how well a model recovers it. We find that the dependence structure of observational play-by-play data substantially inflates the bias and variance of estimators and lowers the effective sample size. Further, to achieve approximately valid marginal coverage, win probability confidence intervals need to be substantially wide. Concisely, these are high variance estimators subject to substantial uncertainty. Our findings are not unique to the particular application of estimating win probability; they are broadly applicable across sports analytics, as myriad other sports datasets are clustered into groups of observations that share the same outcome.

stat.ME↗

Putting Skill as Nearly Indistinguishable from Noise: An Empirical Bayes Analysis of PGA Tour Performance

We revisit a foundational question in golf analytics: how important are the core components of performance--driving, approach play, and putting--in explaining success on the PGA Tour? Building on Mark Broadie's strokes gained analyses, we use an empirical Bayes approach to estimate latent golfer skill and assess statistical significance using a multiple testing procedure that controls the false discovery rate. While tee-to-green skill shows clear and substantial differences across players, putting skill is both less variable and far less reliably estimable. Indeed, putting performance appears nearly indistinguishable from noise.

stat.AP↗

The Loser's Curse and the Critical Role of the Utility Function

A longstanding question in the judgment and decision making literature is whether experts, even in high-stakes environments, exhibit the same cognitive biases observed in controlled experiments with inexperienced participants. Massey and Thaler (2013) claim to have found an example of bias and irrationality in expert decision making: general managers' behavior in the National Football League draft pick trade market. They argue that general managers systematically overvalue top draft picks, which generate less surplus value on average than later first-round picks, a phenomenon known as the loser's curse. Their conclusion hinges on the assumption that general managers should use expected surplus value as their utility function for evaluating draft picks. This assumption, however, is neither explicitly justified nor necessarily aligned with the strategic complexities of constructing a National Football League roster. In this paper, we challenge their framework by considering alternative utility functions, particularly those that emphasize the acquisition of transformational players--those capable of dramatically increasing a team's chances of winning the Super Bowl. Under a decision rule that prioritizes the probability of acquiring elite players, which we construct from a novel Bayesian hierarchical Beta regression model, general managers' draft trade behavior appears rational rather than systematically flawed. More broadly, our findings highlight the critical role of carefully specifying a utility function when evaluating the quality of decisions.

stat.AP↗

Analytics, have some humility: a statistical view of fourth-down decision making

The standard mathematical approach to fourth-down decision making in American football is to make the decision that maximizes estimated win probability. Win probability estimates arise from machine learning models fit from historical data. These models attempt to capture a nuanced relationship between a noisy binary outcome variable and game-state variables replete with interactions and non-linearities from a finite dataset of just a few thousand games. Thus, it is imperative to knit uncertainty quantification into the fourth-down decision procedure; we do so using bootstrapping. We find that uncertainty in the estimated optimal fourth-down decision is far greater than that currently expressed by sports analysts in popular sports media.

stat.AP↗

Moving from Machine Learning to Statistics: the case of Expected Points in American football

Expected points is a value function fundamental to player evaluation and strategic in-game decision-making across sports analytics, particularly in American football. To estimate expected points, football analysts use machine learning tools, which are not equipped to handle certain challenges. They suffer from selection bias, display counter-intuitive artifacts of overfitting, do not quantify uncertainty in point estimates, and do not account for the strong dependence structure of observational football data. These issues are not unique to American football or even sports analytics; they are general problems analysts encounter across various statistical applications, particularly when using machine learning in lieu of traditional statistical models. We explore these issues in detail and devise expected points models that account for them. We also introduce a widely applicable novel methodological approach to mitigate overfitting, using a catalytic prior to smooth our machine learning models.

stat.AP↗

Entropy-Based Strategies for Multi-Bracket Pools

Much work in the parimutuel betting literature has discussed estimating event outcome probabilities or developing optimal wagering strategies, particularly for horse race betting. Some betting pools, however, involve betting not just on a single event, but on a tuple of events. For example, pick six betting in horse racing, March Madness bracket challenges, and predicting a randomly drawn bitstring each involve making a series of individual forecasts. Although traditional optimal wagering strategies work well when the size of the tuple is very small (e.g., betting on the winner of a horse race), they are intractable for more general betting pools in higher dimensions (e.g., March Madness bracket challenges). Hence we pose the multi-brackets problem: supposing we wish to predict a tuple of events and that we know the true probabilities of each potential outcome of each event, what is the best way to tractably generate a set of $n$ predicted tuples? The most general version of this problem is extremely difficult, so we begin with a simpler setting. In particular, we generate $n$ independent predicted tuples according to a distribution having optimal entropy. This entropy-based approach is tractable, scalable, and performs well.

cs.GT↗

Introducing Grid WAR: Rethinking WAR for Starting Pitchers

The baseball statistic "Wins Above Replacement" (WAR) has emerged as one of the most popular evaluation metrics. But it is not readily observed and tabulated; WAR is an estimate of a parameter in a vaguely defined model with all its attendant assumptions. Industry-standard models of WAR for starting pitchers from FanGraphs and Baseball Reference all assume that season-long averages are sufficient statistics for a pitcher's performance. This provides an invalid mathematical foundation for many reasons, especially because WAR should not be linear with respect to any counting statistic. To repair this defect, as well as many others, we devise a new measure, Grid WAR, which accurately estimates a starting pitcher's WAR on a per-game basis. The convexity of Grid WAR diminishes the impact of "blow-up" games and upweights exceptional games, raising the valuation of pitchers like Sandy Koufax, Whitey Ford, and Catfish Hunter who exhibit fundamental game-by-game variance. Grid WAR is designed to accurately measure past performance, but also has predictive value insofar as a pitcher's Grid WAR is better than WAR at predicting future performance. Finally, at https://gridwar.xyz we host a Shiny app which displays the Grid WAR results of each MLB game since 1952, including career, season, and game level results, which updates automatically every morning.

cs.GT↗

A Bayesian analysis of the time through the order penalty in baseball

As a baseball game progresses, batters appear to perform better the more times they face a particular pitcher. The apparent drop-off in pitcher performance from one time through the order to the next, known as the Time Through the Order Penalty (TTOP), is often attributed to within-game batter learning. Although the TTOP has largely been accepted within baseball and influences many managers' in-game decision making, we argue that existing approaches of estimating the size of the TTOP cannot disentangle continuous evolution in pitcher performance over the course of the game from discontinuities between successive times through the order. Using a Bayesian multinomial regression model, we find that, after adjusting for confounders like batter and pitcher quality, handedness, and home field advantage, there is little evidence of strong discontinuity in pitcher performance between times through the order. Our analysis suggests that the start of the third time through the order should not be viewed as a special cutoff point in deciding whether to pull a starting pitcher.

stat.AP↗

Making Sense of Random Forest Probabilities: a Kernel Perspective

A random forest is a popular tool for estimating probabilities in machine learning classification tasks. However, the means by which this is accomplished is unprincipled: one simply counts the fraction of trees in a forest that vote for a certain class. In this paper, we forge a connection between random forests and kernel regression. This places random forest probability estimation on more sound statistical footing. As part of our investigation, we develop a model for the proximity kernel and relate it to the geometry and sparsity of the estimation problem. We also provide intuition and recommendations for tuning a random forest to improve its probability estimates.

stat.ML↗

Explaining the Success of AdaBoost and Random Forests as Interpolating Classifiers

There is a large literature explaining why AdaBoost is a successful classifier. The literature on AdaBoost focuses on classifier margins and boosting's interpretation as the optimization of an exponential likelihood function. These existing explanations, however, have been pointed out to be incomplete. A random forest is another popular ensemble method for which there is substantially less explanation in the literature. We introduce a novel perspective on AdaBoost and random forests that proposes that the two algorithms work for similar reasons. While both classifiers achieve similar predictive accuracy, random forests cannot be conceived as a direct optimization procedure. Rather, random forests is a self-averaging, interpolating algorithm which creates what we denote as a "spikey-smooth" classifier, and we view AdaBoost in the same light. We conjecture that both AdaBoost and random forests succeed because of this mechanism. We provide a number of examples and some theoretical justification to support this explanation. In the process, we question the conventional wisdom that suggests that boosting algorithms for classification require regularization or early stopping and should be limited to low complexity classes of learners, such as decision stumps. We conclude that boosting should be used like random forests: with large decision trees and without direct regularization or early stopping.

stat.ML↗

Estimating an NBA player's impact on his team's chances of winning

Traditional NBA player evaluation metrics are based on scoring differential or some pace-adjusted linear combination of box score statistics like points, rebounds, assists, etc. These measures treat performances with the outcome of the game still in question (e.g. tie score with five minutes left) in exactly the same way as they treat performances with the outcome virtually decided (e.g. when one team leads by 30 points with one minute left). Because they ignore the context in which players perform, these measures can result in misleading estimates of how players help their teams win. We instead use a win probability framework for evaluating the impact NBA players have on their teams' chances of winning. We propose a Bayesian linear regression model to estimate an individual player's impact, after controlling for the other players on the court. We introduce several posterior summaries to derive rank-orderings of players within their team and across the league. This allows us to identify highly paid players with low impact relative to their teammates, as well as players whose high impact is not captured by existing metrics.

stat.AP↗

A statistical analysis of multiple temperature proxies: Are reconstructions of surface temperatures over the last 1000 years reliable?

Predicting historic temperatures based on tree rings, ice cores, and other natural proxies is a difficult endeavor. The relationship between proxies and temperature is weak and the number of proxies is far larger than the number of target data points. Furthermore, the data contain complex spatial and temporal dependence structures which are not easily captured with simple models. In this paper, we assess the reliability of such reconstructions and their statistical significance against various null models. We find that the proxies do not predict temperature significantly better than random series generated independently of temperature. Furthermore, various model specifications that perform similarly at predicting temperature produce extremely different historical backcasts. Finally, the proxies seem unable to forecast the high levels of and sharp run-up in temperature in the 1990s either in-sample or from contiguous holdout blocks, thus casting doubt on their ability to predict such phenomena if in fact they occurred several hundred years ago. We propose our own reconstruction of Northern Hemisphere average annual land temperature over the last millennium, assess its reliability, and compare it to those from the climate science literature. Our model provides a similar reconstruction but has much wider standard errors, reflecting the weak signal and large uncertainty encountered in this setting.

stat.AP↗

Bayesball: A Bayesian hierarchical model for evaluating fielding in major league baseball

The use of statistical modeling in baseball has received substantial attention recently in both the media and academic community. We focus on a relatively under-explored topic: the use of statistical models for the analysis of fielding based on high-resolution data consisting of on-field location of batted balls. We combine spatial modeling with a hierarchical Bayesian structure in order to evaluate the performance of individual fielders while sharing information between fielders at each position. We present results across four seasons of MLB data (2002--2005) and compare our approach to other fielding evaluation procedures.

stat.AP↗