SearcharxivSearch

arXiv subjects

Ryan Elmore

Publications and source records attributed to Ryan Elmore.

6 recordsLinked to original sources

Expected Points Above Average: A Novel NBA Player Metric Based on Bayesian Hierarchical Modeling

In this paper, we propose two novel basketball metrics: ``expected points'' for team-based comparisons and ``expected points above average (EPAA)'' as a player-evaluation tool. Established within the Bayesian hierarchical model framework, teams and players are clustered based on their shooting propensities and abilities using posterior predictive distributions. We illustrate the concepts for the top 100 shot takers over the last decade and offer our metric as an additional metric for evaluating players. We compare our metrics to two traditional NBA player evaluation metrics: player efficiency rating and box plus/minus. Finally, we develop a Shiny web application that allows interested readers to make additional team and player comparisons.

stat.OT

Simulation-Based Decision Making in the NFL using NFLSimulatoR

In this paper, we introduce an R software package for simulating plays and drives using play-by-play data from the National Football League. The simulations are generated by sampling play-by-play data from previous football seasons.The sampling procedure adds statistical rigor to any decisions or inferences arising from examining the simulations. We highlight that the package is particularly useful as a data-driven tool for evaluating potential in-game strategies or rule changes within the league. We demonstrate its utility by evaluating the oft-debated strategy of $\textit{going for it}$ on fourth down and investigating whether or not teams should pass more than the current standard.

stat.AP

The causal effect of a timeout at stopping an opposing run in the NBA

In the summer of 2017, the National Basketball Association reduced the number of total timeouts, along with other rule changes, to regulate the flow of the game. With these rule changes, it becomes increasingly important for coaches to effectively manage their timeouts. Understanding the utility of a timeout under various game scenarios, e.g., during an opposing team's run, is of the utmost importance. There are two schools of thought when the opposition is on a run: (1) call a timeout and allow your team to rest and regroup, or (2) save a timeout and hope your team can make corrections during play. This paper investigates the credence of these tenets using the Rubin causal model framework to quantify the causal effect of a timeout in the presence of an opposing team's run. Too often overlooked, we carefully consider the stable unit-treatment-value assumption (SUTVA) in this context and use SUTVA to motivate our definition of units. To measure the effect of a timeout, we introduce a novel, interpretable outcome based on the score difference to describe broad changes in the scoring dynamics. This outcome is well-suited for situations where the quantity of interest fluctuates frequently, a commonality in many sports analytics applications. We conclude from our analysis that while comebacks frequently occur after a run, it is slightly disadvantageous to call a timeout during a run by the opposing team and further demonstrate that the magnitude of this effect varies by franchise.

stat.AP

Modeling Sums of Exchangeable Binary Variables

We introduce a new model for sums of exchangeable binary random variables. The proposed distribution is an approximation to the exact distributional form, and relies on the theory of completely monotone functions and the Laplace transform of a gamma distribution function. Using Monte Carlo methods, we show that this new model compares favorably to the beta binomial model with respect to estimating the success probability of the Bernoulli trials and the correlation between any two variables in the exchangeable set. We apply the new methodology to two classic data sets and the results are summarized.

stat.ME

Prioritized Data Compression using Wavelets

The volume of data and the velocity with which it is being generated by com- putational experiments on high performance computing (HPC) systems is quickly outpacing our ability to effectively store this information in its full fidelity. There- fore, it is critically important to identify and study compression methodologies that retain as much information as possible, particularly in the most salient regions of the simulation space. In this paper, we cast this in terms of a general decision-theoretic problem and discuss a wavelet-based compression strategy for its solution. We pro- vide a heuristic argument as justification and illustrate our methodology on several examples. Finally, we will discuss how our proposed methodology may be utilized in an HPC environment on large-scale computational experiments.

stat.CO

Keeping greed good: sparse regression under design uncertainty with application to biomass characterization

In this paper, we consider the classic measurement error regression scenario in which our independent, or design, variables are observed with several sources of additive noise. We will show that our motivating example's replicated measurements on both the design and dependent variables may be leveraged to enhance a sparse regression algorithm. Specifically, we estimate the variance and use it to scale our design variables. We demonstrate the efficacy of scaling from several points of view and validate it empirically with a biomass characterization data set using two of the most widely used sparse algorithms: least angle regression (LARS) and the Dantzig selector (DS).

stat.AP