SearcharxivSearch

arXiv subjects

Shane T. Jensen

Publications and source records attributed to Shane T. Jensen.

At least 19 recordsLinked to original sources

The Hall of Fame Cut in Major League Baseball

I present a simple and transparent standard for career greatness in baseball: any major league player with H > 2500 or HR > 350 or K > 2800 or W > 240 makes my Hall of Fame Cut. Rate statistics are avoided due to small sample issues and to ensure the standard is permanent once achieved. Hits and home runs were chosen to represent the two extremes of batting styles. Strikeouts are chosen as the most fundamental unit of pitching performance whereas wins are included in deference to their historical importance as a benchmark. Most major league batters and pitchers in the elected Hall of Fame also make my Hall of Fame Cut but my quantitative standard shifts attention to several under-appreciated players, such as Johnny Damon and Bartolo Colon, and allows us to celebrate recent and active players without the waiting period (5 years post-retirement) needed for Hall of Fame election. My Hall of Fame Cut is also agnostic to performance enhancement or off-field issues and strongly favors longevity over peak performance.

stat.OT

Spatial Analysis of the Association between School Proximity and Crime in Philadelphia

We use high resolution data to investigate the association between crime incidence and proximity to different types of public schools over the past fifteen years in the city of Philadelphia. We employ two statistical methods, regression modeling and propensity score matching, in order to better isolate the association between crime and school proximity while controlling for the demographic, economic, land use and disorder characteristics of the surrounding neighborhood. With both of these approaches, we find significantly increased crime incidence near to public schools regardless of crime outcome, educational level and time period. The effect of school proximity on crime varies substantially depending on whether or not school is in session, as well as between different types of crime and educational levels of the school. We see the largest effects of school proximity on crime for violent crimes near to high schools during their in-session time periods. Our results support several theories which suggest that crime should be elevated near to schools, as well as finding significant associations between crime and other aspects of the built environment.

stat.AP

Clustering Areal Units at Multiple Levels of Resolution to Model Crime in Philadelphia

Estimation of the spatial heterogeneity in crime incidence across an entire city is an important step towards reducing crime and increasing our understanding of the physical and social functioning of urban environments. This is a difficult modeling endeavor since crime incidence can vary smoothly across space and time but there also exist physical and social barriers that result in discontinuities in crime rates between different regions within a city. A further difficulty is that there are different levels of resolution that can be used for defining regions of a city in order to analyze crime. To address these challenges, we develop a Bayesian non-parametric approach for the clustering of urban areal units at different levels of resolution simultaneously. Our approach is evaluated with an extensive synthetic data study and then applied to the estimation of crime incidence at various levels of resolution in the city of Philadelphia.

stat.ME

Crime in Philadelphia: Bayesian Clustering with Particle Optimization

Accurate estimation of the change in crime over time is a critical first step towards better understanding of public safety in large urban environments. Bayesian hierarchical modeling is a natural way to study spatial variation in urban crime dynamics at the neighborhood level, since it facilitates principled ``sharing of information'' between spatially adjacent neighborhoods. Typically, however, cities contain many physical and social boundaries that may manifest as spatial discontinuities in crime patterns. In this situation, standard prior choices often yield overly-smooth parameter estimates, which can ultimately produce mis-calibrated forecasts. To prevent potential over-smoothing, we introduce a prior that partitions the set of neighborhoods into several clusters and encourages spatial smoothness within each cluster. In terms of model implementation, conventional stochastic search techniques are computationally prohibitive, as they must traverse a combinatorially vast space of partitions. We introduce an ensemble optimization procedure that simultaneously identifies several high probability partitions by solving one optimization problem using a new local search strategy. We then use the identified partitions to estimate crime trends in Philadelphia between 2006 and 2017. On simulated and real data, our proposed method demonstrates good estimation and partition selection performance.

stat.AP

Bayesian Learning of Play Styles in Multiplayer Video Games

The complexity of game play in online multiplayer games has generated strong interest in modeling the different play styles or strategies used by players for success. We develop a hierarchical Bayesian regression approach for the online multiplayer game Battlefield 3 where performance is modeled as a function of the roles, game type, and map taken on by that player in each of their matches. We use a Dirichlet process prior that enables the clustering of players that have similar player-specific coefficients in our regression model, which allows us to discover common play styles amongst our sample of Battlefield 3 players. This Bayesian semi-parametric clustering approach has several advantages: the number of common play styles do not need to be specified, players can move between multiple clusters, and the resulting groupings often have a straight-forward interpretations. We examine the most common play styles among Battlefield 3 players in detail and find groups of players that exhibit overall high performance, as well as groupings of players that perform particularly well in specific game types, maps and roles. We are also able to differentiate between players that are stable members of a particular play style from hybrid players that exhibit multiple play styles across their matches. Modeling this landscape of different play styles will aid game developers in developing specialized tutorials for new participants as well as improving the construction of complementary teams in their online matching queues.

cs.LG

The Effects of Vacant Lot Greening and the Impact of Land Use and Business Presence on Crime

We examine the effect of the Philadelphia LandCare (PLC) vacant lot greening initiative on crime and the extent to which surrounding land uses and business types moderate this intervention. We rely on a propensity score matching analysis to account for substantial differences in demographic, economic, land use, and business characteristics between greened and ungreened vacant lots. We estimate larger and more significant crime reductions around vacant lots that are greened in our matched pairs analysis compared to unmatched analyses. The effects of vacant lot greening on crime are larger in areas with high residential and low commercial land use and are moderated by the presence of different types of nearby businesses.

stat.OT

Community Vibrancy and its Relationship with Safety in Philadelphia

To what extent can the strength of a local urban community impact neighborhood safety? We construct measures of community vibrancy based on a unique dataset of block party permit approvals from the City of Philadelphia. Our first measure captures the overall volume of block party events in a neighborhood whereas our second measure captures differences in the type (regular versus spontaneous) of block party activities. We use both regression modeling and propensity score matching to control for the economic, demographic and land use characteristics of the surrounding neighborhood when examining the relationship between crime and our two measures of community vibrancy. We conduct our analysis on aggregate levels of crime and community vibrancy from 2006 to 2015 as well as the trends in community vibrancy and crime over this time period. We find that neighborhoods with a higher number of block parties have a significantly higher crime rate, while those holding a greater proportion of spontaneous block party events have a significantly lower crime rate. We also find that neighborhoods which have an increase in the proportion of spontaneous block parties over time are significantly more likely to have a decreasing trend in total crime incidence over that same time period.

stat.AP

Spatial Modeling of Trends in Crime over Time in Philadelphia

Understanding the relationship between change in crime over time and the geography of urban areas is an important problem for urban planning. Accurate estimation of changing crime rates throughout a city would aid law enforcement as well as enable studies of the association between crime and the built environment. Bayesian modeling is a promising direction since areal data require principled sharing of information to address spatial autocorrelation between proximal neighborhoods. We develop several Bayesian approaches to spatial sharing of information between neighborhoods while modeling trends in crime counts over time. We apply our methodology to estimate changes in crime throughout Philadelphia over the 2006-15 period, while also incorporating spatially-varying economic and demographic predictors. We find that the local shrinkage imposed by a conditional autoregressive model has substantial benefits in terms of out-of-sample predictive accuracy of crime. We also explore the possibility of spatial discontinuities between neighborhoods that could represent natural barriers or aspects of the built environment.

stat.AP

Urban Vibrancy and Safety in Philadelphia

Statistical analyses of urban environments have been recently improved through publicly available high resolution data and mapping technologies that have been adopted across industries. These technologies allow us to create metrics to empirically investigate urban design principles of the past half-century. Philadelphia is an interesting case study for this work, with its rapid urban development and population increase in the last decade. We outline a data analysis pipeline for exploring the association between safety and local neighborhood features such as population, economic health and the built environment. As a particular example of our analysis pipeline, we focus on quantitative measures of the built environment that serve as proxies for vibrancy: the amount of human activity in a local area. Historically, vibrancy has been very challenging to measure empirically. Measures based on land use zoning are not an adequate description of local vibrancy and so we construct a database and set of measures of business activity in each neighborhood. We employ several matching analyses to explore the relationship between neighborhood vibrancy and safety, such as comparing high crime versus low crime locations within the same neighborhood. As additional sources of urban data become available, our analysis pipeline can serve as the template for further investigations into the relationships between safety, economic factors and the built environment at the local neighborhood level.

stat.AP

Tree-Structured Boosting: Connections Between Gradient Boosted Stumps and Full Decision Trees

Additive models, such as produced by gradient boosting, and full interaction models, such as classification and regression trees (CART), are widely used algorithms that have been investigated largely in isolation. We show that these models exist along a spectrum, revealing never-before-known connections between these two approaches. This paper introduces a novel technique called tree-structured boosting for creating a single decision tree, and shows that this method can produce models equivalent to CART or gradient boosted stumps at the extremes by varying a single parameter. Although tree-structured boosting is designed primarily to provide both the model interpretability and predictive performance needed for high-stake applications like medicine, it also can produce decision trees represented by hybrid models between CART and boosted stumps that can outperform either of these approaches.

stat.ML

Partial Information Framework: Model-Based Aggregation of Estimates from Diverse Information Sources

Prediction polling is an increasingly popular form of crowdsourcing in which multiple participants estimate the probability or magnitude of some future event. These estimates are then aggregated into a single forecast. Historically, randomness in scientific estimation has been generally assumed to arise from unmeasured factors which are viewed as measurement noise. However, when combining subjective estimates, heterogeneity stemming from differences in the participants' information is often more important than measurement noise. This paper formalizes information diversity as an alternative source of such heterogeneity and introduces a novel modeling framework that is particularly well-suited for prediction polls. A practical specification of this framework is proposed and applied to the task of aggregating probability and point estimates from two real-world prediction polls. In both cases our model outperforms standard measurement-error-based aggregators, hence providing evidence in favor of information diversity being the more important source of heterogeneity.

stat.ME

Estimating an NBA player's impact on his team's chances of winning

Traditional NBA player evaluation metrics are based on scoring differential or some pace-adjusted linear combination of box score statistics like points, rebounds, assists, etc. These measures treat performances with the outcome of the game still in question (e.g. tie score with five minutes left) in exactly the same way as they treat performances with the outcome virtually decided (e.g. when one team leads by 30 points with one minute left). Because they ignore the context in which players perform, these measures can result in misleading estimates of how players help their teams win. We instead use a win probability framework for evaluating the impact NBA players have on their teams' chances of winning. We propose a Bayesian linear regression model to estimate an individual player's impact, after controlling for the other players on the court. We introduce several posterior summaries to derive rank-orderings of players within their team and across the league. This allows us to identify highly paid players with low impact relative to their teammates, as well as players whose high impact is not captured by existing metrics.

stat.AP

Power Weighted Densities for Time Series Data

While time series prediction is an important, actively studied problem, the predictive accuracy of time series models is complicated by non-stationarity. We develop a fast and effective approach to allow for non-stationarity in the parameters of a chosen time series model. In our power-weighted density (PWD) approach, observations in the distant past are down-weighted in the likelihood function relative to more recent observations, while still giving the practitioner control over the choice of data model. One of the most popular non-stationary techniques in the academic finance community, rolling window estimation, is a special case of our PWD approach. Our PWD framework is a simpler alternative compared to popular state-space methods that explicitly model the evolution of an underlying state vector. We demonstrate the benefits of our PWD approach in terms of predictive performance compared to both stationary models and alternative non-stationary methods. In a financial application to thirty industry portfolios, our PWD method has a significantly favorable predictive performance and draws a number of substantive conclusions about the evolution of the coefficients and the importance of market factors over time.

stat.AP

Locating recombination hot spots in genomic sequences through the singular value decomposition

Locating recombination hotspots in genomic data is an important but difficult task. Current methods frequently rely on estimating complicated models at high computational cost. In this paper we develop an extremely fast, scalable method for inferring recombination hot spots in a population of genomic sequences that is based on the singular value decomposition. Our method performs well in several synthetic data scenarios. We also apply our technique to a real data investigation of the evolution of drug therapy resistance in a population of HIV genomic sequences. Finally, we compare our method both on real and simulated data to a state of the art algorithm.

stat.AP

Probabilistic Approach for Evaluating Metabolite Sample Integrity

The success of metabolomics studies depends upon the "fitness" of each biological sample used for analysis: it is critical that metabolite levels reported for a biological sample represent an accurate snapshot of the studied organism's metabolite profile at time of sample collection. Numerous factors may compromise metabolite sample fitness, including chemical and biological factors which intervene during sample collection, handling, storage, and preparation for analysis. We propose a probabilistic model for the quantitative assessment of metabolite sample fitness. Collection and processing of nuclear magnetic resonance (NMR) and ultra-performance liquid chromatography (UPLC-MS) metabolomics data is discussed. Feature selection methods utilized for multivariate data analysis are briefly reviewed, including feature clustering and computation of latent vectors using spectral methods. We propose that the time-course of metabolite changes in samples stored at different temperatures may be utilized to identify changing-metabolite-to-stable-metabolite ratios as markers of sample fitness. Tolerance intervals may be computed to characterize these ratios among fresh samples. In order to discover additional structure in the data relevant to sample fitness, we propose using data labeled according to these ratios to train a Dirichlet process mixture model (DPMM) for assessing sample fitness. DPMMs are highly intuitive since they model the metabolite levels in a sample as arising from a combination of processes including, e.g., normal biological processes and degradation- or contamination-inducing processes. The outputs of a DPMM are probabilities that a sample is associated with a given process, and these probabilities may be incorporated into a final classifier for sample fitness.

q-bio.QM

openWAR: An Open Source System for Evaluating Overall Player Performance in Major League Baseball

Within baseball analytics, there is substantial interest in comprehensive statistics intended to capture overall player performance. One such measure is Wins Above Replacement (WAR), which aggregates the contributions of a player in each facet of the game: hitting, pitching, baserunning, and fielding. However, current versions of WAR depend upon proprietary data, ad hoc methodology, and opaque calculations. We propose a competitive aggregate measure, openWAR, that is based upon public data and methodology with greater rigor and transparency. We discuss a principled standard for the nebulous concept of a "replacement" player. Finally, we use simulation-based techniques to provide interval estimates for our openWAR measure.

stat.AP

Variable selection for BART: An application to gene regulation

We consider the task of discovering gene regulatory networks, which are defined as sets of genes and the corresponding transcription factors which regulate their expression levels. This can be viewed as a variable selection problem, potentially with high dimensionality. Variable selection is especially challenging in high-dimensional settings, where it is difficult to detect subtle individual effects and interactions between predictors. Bayesian Additive Regression Trees [BART, Ann. Appl. Stat. 4 (2010) 266-298] provides a novel nonparametric alternative to parametric regression approaches, such as the lasso or stepwise regression, especially when the number of relevant predictors is sparse relative to the total number of available predictors and the fundamental relationships are nonlinear. We develop a principled permutation-based inferential approach for determining when the effect of a selected predictor is likely to be real. Going further, we adapt the BART procedure to incorporate informed prior information about variable importance. We present simulations demonstrating that our method compares favorably to existing parametric and nonparametric procedures in a variety of data settings. To demonstrate the potential of our approach in a biological context, we apply it to the task of inferring the gene regulatory network in yeast (Saccharomyces cerevisiae). We find that our BART-based procedure is best able to recover the subset of covariates with the largest signal compared to other variable selection methods. The methods developed in this work are readily available in the R package bartMachine.

stat.ME

Probability aggregation in time-series: Dynamic hierarchical modeling of sparse expert beliefs

Most subjective probability aggregation procedures use a single probability judgment from each expert, even though it is common for experts studying real problems to update their probability estimates over time. This paper advances into unexplored areas of probability aggregation by considering a dynamic context in which experts can update their beliefs at random intervals. The updates occur very infrequently, resulting in a sparse data set that cannot be modeled by standard time-series procedures. In response to the lack of appropriate methodology, this paper presents a hierarchical model that takes into account the expert's level of self-reported expertise and produces aggregate probabilities that are sharp and well calibrated both in- and out-of-sample. The model is demonstrated on a real-world data set that includes over 2300 experts making multiple probability forecasts over two years on different subsets of 166 international political events.

stat.AP