SearcharxivSearch

arXiv subjects

Lee Kennedy-Shaffer

Publications and source records attributed to Lee Kennedy-Shaffer.

8 recordsLinked to original sources

Emulating Stepped-Wedge Cluster Randomized Trials to Evaluate Health Policies and Interventions

Both cluster randomized trials and quasi-experimental designs are used to evaluate the impact of health and social policies and interventions. Stepped-wedge cluster randomized trials randomize a staggered adoption approach, while recent difference-in-differences methods allow analysis of non-randomized settings where similar policies are adopted at different time points. These approaches have become common, but the sheer variety of methods for analyzing observational studies with staggered adoption makes it challenging to clearly design and report such studies. We propose that observational and quasi-experimental study investigators can address these challenges by emulating stepped-wedge cluster randomized trials in the target trial emulation framework. The conceptual framework and reporting standards of trial emulation will encourage consideration of key features of these designs, such as policy heterogeneity and time-varying effects, and clear reporting of the estimand and assumptions. It also highlights areas where those interested in randomized trials and quasi-experimental designs can benefit from one another's experience by bringing insights across disciplines. Questions of treatment effect heterogeneity, power, spillovers, and anticipation effects, among others, are common to both fields and can benefit from cross-pollination. This article also demonstrates how trial emulation can identify settings that are not well-served by either approach, thereby avoiding studies unlikely to generate high-quality causal evidence. Finally, it informs the bias-variance-generalizability trade-off that arises with design and analysis choices made in these settings, supporting better evidence generation and interpretation in settings where important questions can be answered.

stat.ME

One Person, How Many Votes? Demographic Distortions in United States Elections

Representative democracy in the United States relies on election systems that transmit votes into representatives in three key bodies: the two chambers of the federal legislature (House of Representatives and Senate) and the Electoral College, which selects the President and Vice-President. This happens through a process of re-weighting based on geographic units (congressional districts and states) that can introduce substantial distortion. In this paper, I propose quantitative measures of this distortion that can be applied to demographic groups, using Census data, to assess and visualize these distortive effects. These include the absolute weight of votes under these systems and the excess population represented in the bodies through the distortions. Visualizing these metrics from 2000 -- 2020 shows persistent malapportionment in key demographic categories. White (non-Hispanic) residents, residents of rural areas, and owner-occupied households are overrepresented in the Senate and Electoral College; Black and Hispanic people, urban dwellers, and renter-occupied households are underrepresented. For urban residents, this underrepresentation is the equivalent of 25 million fewer residents in the Senate and nearly 5 million in the Electoral College. I discuss implications for further research on the effects of these distortions and their interactions with other features of the electoral system.

stat.AP

Spillovers and Effect Attenuation in Firearm Policy Research in the United States

In the United States, firearm-related deaths and injuries are a major public health issue. Because of limited federal action, state policies are particularly important, and their evaluation informs the actions of other policymakers. The movement of firearms across state and local borders, however, can undermine the effectiveness of these policies and have statistical consequences for their empirical evaluation. This movement causes spillover and bypass effects of policies, wherein interventions affect nearby control states and the lack of intervention in nearby states reduces the effectiveness in the intervention states. While some causal inference methods exist to account for spillover effects and reduce bias, these do not necessarily align well with the data available for firearm research or with the most policy-relevant estimands. Integrated data infrastructure and new methods are necessary for a better understanding of the effects these policies would have if widely adopted. In the meantime, appropriately understanding and interpreting effect estimates from quasi-experimental analyses is crucial for ensuring that effective policies are not dismissed due to these statistical challenges.

stat.AP

An Undergraduate Course on the Statistical Principles of Research Study Design

The undergraduate curriculum in statistics and data science is undergoing changes to accommodate new methods, newly interested students, and the changing role of statistics in society. Because of this, it is more important than ever that students understand the role of study design and how to formulate meaningful scientific and statistical research questions. While the traditional Design of Experiments course is still extremely valuable for students heading to industry and research careers, a broader study design course that incorporates survey sampling, observational studies, and the basics of causal inference with randomized experiment design is particularly useful for students with a wide range of applied interests. Here, I describe such a course at a small liberal arts college, along with ways to adapt it to meet different student and instructor background and interests. The course serves as a valuable bridge to advanced statistical coursework, meets key statistical literacy and communication learning goals, and can be tailored to the desired level of computational and mathematical fluency. Through reading, discussing, and critiquing actual published research studies, students learn that statistics is a living discipline with real consequences and become better consumers and producers of scientific research and data-driven insights.

stat.OT

The Effects of Major League Baseball's Ban on Infield Shifts: A Quasi-Experimental Analysis

From 2020 to 2023, Major League Baseball changed rules affecting team composition, player positioning, and game time. Understanding the effects of these rules is crucial for leagues, teams, players, and other relevant parties to assess their impact and to advocate either for further changes or undoing previous ones. Panel data and quasi-experimental methods provide useful tools for causal inference in these settings. I demonstrate this potential by analyzing the effect of the 2023 shift ban at both the league-wide and player-specific levels. Using difference-in-differences analysis, I show that the policy increased batting average on balls in play and on-base percentage for left-handed batters by a modest amount (nine points). For individual players, synthetic control analyses identify several players whose offensive performance (on-base percentage, on-base plus slugging percentage, and weighted on-base average) improved substantially (over 70 points in several cases) because of the rule change, and other players with previously high shift rates for whom it had little effect. This article both estimates the impact of this specific rule change and demonstrates how these methods for causal inference are potentially valuable for sports analytics -- at the player, team, and league levels -- more broadly.

stat.AP

A Generalized Difference-in-Differences Estimator for Randomized Stepped-Wedge and Observational Staggered Adoption Settings

Staggered treatment adoption arises in the evaluation of policy impact and implementation in many settings, including both randomized stepped-wedge trials and non-randomized quasi-experiments with panel data. In both settings, getting an interpretable, unbiased effect estimate requires careful consideration of the target estimand and possible treatment effect heterogeneities. This paper proposes a novel non-parametric approach to this estimation for either setting. By constructing an estimator using weighted averages of two-by-two difference-in-differences comparisons as building blocks, the investigator can target the desired estimand for any assumed treatment effect heterogeneities. This provides desirable bias and interpretation properties while using the comparisons efficiently to mitigate the loss of precision, without requiring correct variance specification. The methods are demonstrated for both a randomized stepped-wedge trial on the impact of novel tuberculosis diagnostic tools and an observational staggered adoption study on the effects of COVID-19 vaccine financial incentive lotteries in U.S. states; these are compared to analyses using previous methods. A full algorithm with R code is provided to implement this method and to compare against existing methods. The proposed method allows for high flexibility and clear targeting of desired effects, providing one solution to the bias-variance-generalizability tradeoff.

stat.ME

Novel Methods for the Analysis of Stepped Wedge Cluster Randomized Trials

Stepped wedge cluster randomized trials (SW-CRTs) have become increasingly popular and are used for a variety of interventions and outcomes, often chosen for their feasibility advantages. SW-CRTs must account for time trends in the outcome because of the staggered rollout of the intervention inherent in the design. Robust inference procedures and non-parametric analysis methods have recently been proposed to handle such trends without requiring strong parametric modeling assumptions, but these are less powerful than model-based approaches. We propose several novel analysis methods that reduce reliance on modeling assumptions while preserving some of the increased power provided by the use of mixed effects models. In one method, we use the synthetic control approach to find the best matching clusters for a given intervention cluster. This approach can improve the power of the analysis but is fully non-parametric. Another method makes use of within-cluster crossover information to construct an overall estimator. We also consider methods that combine these approaches to further improve power. We test these methods on simulated SW-CRTs and identify settings for which these methods gain robustness to model misspecification while retaining some of the power advantages of mixed effects models. Finally, we propose avenues for future research on the use of these methods; motivation for such research arises from their flexibility, which allows the identification of specific causal contrasts of interest, their robustness, and the potential for incorporating covariates to further increase power. Investigators conducting SW-CRTs might well consider such methods when common modeling assumptions may not hold.

stat.ME

Nonpositive Eigenvalues of the Adjacency Matrix and Lower Bounds for Laplacian Eigenvalues

Let $NPO(k)$ be the smallest number $n$ such that the adjacency matrix of any undirected graph with $n$ vertices or more has at least $k$ nonpositive eigenvalues. We show that $NPO(k)$ is well-defined and prove that the values of $NPO(k)$ for $k=1,2,3,4,5$ are $1,3,6,10,16$ respectively. In addition, we prove that for all $k \geq 5$, $R(k,k+1) \ge NPO(k) > T_k$, in which $R(k,k+1)$ is the Ramsey number for $k$ and $k+1$, and $T_k$ is the $k^{th}$ triangular number. This implies new lower bounds for eigenvalues of Laplacian matrices: the $k$-th largest eigenvalue is bounded from below by the $NPO(k)$-th largest degree, which generalizes some prior results.

math.CO