SearcharxivSearch

arXiv subjects

Dylan Small

Publications and source records attributed to Dylan Small.

15 recordsLinked to original sources

Randomization inference for treatment effects on survival outcomes

The log-rank test and Kaplan--Meier plot are standard tools for analyzing time-to-event data in randomized clinical trials, yet neither provides a summary of the magnitude of the treatment effect. Practitioners typically fill this gap by reporting a hazard ratio from a Cox proportional-hazards model or an acceleration factor from an accelerated failure time (AFT) model, but both require assumptions beyond those needed for the log-rank test or Kaplan--Meier estimator. We propose two nonparametric confidence intervals for scalar effect-size summaries, an additive shift c and a multiplicative factor $\rho$, obtained by inverting the log-rank test under sharp null hypotheses of constant treatment effects. Building on the randomization-inference framework of Li and Small (2023), both intervals are valid under the randomization distribution alone, requiring no assumptions for the event-time distribution. We evaluate the proposed multiplicative interval via simulation, finding that it maintains nominal coverage across a range of censoring rates and sample sizes, including under data-generating processes that misspecify a parametric AFT model, while incurring only a modest efficiency loss compared to parametric AFT inference under correct specification. We illustrate the approach using data from a randomized trial of rhDNase for cystic fibrosis and provide R code and a Shiny application for ease of implementation.

stat.ME

Identification and Inference with Invalid Instruments

Instrumental variables (IVs) are widely used to study the causal effect of an exposure on an outcome in the presence of unmeasured confounding. IVs require an instrument, a variable that is (A1) associated with the exposure, (A2) has no direct effect on the outcome except through the exposure, and (A3) is not related to unmeasured confounders. Unfortunately, finding variables that satisfy conditions (A2) or (A3) can be challenging in practice. This paper reviews works where instruments may not satisfy conditions (A2) or (A3), which we refer to as invalid instruments. We review identification and inference under different violations of (A2) or (A3), specifically under linear models, non-linear models, and heteroskedatic models. We conclude with an empirical comparison of various methods by re-analyzing the effect of body mass index on systolic blood pressure from the UK Biobank.

stat.ME

Adolescent sports participation and health in early adulthood: An observational study

We study the impact of teenage sports participation on early-adulthood health using longitudinal data from the National Study of Youth and Religion. We focus on two primary outcomes measured at ages 23--28 -- self-rated health and total score on the PHQ9 Patient Depression Questionnaire -- and control for several potential confounders related to demographics and family socioeconomic status. To probe the possibility that certain types of sports participation may have larger effects on health than others, we conduct a matched observational study at each level within a hierarchy of exposures. Our hierarchy ranges from broadly defined exposures (e.g., participation in any organized after-school activity) to narrow (e.g., participation in collision sports). We deployed an ordered testing approach that exploits the hierarchical relationships between our exposure definitions to perform our analyses while maintaining a fixed family-wise error rate. Compared to teenagers who did not participate in any after-school activities, those who participated in sports had statistically significantly better self-rated and mental health outcomes in early adulthood.

stat.AP

Sensitivity Analysis for Matched Observational Studies with Continuous Exposures and Binary Outcomes

Matching is one of the most widely used study designs for adjusting for measured confounders in observational studies. However, unmeasured confounding may exist and cannot be removed by matching. Therefore, a sensitivity analysis is typically needed to assess a causal conclusion's sensitivity to unmeasured confounding. Sensitivity analysis frameworks for binary exposures have been well-established for various matching designs and are commonly used in various studies. However, unlike the binary exposure case, there still lacks valid and general sensitivity analysis methods for continuous exposures, except in some special cases such as pair matching. To fill this gap in the binary outcome case, we develop a sensitivity analysis framework for general matching designs with continuous exposures and binary outcomes. First, we use probabilistic lattice theory to show our sensitivity analysis approach is finite-population-exact under Fisher's sharp null. Second, we prove a novel design sensitivity formula as a powerful tool for asymptotically evaluating the performance of our sensitivity analysis approach. Third, to allow effect heterogeneity with binary outcomes, we introduce a framework for conducting asymptotically exact inference and sensitivity analysis on generalized attributable effects with binary outcomes via mixed-integer programming. Fourth, for the continuous outcomes case, we show that conducting an asymptotically exact sensitivity analysis in matched observational studies when both the exposures and outcomes are continuous is generally NP-hard, except in some special cases such as pair matching. As a real data application, we apply our new methods to study the effect of early-life lead exposure on juvenile delinquency. We also develop a publicly available R package for implementation of the methods in this work.

stat.ME

Test-negative designs with various reasons for testing: statistical bias and solution

Test-negative designs are widely used for post-market evaluation of vaccine effectiveness, particularly in cases when randomized trials are not feasible. Differing from classical test-negative designs where only healthcare-seekers with symptoms are included, recent test-negative designs have involved individuals with various reasons for testing, especially in an outbreak setting. While including these data can increase sample size and hence improve precision, concerns have been raised about whether they introduce bias into the current framework of test-negative designs, thereby demanding a formal statistical examination of this modified design. In this article, using statistical derivations, causal graphs, and numerical demonstrations, we show that the standard odds ratio estimator may be biased if various reasons for testing are not accounted for. To eliminate this bias, we identify three categories of reasons for testing, including symptoms, mandatory screening, and case contact tracing, and characterize associated statistical properties and estimands. Based on our characterization, we show how to consistently estimate each estimand via stratification. Furthermore, we describe when these estimands correspond to the same vaccine effectiveness parameter, and, when appropriate, propose a stratified estimator that can incorporate multiple reasons for testing and improve precision. The performance of our proposed method is demonstrated through simulation studies.

stat.ME

Pre-analysis protocol for an observational study on the effects of adolescent sports participation on health in early adulthood

We will study the impact of adolescent sports participation on early-adulthood health using longitudinal data from the National Study of Youth and Religion. We focus on two primary outcomes measured at ages 23--28 -- self-rated health and total score on the PHQ9 Patient Depression Questionnaire -- and control for several potential confounders related to demographics and family socioeconomic status. Comparing outcomes between sports participants and matched non-sports participants with similar confounders is straightforward. Unfortunately, an analysis based on such a broad exposure cannot probe the possibility that participation in certain types of sports (e.g., collision sports like football or soccer) may have larger effects on health than others. In this study, we introduce a hierarchy of exposure definitions, ranging from broad (participation in any after-school organized activity) to narrow (e.g., participation in limited-contact sports). We will perform separate matched observational studies, one for each definition, to estimate the health effects of several levels of sports participation. In order to conduct these studies while maintaining a fixed family-wise error rate, we deployed an ordered testing approach that exploits the logical relationships between exposure definitions. Our study will also consider several secondary outcomes including body mass index, life satisfaction, and problematic drinking behavior.

stat.AP

Inference for a test-negative case-control study with added controls

Test-negative designs with added controls have recently been proposed to study COVID-19. An individual is test-positive or test-negative accordingly if they took a test for a disease but tested positive or tested negative. Adding a control group to a comparison of test-positives vs test-negatives is useful since additional comparison of test-positives vs controls can have potential biases different from the first comparison. Bonferroni correction ensures necessary type-I error control for these two comparisons done simultaneously. We propose two new methods for inference which have better interpretability and higher statistical power for these designs. These methods add a third comparison that is essentially independent of the first comparison, but our proposed second method often pays much less for these three comparisons than what a Bonferroni correction would pay for the two comparisons.

stat.ME

Technical Preprint: Rationale and Design of a Planned Observational Study to Evaluate the Impact of Hydrocodone Rescheduling on Opioid Prescribing After Surgery

In October 2014, the US Drug Enforcement Agency (DEA) reclassified hydrocodone from Schedule III to Schedule II of the Controlled Substances Act, resulting in a prohibition on refills in the initial prescription. While this schedule change was associated with overall decreases in the rate of filled hydrocodone prescriptions and opioid dispensing, available studies conflict regarding its impact on acute opioid prescribing among surgical patients. Here, we present the rationale and design of a planned study to measure the effect of hydrocodone rescheduling using a difference-in-differences design that leverages anticipated variation in the relative impact of this policy on patients treated by surgeons that more or less frequently prescribed hydrocodone products versus other opioids prior to the schedule change. Additionally, we present findings from preliminary study conducted on a subset of our full planned sample to assess for potential differences in outcome trends over the 3 years prior to rescheduling among patients treated by surgeons who commonly prescribed hydrocodone versus those treated by surgeons who rarely prescribed hydrocodone.

stat.AP

A calibrated sensitivity analysis for matched observational studies with application to the effect of second-hand smoke exposure on blood lead levels in U.S. children

Matched observational studies are commonly used to study treatment effects in non-randomized data. After matching for observed confounders, there could remain bias from unobserved confounders. A standard way to address this problem is to do a sensitivity analysis. A sensitivity analysis asks how sensitive the result is to a hypothesized unmeasured confounder U. One method, known as simultaneous sensitivity analysis, has two sensitivity parameters: one relating U to treatment assignment and the other to response. This method assumes that in each matched set, U is distributed to make the bias worst. This approach has two concerning features. First, this worst case distribution of U in each matched set does not correspond to a realistic distribution of U in the population. Second, sensitivity parameters are in absolute scales which are hard to compare to observed covariates. We address these concerns by introducing a method that endows U with a probability distribution in the population and calibrates the unmeasured confounder to the observed covariates. We compare our method to simultaneous sensitivity analysis in simulations and in a study of the effect of second-hand smoke exposure on blood lead levels in U.S. children.

stat.ME

Comparing Covariate Prioritization via Matching to Machine Learning Methods for Causal Inference using Five Empirical Applications

When investigators seek to estimate causal effects, they often assume that selection into treatment is based only on observed covariates. Under this identification strategy, analysts must adjust for observed confounders. While basic regression models have long been the dominant method of statistical adjustment, more robust methods based on matching or weighting have become more common. Of late, even more flexible methods based on machine learning methods have been developed for statistical adjustment. These machine learning methods are designed to be black box methods with little input from the researcher. Recent research used a data competition to evaluate various methods of statistical adjustment and found that black box methods out performed all other methods of statistical adjustment. Matching methods with covariate prioritization are designed for direct input from substantive investigators in direct contrast to black methods. In this article, we use a different research design to compare matching with covariate prioritization to black box methods. We use black box methods to replicate results from five studies where matching with covariate prioritization was used to customize the statistical adjustment in direct response to substantive expertise. We find little difference across the methods. We conclude with advice for investigators.

stat.AP

Urban Vibrancy and Safety in Philadelphia

Statistical analyses of urban environments have been recently improved through publicly available high resolution data and mapping technologies that have been adopted across industries. These technologies allow us to create metrics to empirically investigate urban design principles of the past half-century. Philadelphia is an interesting case study for this work, with its rapid urban development and population increase in the last decade. We outline a data analysis pipeline for exploring the association between safety and local neighborhood features such as population, economic health and the built environment. As a particular example of our analysis pipeline, we focus on quantitative measures of the built environment that serve as proxies for vibrancy: the amount of human activity in a local area. Historically, vibrancy has been very challenging to measure empirically. Measures based on land use zoning are not an adequate description of local vibrancy and so we construct a database and set of measures of business activity in each neighborhood. We employ several matching analyses to explore the relationship between neighborhood vibrancy and safety, such as comparing high crime versus low crime locations within the same neighborhood. As additional sources of urban data become available, our analysis pipeline can serve as the template for further investigations into the relationships between safety, economic factors and the built environment at the local neighborhood level.

stat.AP

Control Function Instrumental Variable Estimation of Nonlinear Causal Effect Models

The instrumental variable method consistently estimates the effect of a treatment when there is unmeasured confounding and a valid instrumental variable. A valid instrumental variable is a variable that is independent of unmeasured confounders and affects the treatment but does not have a direct effect on the outcome beyond its effect on the treatment. Two commonly used estimators for using an instrumental variable to estimate a treatment effect are the two stage least squares estimator and the control function estimator. For linear causal effect models, these two estimators are equivalent, but for nonlinear causal effect models, the estimators are different. We provide a systematic comparison of these two estimators for nonlinear causal effect models and develop an approach to combing the two estimators that generally performs better than either one alone. We show that the control function estimator is a two stage least squares estimator with an augmented set of instrumental variables. If these augmented instrumental variables are valid, then the control function estimator can be much more efficient than usual two stage least squares without the augmented instrumental variables while if the augmented instrumental variables are not valid, then the control function estimator may be inconsistent while the usual two stage least squares remains consistent. We apply the Hausman test to test whether the augmented instrumental variables are valid and construct a pretest estimator based on this test. The pretest estimator is shown to work well in a simulation study. An application to the effect of exposure to violence on time preference is considered.

stat.ME

Selection bias when using instrumental variable methods to compare two treatments but more than two treatments are available

Instrumental variable (IV) methods are widely used to adjust for the bias in estimating treatment effects caused by unmeasured confounders in observational studies. In this manuscript, we provide empirical and theoretical evidence that the IV methods may result in biased treatment effects if applied on a data set in which subjects are preselected based on their received treatments. We frame this as a selection bias problem and propose a procedure that identifies the treatment effect of interest as a function of a vector of sensitivity parameters. We also list assumptions under which analyzing the preselected data does not lead to a biased treatment effect estimate. The performance of the proposed method is examined using simulation studies. We applied our method on The Health Improvement Network (THIN) database to estimate the comparative effect of metformin and sulfonylureas on weight gain among diabetic patients.

stat.ME

Instrumental Variable Estimation When Compliance is not Deterministic: The Stochastic Monotonicity Assumption

The instrumental variables (IV) method is a method for making causal inferences about the effect of a treatment based on an observational study in which there are unmeasured confounding variables. The method requires a valid IV, a variable that is independent of the unmeasured confounding variables and is associated with the treatment but which has no effect on the outcome beyond its effect on the treatment. An additional assumption that is often made for the IV method is deterministic monotonicity, which is an assumption that for each subject, the level of the treatment that a subject would take if given a level of the IV is a monotonic increasing function of the level of the IV. Under deterministic monotonicity, the IV method identifies the average treatment effect for the compliers (those subject who would take the treatment if encouraged to do so by the IV and not take the treatment if not encouraged). However, deterministic monotonicity is sometimes not realistic. We introduce a stochastic monotonicity condition which relaxes deterministic monotonicity in that it does not require that a monotonic increasing relationship hold within subjects between the levels of the IV and the level of the treatment that the subject would take if given a level of the IV, but only that a monotonic increasing relationship hold across subjects between the IV and the treatment in a certain manner. We show that under stochastic monotonicity, the IV method identifies a weighted average of treatment effects with greater weight on subgroups of subjects on whom the IV has a stronger effect. We provide bounds on the global average treatment effect under stochastic monotonicity and a sensitivity analysis for violations of the stochastic monotonicity assumption.

stat.ME

Defining and Estimating Intervention Effects for Groups that will Develop an Auxiliary Outcome

It has recently become popular to define treatment effects for subsets of the target population characterized by variables not observable at the time a treatment decision is made. Characterizing and estimating such treatment effects is tricky; the most popular but naive approach inappropriately adjusts for variables affected by treatment and so is biased. We consider several appropriate ways to formalize the effects: principal stratification, stratification on a single potential auxiliary variable, stratification on an observed auxiliary variable and stratification on expected levels of auxiliary variables. We then outline identifying assumptions for each type of estimand. We evaluate the utility of these estimands and estimation procedures for decision making and understanding causal processes, contrasting them with the concepts of direct and indirect effects. We motivate our development with examples from nephrology and cancer screening, and use simulated data and real data on cancer screening to illustrate the estimation methods.

math.ST