SearcharxivSearch

arXiv subjects

Forrest W. Crawford

Publications and source records attributed to Forrest W. Crawford.

At least 19 recordsLinked to original sources

Time-varying confounding in epidemic intervention evaluations

Estimating the causal effect of a time-varying public health intervention on the course of an infectious disease epidemic is an important methodological challenge. During the COVID-19 pandemic, researchers attempted to estimate the effects of social distancing policies, stay-at-home orders, school closures, mask mandates, vaccination programs, and many other interventions on population-level infection outcomes. However, measuring the effect of these interventions is complicated by time-varying confounding: public health interventions are causal consequences of prior outcomes and interventions, as well as causes of future outcomes and interventions. Researchers have shown repeatedly that neglecting time-varying confounding for individual-level longitudinal interventions can result in profoundly biased estimates of causal effects. However, the issue with time-varying confounding bias has often been overlooked in population-level epidemic intervention evaluations. In this paper, we explain why associational modeling to estimate the effects of interventions on epidemic outcomes based on observations can be prone to time-varying confounding bias. Using causal reasoning and model-based simulation, we show how directional bias due to time-varying confounding arises in associational modeling and the misleading conclusions it induces.

stat.AP

Communication network dynamics in a large organizational hierarchy

Most businesses impose a supervisory hierarchy on employees to facilitate management, decision-making, and collaboration, yet routine inter-employee communication patterns within workplaces tend to emerge more naturally as a consequence of both supervisory relationships and the needs of the organization. What then is the relationship between a formal organizational structure and the emergent communications between its employees? Understanding the nature of this relationship is critical for the successful management of an organization. While scholars of organizational management have proposed theories relating organizational trees to communication dynamics, and separately, network scientists have studied the topological structure of communication patterns in different types of organizations, existing empirical analyses are both lacking in representativeness and limited in size. In fact, much of the methodology used to study the relationship between organizational hierarchy and communication patterns comes from analyses of the Enron email corpus, reflecting a uniquely dysfunctional corporate environment. In this paper, we develop new methodology for assessing the relationship between organizational hierarchy and communication dynamics and apply it to Microsoft Corporation, currently the highest valued company in the world, consisting of approximately 200,000 employees divided into 88 teams. This reveals distinct communication network structures within and between teams. We then characterize the relationship of routine employee communication patterns to these team supervisory hierarchies, while empirically evaluating several theories of organizational management and performance. To do so, we propose new measures of communication reciprocity and new shortest-path distances for trees to track the frequency of messages passed up, down, and across the organizational hierarchy.

stat.AP

Mutually Exciting Point Processes for Crowdfunding Platform Dynamics

Crowdfunding is a powerful tool for individuals or organizations seeking financial support from a vast audience. Despite widespread adoption, managers often lack information about dynamics of their platforms. Hawkes processes have been used to represent self-exciting behavior in a wide variety of empirical fields, but have not been applied to crowdfunding platforms in a way that could help managers understand the dynamics of users' engagement with the platform. In this paper, we extend the Hawkes process to capture important features of crowdfunding platform contributions and apply the model to analyze data from two donation-based platforms. For each user-item pair, the continuous-time conditional intensity is modeled as the superposition of a self-exciting baseline rate and a mutual excitation by preferential attachment, both depending on prior user engagement, and attenuated by a power law decay of user interest. The model is thus structured around two time-varying features -- contribution count and item popularity. We estimate parameters that govern the dynamics of contributions from 2,000 items and 164,000 users over several years. We identify a bottleneck in the user contribution pipeline, measure the force of item popularity, and characterize the decline in user interest over time. A contagion effect is introduced to assess the effect of item popularity on contribution rates. This mechanistic model lays the groundwork for enhanced crowdfunding platform monitoring based on evaluation of counterfactual scenarios and formulation of dynamics-aware recommendations.

stat.AP

The role of discretization scales in causal inference with continuous-time treatment

There are well-established methods for identifying the causal effect of a time-varying treatment applied at discrete time points. However, in the real world, many treatments are continuous or have a finer time scale than the one used for measurement or analysis. While researchers have investigated the discrepancies between estimates under varying discretization scales using simulations and empirical data, it is still unclear how the choice of discretization scale affects causal inference. To address this gap, we present a framework to understand how discretization scales impact the properties of causal inferences about the effect of a time-varying treatment. We introduce the concept of "identification bias", which is the difference between the causal estimand for a continuous-time treatment and the purported estimand of a discretized version of the treatment. We show that this bias can persist even with an infinite number of longitudinal treatment-outcome trajectories. We specifically examine the identification problem in a class of linear stochastic continuous-time data-generating processes and demonstrate the identification bias of the g-formula in this context. Our findings indicate that discretization bias can significantly impact empirical analysis, especially when there are limited repeated measurements. Therefore, we recommend that researchers carefully consider the choice of discretization scale and perform sensitivity analysis to address this bias. We also propose a simple and heuristic quantitative measure for sensitivity concerning discretization and suggest that researchers report this measure along with point and interval estimates in their work. By doing so, researchers can better understand and address the potential impact of discretization bias on causal inference.

stat.ME

Causal identification for continuous-time stochastic processes

Many real-world processes are trajectories that may be regarded as continuous-time "functional data". Examples include patients' biomarker concentrations, environmental pollutant levels, and prices of stocks. Corresponding advances in data collection have yielded near continuous-time measurements, from e.g. physiological monitors, wearable digital devices, and environmental sensors. Statistical methodology for estimating the causal effect of a time-varying treatment, measured discretely in time, is well developed. But discrete-time methods like the g-formula, structural nested models, and marginal structural models do not generalize easily to continuous time, due to the entanglement of uncountably infinite variables. Moreover, researchers have shown that the choice of discretization time scale can seriously affect the quality of causal inferences about the effects of an intervention. In this paper, we establish causal identification results for continuous-time treatment-outcome relationships for general cadlag stochastic processes under continuous-time confounding, through orthogonalization and weighting. We use three concrete running examples to demonstrate the plausibility of our identification assumptions, as well as their connections to the discrete-time g methods literature.

math.ST

Dependence-robust confidence intervals for capture-recapture surveys

Capture-recapture (CRC) surveys are used to estimate the size of a population whose members cannot be enumerated directly. CRC surveys have been used to estimate the number of Covid-19 infections, people who use drugs, sex workers, conflict casualties, and trafficking victims. When $k$ capture samples are obtained, counts of unit captures in subsets of samples are represented naturally by a $2^k$ contingency table in which one element -- the number of individuals appearing in none of the samples -- remains unobserved. In the absence of additional assumptions, the population size is not identifiable (i.e. point-identified). Stringent assumptions about the dependence between samples are often used to achieve point-identification. However, real-world CRC surveys often use convenience samples in which the assumed dependence cannot be guaranteed, and population size estimates under these assumptions may lack empirical credibility. In this work, we apply the theory of partial identification to show that weak assumptions or qualitative knowledge about the nature of dependence between samples can be used to characterize a non-trivial confidence set for the true population size. We construct confidence sets under bounds on pairwise capture probabilities using two methods: test inversion bootstrap confidence intervals, and profile likelihood confidence intervals. Simulation results demonstrate well-calibrated confidence sets for each method. In an extensive real-world study, we apply the new methodology to the problem of using heterogeneous survey data to estimate the number of people who inject drugs in Brussels, Belgium.

stat.ME

A sample size heuristic for network scale-up studies

The network scale-up method (NSUM) is a survey-based method for estimating the number of individuals in a hidden or hard-to-reach subgroup of a general population. In NSUM surveys, sampled individuals report how many others they know in the subpopulation of interest (e.g. "How many sex workers do you know?") and how many others they know in subpopulations of the general population (e.g. "How many bus drivers do you know?"). NSUM is widely used to estimate the size of important epidemiological risk groups, including men who have sex with men, sex workers, HIV+ individuals, and drug users. Unlike several other methods for population size estimation, NSUM requires only a single random sample and the estimator has a conveniently simple form. Despite its popularity, there are no published guidelines for the minimum sample size calculation to achieve a desired statistical precision. Here, we provide a sample size formula that can be employed in any NSUM survey. We show analytically and by simulation that the sample size controls error at the nominal rate and is robust to some forms of network model mis-specification. We apply this methodology to study the minimum sample size and relative error properties of several published NSUM surveys.

stat.ME

Causal identification of infectious disease intervention effects in a clustered population

Causal identification of treatment effects for infectious disease outcomes in interconnected populations is challenging because infection outcomes may be transmissible to others, and treatment given to one individual may affect others' outcomes. Contagion, or transmissibility of outcomes, complicates standard conceptions of treatment interference in which an intervention delivered to one individual can affect outcomes of others. Several statistical frameworks have been proposed to measure causal treatment effects in this setting, including structural transmission models, mediation-based partnership models, and randomized trial designs. However, existing estimands for infectious disease intervention effects are of limited conceptual usefulness: Some are parameters in a structural model whose causal interpretation is unclear, others are causal effects defined only in a restricted two-person setting, and still others are nonparametric estimands that arise naturally in the context of a randomized trial but may not measure any biologically meaningful effect. In this paper, we describe a unifying formalism for defining nonparametric structural causal estimands and an identification strategy for learning about infectious disease intervention effects in clusters of interacting individuals when infection times are observed. The estimands generalize existing quantities and provide a framework for causal identification in randomized and observational studies, including situations where only binary infection outcomes are observed. A semiparametric class of pairwise Cox-type transmission hazard models is used to facilitate statistical inference in finite samples. A comprehensive simulation study compares existing and proposed estimands under a variety of randomized and observational vaccine trial designs.

stat.ME

Identification of causal intervention effects under contagion

Defining and identifying causal intervention effects for transmissible infectious disease outcomes is challenging because a treatment -- such as a vaccine -- given to one individual may affect the infection outcomes of others. Epidemiologists have proposed causal estimands to quantify effects of interventions under contagion using a two-person partnership model. These simple conceptual models have helped researchers develop causal estimands relevant to clinical evaluation of vaccine effects. However, many of these partnership models are formulated under structural assumptions that preclude realistic infectious disease transmission dynamics, limiting their conceptual usefulness in defining and identifying causal treatment effects in empirical intervention trials. In this paper, we propose causal intervention effects in two-person partnerships under arbitrary infectious disease transmission dynamics, and give nonparametric identification results showing how effects can be estimated in empirical trials using time-to-infection or binary outcome data. The key insight is that contagion is a causal phenomenon that induces conditional independencies on infection outcomes that can be exploited for the identification of clinically meaningful causal estimands. These new estimands are compared to existing quantities, and results are illustrated using a realistic simulation of an HIV vaccine trial.

stat.AP

Randomization for the susceptibility effect of an infectious disease intervention

Randomized trials of infectious disease interventions, such as vaccines, often focus on groups of connected or potentially interacting individuals. When the pathogen of interest is transmissible between study subjects, interference may occur: individual infection outcomes may depend on treatments received by others. Epidemiologists have defined the primary causal effect of interest -- called the "susceptibility effect" -- as a contrast in infection risk under treatment versus no treatment, while holding exposure to infectiousness constant. A related quantity -- the "direct effect" -- is defined as an unconditional contrast between the infection risk under treatment versus no treatment. The purpose of this paper is to show that under a widely recommended randomization design, the direct effect may fail to recover the sign of the true susceptibility effect of the intervention in a randomized trial when outcomes are contagious. The analytical approach uses structural features of infectious disease transmission to define the susceptibility effect. A new probabilistic coupling argument reveals stochastic dominance relations between potential infection outcomes under different treatment allocations. The results suggest that estimating the direct effect under randomization may provide misleading inferences about the effect of an intervention -- such as a vaccine -- when outcomes are contagious.

stat.AP

Efficient and minimal length parametric conformal prediction regions

Conformal prediction methods construct prediction regions for iid data that are valid in finite samples. We provide two parametric conformal prediction regions that are applicable for a wide class of continuous statistical models. This class of statistical models includes generalized linear models (GLMs) with continuous outcomes. Our parametric conformal prediction regions possesses finite sample validity, even when the model is misspecified, and are asymptotically of minimal length when the model is correctly specified. The first parametric conformal prediction region is constructed through binning of the predictor space, guarantees finite-sample local validity and is asymptotically minimal at the $\sqrt{\log(n)/n}$ rate when the dimension $d$ of the predictor space is one or two, and converges at the $O\{(\log(n)/n)^{1/d}\}$ rate when $d > 2$. The second parametric conformal prediction region is constructed by transforming the outcome variable to a common distribution via the probability integral transform, guarantees finite-sample marginal validity, and is asymptotically minimal at the $\sqrt{\log(n)/n}$ rate. We develop a novel concentration inequality for maximum likelihood estimation that induces these convergence rates. We analyze prediction region coverage properties, large-sample efficiency, and robustness properties of four methods for constructing conformal prediction intervals for GLMs: fully nonparametric kernel-based conformal, residual based conformal, normalized residual based conformal, and parametric conformal which uses the assumed GLM density as a conformity measure. Extensive simulations compare these approaches to standard asymptotic prediction regions. The utility of the parametric conformal prediction region is demonstrated in an application to interval prediction of glycosylated hemoglobin levels, a blood measurement used to diagnose diabetes.

stat.ME

Estimating the size of a hidden finite set: large-sample behavior of estimators

A finite set is "hidden" if its elements are not directly enumerable or if its size cannot be ascertained via a deterministic query. In public health, epidemiology, demography, ecology and intelligence analysis, researchers have developed a wide variety of indirect statistical approaches, under different models for sampling and observation, for estimating the size of a hidden set. Some methods make use of random sampling with known or estimable sampling probabilities, and others make structural assumptions about relationships (e.g. ordering or network information) between the elements that comprise the hidden set. In this review, we describe models and methods for learning about the size of a hidden finite set, with special attention to asymptotic properties of estimators. We study the properties of these methods under two asymptotic regimes, "infill" in which the number of fixed-size samples increases, but the population size remains constant, and "outfill" in which the sample size and population size grow together. Statistical properties under these two regimes can be dramatically different.

math.ST

Interpretation of the individual effect under treatment spillover

Some interventions may include important spillover or dissemination effects between study participants. For example, vaccines, cash transfers, and education programs may exert a causal effect on participants beyond those to whom individual treatment is assigned. In a recent paper, Buchanan et al. provide a causal definition of the "individual effect" of an intervention in networks of people who inject drugs. In this short note, we discuss the interpretation of the individual effect when a spillover or dissemination effect exists.

stat.AP

Direct likelihood-based inference for discretely observed stochastic compartmental models of infectious disease

Stochastic compartmental models are important tools for understanding the course of infectious diseases epidemics in populations and in prospective evaluation of intervention policies. However, calculating the likelihood for discretely observed data from even simple models -- such as the ubiquitous susceptible-infectious-removed (SIR) model -- has been considered computationally intractable, since its formulation almost a century ago. Recently researchers have proposed methods to circumvent this limitation through data augmentation or approximation, but these approaches often suffer from high computational cost or loss of accuracy. We develop the mathematical foundation and an efficient algorithm to compute the likelihood for discretely observed data from a broad class of stochastic compartmental models. We also give expressions for the derivatives of the transition probabilities using the same technique, making possible inference via Hamiltonian Monte Carlo (HMC). We use the 17th century plague in Eyam, a classic example of the SIR model, to compare our recursion method to sequential Monte Carlo, analyze using HMC, and assess the model assumptions. We also apply our direct likelihood evaluation to perform Bayesian inference for the 2014-2015 Ebola outbreak in Guinea. The results suggest that the epidemic infectious rates have decreased since October 2014 in the Southeast region of Guinea, while rates remain the same in other regions, facilitating understanding of the outbreak and the effectiveness of Ebola control interventions.

stat.CO

Birth/birth-death processes and their computable transition probabilities with biological applications

Birth-death processes track the size of a univariate population, but many biological systems involve interaction between populations, necessitating models for two or more populations simultaneously. A lack of efficient methods for evaluating finite-time transition probabilities of bivariate processes, however, has restricted statistical inference in these models. Researchers rely on computationally expensive methods such as matrix exponentiation or Monte Carlo approximation, restricting likelihood-based inference to small systems, or indirect methods such as approximate Bayesian computation. In this paper, we introduce the birth(death)/birth-death process, a tractable bivariate extension of the birth-death process. We develop an efficient and robust algorithm to calculate the transition probabilities of birth(death)/birth-death processes using a continued fraction representation of their Laplace transforms. Next, we identify several exemplary models arising in molecular epidemiology, macro-parasite evolution, and infectious disease modeling that fall within this class, and demonstrate advantages of our proposed method over existing approaches to inference in these models. Notably, the ubiquitous stochastic susceptible-infectious-removed (SIR) model falls within this class, and we emphasize that computable transition probabilities newly enable direct inference of parameters in the SIR model. We also propose a very fast method for approximating the transition probabilities under the SIR model via a novel branching process simplification, and compare it to the continued fraction representation method with application to the 17th century plague in Eyam. Although the two methods produce similar maximum a posteriori estimates, the branching process approximation fails to capture the correlation structure in the joint posterior distribution.

stat.CO

Risk ratios for contagious outcomes

The risk ratio is a popular tool for summarizing the relationship between a binary covariate and outcome, even when outcomes may be dependent. Investigations of infectious disease outcomes in cohort studies of individuals embedded within clusters -- households, villages, or small groups -- often report risk ratios. Epidemiologists have warned that risk ratios may be misleading when outcomes are contagious, but the nature and severity of this error is not well understood. In this study, we assess the epidemiologic meaning of the risk ratio when outcomes are contagious. We first give a structural definition of infectious disease transmission within clusters, based on the canonical susceptible-infective epidemic model. From this standard characterization, we define the individual-level ratio of instantaneous risks (hazard ratio) as the inferential target, and evaluate the properties of the risk ratio as an estimate of this quantity. We exhibit analytically and by simulation the circumstances under which the risk ratio implies an effect whose direction is opposite that of the true individual-level hazard ratio. In particular, the risk ratio can be greater than one even when the covariate of interest reduces both individual-level susceptibility to infection, and transmissibility once infected. We explain these findings in the epidemiologic language of confounding and relate the direction bias to Simpson's paradox.

stat.ME

Estimating the Size of a Large Network and its Communities from a Random Sample

Most real-world networks are too large to be measured or studied directly and there is substantial interest in estimating global network properties from smaller sub-samples. One of the most important global properties is the number of vertices/nodes in the network. Estimating the number of vertices in a large network is a major challenge in computer science, epidemiology, demography, and intelligence analysis. In this paper we consider a population random graph G = (V;E) from the stochastic block model (SBM) with K communities/blocks. A sample is obtained by randomly choosing a subset W and letting G(W) be the induced subgraph in G of the vertices in W. In addition to G(W), we observe the total degree of each sampled vertex and its block membership. Given this partial information, we propose an efficient PopULation Size Estimation algorithm, called PULSE, that correctly estimates the size of the whole population as well as the size of each community. To support our theoretical analysis, we perform an exhaustive set of experiments to study the effects of sample size, K, and SBM model parameters on the accuracy of the estimates. The experimental results also demonstrate that PULSE significantly outperforms a widely-used method called the network scale-up estimator in a wide variety of scenarios. We conclude with extensions and directions for future work.

stat.ML

Confidence intervals for means under constrained dependence

We develop a general framework for conducting inference on the mean of dependent random variables given constraints on their dependency graph. We establish the consistency of an oracle variance estimator of the mean when the dependency graph is known, along with an associated central limit theorem. We derive an integer linear program for finding an upper bound for the estimated variance when the graph is unknown, but topological and degree-based constraints are available. We develop alternative bounds, including a closed-form bound, under an additional homoskedasticity assumption. We establish a basis for Wald-type confidence intervals for the mean that are guaranteed to have asymptotically conservative coverage. We apply the approach to inference from a social network link-tracing study and provide statistical software implementing the approach.

math.ST