SearcharxivSearch

arXiv subjects

Eben Kenah

Publications and source records attributed to Eben Kenah.

At least 19 recordsLinked to original sources

Simultaneous confidence bands for cumulative hazard via exchangeable bootstrap and box calibration

Resampling-based simultaneous confidence bands for cumulative hazard functions often undercover in finite samples with right censoring. We study two aspects of the construction that can contribute to this gap, the resampling scheme and the calibration statistic, and propose a procedure that intervenes on both. The exchangeable bootstrap reweights the numerator and the denominator of the Nelson-Aalen ratio, preserving its ratio structure. The box-calibrated discrepancy constructs lower and upper step envelopes from adjacent values of the original and resampled Nelson-Aalen estimators and measures the resulting vertical discrepancy. We establish conditional weak convergence of the exchangeable bootstrap, prove that box calibration is first-order asymptotically equivalent to grid calibration, and show that the resulting band attains nominal coverage asymptotically. The box correction uses the same bootstrap paths and event-time grid as grid calibration; after each bootstrap path is formed, it requires only an additional linear pass over the event-time grid and therefore has negligible computational overhead. In simulations across a range of hazard shapes and censoring levels, the exchangeable bootstrap with box calibration is, in most configurations, closest to nominal coverage among the methods considered. A notable consequence is a ranking reversal: the ratio-preserving exchangeable bootstrap has the lowest coverage under grid calibration, yet is usually closest to the nominal level after box calibration. A melanoma data example illustrates the practical effect on the cumulative hazard bands. The proposed procedure operates on the original cumulative-hazard scale, requires no variance-stabilizing transformation, and permits inference from time zero.

stat.ME

Rothman diagrams: the geometry of association measure modification and collapsibility

Here, we outline how Rothman diagrams provide a geometric perspective that can help epidemiologists understand the relationships between effect measure modification (which we call association measure modification), collapsibility, and confounding. A Rothman diagram plots the risk of disease in the unexposed on the x-axis and the risk in the exposed on the y-axis. Crude and stratum-specific risks in the two exposure groups define points in the unit square. When there is modification of a measure of association $M$ by a covariate $C$, the stratum-specific values of $M$ differ across strata defined by $C$, so the stratum-specific points are on different contour lines of $M$. We show how collapsibility can be defined in terms of standardization instead of no confounding, and we show that a measure of association is collapsible if and only if all its contour lines are straight. We illustrate these ideas using data from a study in Newcastle, United Kingdom, where the causal effect of smoking on 20-year mortality was confounded by age. From this perspective, it is clear that association measure modification and collapsibility are logically independent of confounding. This distinction can be obscured when these concepts are taught using regression models.

stat.AP

Graphical tools for detection and control of selection bias with multiple exposures and samples

Among recent developments in definitions and analysis of selection bias is the potential outcomes approach of Kenah (Epidemiology, 2023), which allows non-parametric analysis using single-world intervention graphs, linking selection of study participants to identification of causal effects. Mohan & Pearl (JASA, 2021) provide a framework for missing data via directed acyclic graphs augmented with nodes indicating missingness for each sometimes-missing variable, which allows for analysis of more general missing data problems but cannot easily encode scenarios in which different groups of variables are observed in specific subsamples. We give an alternative formulation of the potential outcomes framework based on conditional separable effects and indicators for selection into subsamples. This is practical for problems between the single-sample scenarios considered by Kenah and the variable-wise missingness considered by Mohan & Pearl. This simplifies identification conditions and admits generalizations to scenarios with multiple, potentially nested or overlapping study samples, as well as multiple or time-dependent exposures. We give examples of identifiability arguments for case-cohort studies, multiple or time-dependent exposures, and direct effects of selection.

stat.ME

Rothman diagrams: the geometry of causal inference in epidemiology

Here, we explain and illustrate a geometric perspective on causal inference in cohort studies that can help epidemiologists understand the role of standardization in causal inference as well as the distinctions between confounding, effect modification, and noncollapsibility. For simplicity, we focus on a binary exposure X, a binary outcome D, and a binary confounder C that is not causally affected by X. Rothman diagrams plot risk in the unexposed on the x-axis and risk in the exposed on the y-axis. The crude risks define one point in the unit square, and the stratum-specific risks define two other points in the unit square. These three points can be used to identify confounding and effect modification, and we show briefly how these concepts generalize to confounders with more than two levels. We propose a simplified but equivalent definition of collapsibility in terms of standardization, and we show that a measure of association is collapsible if and only if all of its contour lines are straight. We illustrate these ideas using data from a study conducted in Newcastle upon Tyne, United Kingdom, where the causal effect of smoking on 20-year mortality was confounded by age. We conclude that causal inference should be taught using geometry before using regression models.

math.ST

A potential outcomes approach to selection bias

We propose a novel definition of selection bias in analytic epidemiology using potential outcomes. This definition captures selection bias under both the structural approach (where conditioning on selection into the study opens a noncausal path from exposure to disease in a directed acyclic graph) and the traditional definition (where a given measure of association differs between the study sample and the population eligible for inclusion). It is nonparametric, and selection bias under this approach can be analyzed using single-world intervention graphs both under and away from the null hypothesis. It allows the simultaneous analysis of confounding and selection bias, it explicitly links the selection of study participants to the estimation of causal effects using study data, and it can be adapted to handle selection bias in descriptive epidemiology. Through examples, we show that this approach provides a novel perspective on the variety of mechanisms that can generate selection bias and simplifies the analysis of selection bias in matched studies and case-cohort studies.

stat.ME

Necessary and sufficient conditions for exact closures of epidemic equations on configuration model networks

We prove that the exact closure of SIR pairwise epidemic equations on a configuration model network is possible if and only if the degree distribution is Poisson, Binomial, or Negative Binomial. The proof relies on establishing, for these specific degree distributions, the equivalence of the closed pairwise model and the so-called dynamical survival analysis (DSA) edge-based model which was previously shown to be exact. Indeed, as we show here, the DSA model is equivalent to the well-known edge-based Volz model. We use this result to provide reductions of the closed pairwise and Volz models to the same single equation involving only susceptibles, which has a useful statistical interpretation in terms of the times to infection. We illustrate our findings with some numerical examples.

q-bio.PE

Dynamic Survival Analysis for non-Markovian Epidemic Models

We present a new method for analyzing stochastic epidemic models under minimal assumptions. The method, dubbed DSA, is based on a simple yet powerful observation, namely that population-level mean-field trajectories described by a system of PDE may also approximate individual-level times of infection and recovery. This idea gives rise to a certain non-Markovian agent-based model and provides an agent-level likelihood function for a random sample of infection and/or recovery times. Extensive numerical analyses on both synthetic and real epidemic data from the FMD in the United Kingdom and the COVID-19 in India show good accuracy and confirm method's versatility in likelihood-based parameter estimation. The accompanying software package gives prospective users a practical tool for modeling, analyzing and interpreting epidemic data with the help of the DSA approach.

q-bio.PE

Causal identification of infectious disease intervention effects in a clustered population

Causal identification of treatment effects for infectious disease outcomes in interconnected populations is challenging because infection outcomes may be transmissible to others, and treatment given to one individual may affect others' outcomes. Contagion, or transmissibility of outcomes, complicates standard conceptions of treatment interference in which an intervention delivered to one individual can affect outcomes of others. Several statistical frameworks have been proposed to measure causal treatment effects in this setting, including structural transmission models, mediation-based partnership models, and randomized trial designs. However, existing estimands for infectious disease intervention effects are of limited conceptual usefulness: Some are parameters in a structural model whose causal interpretation is unclear, others are causal effects defined only in a restricted two-person setting, and still others are nonparametric estimands that arise naturally in the context of a randomized trial but may not measure any biologically meaningful effect. In this paper, we describe a unifying formalism for defining nonparametric structural causal estimands and an identification strategy for learning about infectious disease intervention effects in clusters of interacting individuals when infection times are observed. The estimands generalize existing quantities and provide a framework for causal identification in randomized and observational studies, including situations where only binary infection outcomes are observed. A semiparametric class of pairwise Cox-type transmission hazard models is used to facilitate statistical inference in finite samples. A comprehensive simulation study compares existing and proposed estimands under a variety of randomized and observational vaccine trial designs.

stat.ME

Estimating and interpreting secondary attack risk: Binomial considered harmful

The household secondary attack risk (SAR), often called the secondary attack rate or secondary infection risk, is the probability of infectious contact from an infectious household member A to a given household member B, where we define infectious contact to be a contact sufficient to infect B if he or she is susceptible. Estimation of the SAR is an important part of understanding and controlling the transmission of infectious diseases. In practice, it is most often estimated using binomial models such as logistic regression, which implicitly attribute all secondary infections in a household to the primary case. In the simplest case, the number of secondary infections in a household with m susceptibles and a single primary case is modeled as a binomial(m, p) random variable where p is the SAR. Although it has long been understood that transmission within households is not binomial, it is thought that multiple generations of transmission can be safely neglected when p is small. We use probability generating functions and simulations to show that this is a mistake. The proportion of susceptible household members infected can be substantially larger than the SAR even when p is small. As a result, binomial estimates of the SAR are biased upward and their confidence intervals have poor coverage probabilities even if adjusted for clustering. Accurate point and interval estimates of the SAR can be obtained using longitudinal chain binomial models or pairwise survival analysis, which account for multiple generations of transmission within households, the ongoing risk of infection from outside the household, and incomplete follow-up. We illustrate the practical implications of these results in an analysis of household surveillance data collected by the Los Angeles County Department of Public Health during the 2009 influenza A (H1N1) pandemic.

q-bio.QM

Incorporating age and delay into models for biophysical systems

In many biological systems, chemical reactions or changes in a physical state are assumed to occur instantaneously. For describing the dynamics of those systems, Markov models that require exponentially distributed inter-event times have been used widely. However, some biophysical processes such as gene transcription and translation are known to have a significant gap between the initiation and the completion of the processes, which renders the usual assumption of exponential distribution untenable. In this paper, we consider relaxing this assumption by incorporating age-dependent random time delays into the system dynamics. We do so by constructing a measure-valued Markov process on a more abstract state space, which allows us to keep track of the "ages" of molecules participating in a chemical reaction. We study the large-volume limit of such age-structured systems. We show that, when appropriately scaled, the stochastic system can be approximated by a system of Partial Differential Equations (PDEs) in the large-volume limit, as opposed to Ordinary Differential Equations (ODEs) in the classical theory. We show how the limiting PDE system can be used for the purpose of further model reductions and for devising efficient simulation algorithms. In order to describe the ideas, we use a simple transcription process as a running example. We, however, note that the methods developed in this paper apply to a wide class of biophysical systems.

q-bio.PE

Pairwise accelerated failure time regression models for infectious disease transmission in close-contact groups with external sources of infection

Many important questions in infectious disease epidemiology involve the effects of covariates (e.g., age or vaccination status) on infectiousness and susceptibility, which can be measured in studies of transmission in households or other close-contact groups. Because the transmission of disease produces dependent outcomes, these questions are difficult or impossible to address using standard regression models from biostatistics. Pairwise survival analysis handles dependent outcomes by calculating likelihoods in terms of contact interval distributions in ordered pairs of individuals. The contact interval in the ordered pair ij is the time from the onset of infectiousness in i to infectious contact from i to j, where an infectious contact is sufficient to infect j if they are susceptible. Here, we introduce a pairwise accelerated failure time regression model for infectious disease transmission that allows the rate parameter of the contact interval distribution to depend on infectiousness covariates for i, susceptibility covariates for j, and pairwise covariates. This model can simultaneously handle internal infections (caused by transmission between individuals under observation) and external infections (caused by environmental or community sources of infection). In a simulation study, we show that these models produce valid point and interval estimates of parameters governing the contact interval distributions. We also explore the role of epidemiologic study design and the consequences of model misspecification. We use this regression model to analyze household data from Los Angeles County during the 2009 influenza A (H1N1) pandemic, where we find that the ability to account for external sources of infection is critical to estimating the effect of antiviral prophylaxis.

stat.AP

Survival Dynamical Systems for the Population-level Analysis of Epidemics

Motivated by the classical Susceptible-Infected-Recovered (SIR) epidemic models proposed by Kermack and Mckendrick, we consider a class of stochastic compartmental dynamical systems with a notion of partial ordering among the compartments. We call such systems unidirectional Mass Transfer Models (MTMs). We show that there is a natural way of interpreting a uni-directional MTM as a Survival Dynamical System (SDS) that is described in terms of survival functions instead of population counts. This SDS interpretation allows us to employ tools from survival analysis to address various issues with data collection and statistical inference of unidirectional MTMs. In particular, we propose and numerically validate a statistical inference procedure based on SDS-likelihoods. We use the SIR model as a running example throughout the paper to illustrate the ideas.

q-bio.PE

Molecular Infectious Disease Epidemiology: Survival Analysis and Algorithms Linking Phylogenies to Transmission Trees

Recent work has attempted to use whole-genome sequence data from pathogens to reconstruct the transmission trees linking infectors and infectees in outbreaks. However, transmission trees from one outbreak do not generalize to future outbreaks. Reconstruction of transmission trees is most useful to public health if it leads to generalizable scientific insights about disease transmission. In a survival analysis framework, estimation of transmission parameters is based on sums or averages over the possible transmission trees. A phylogeny can increase the precision of these estimates by providing partial information about who infected whom. The leaves of the phylogeny represent sampled pathogens, which have known hosts. The interior nodes represent common ancestors of sampled pathogens, which have unknown hosts. Starting from assumptions about disease biology and epidemiologic study design, we prove that there is a one-to-one correspondence between the possible assignments of interior node hosts and the transmission trees simultaneously consistent with the phylogeny and the epidemiologic data on person, place, and time. We develop algorithms to enumerate these transmission trees and show these can be used to calculate likelihoods that incorporate both epidemiologic data and a phylogeny. A simulation study confirms that this leads to more efficient estimates of hazard ratios for infectiousness and baseline hazards of infectious contact, and we use these methods to analyze data from a foot-and-mouth disease virus outbreak in the United Kingdom in 2001. These results demonstrate the importance of data on individuals who escape infection, which is often overlooked. The combination of survival analysis and algorithms linking phylogenies to transmission trees is a rigorous but flexible statistical foundation for molecular infectious disease epidemiology.

q-bio.QM

Semiparametric Relative-risk Regression for Infectious Disease Data

This paper introduces semiparametric relative-risk regression models for infectious disease data based on contact intervals, where the contact interval from person i to person j is the time between the onset of infectiousness in i and infectious contact from i to j. The hazard of infectious contact from i to j is λ_0(τ)r(β_0^T X_{ij}), where λ_0(τ) is an unspecified baseline hazard function, r is a relative risk function, β_0 is an unknown covariate vector, and X_{ij} is a covariate vector. When who-infects-whom is observed, the Cox partial likelihood is a profile likelihood for βmaximized over all possible λ_0(τ). When who-infects-whom is not observed, we use an EM algorithm to maximize the profile likelihood for βintegrated over all possible combinations of who-infected-whom. This extends the most important class of regression models in survival analysis to infectious disease epidemiology.

stat.ME

Nonparametric survival analysis of epidemic data

This paper develops nonparametric methods for the survival analysis of epidemic data based on contact intervals. The contact interval from person i to person j is the time between the onset of infectiousness in i and infectious contact from i to j, where we define infectious contact as a contact sufficient to infect a susceptible individual. We show that the Nelson-Aalen estimator produces an unbiased estimate of the contact interval cumulative hazard function when who-infects-whom is observed. When who-infects-whom is not observed, we average the Nelson-Aalen estimates from all transmission networks consistent with the observed data using an EM algorithm. This converges to a nonparametric MLE of the contact interval cumulative hazard function that we call the marginal Nelson-Aalen estimate. We study the behavior of these methods in simulations and use them to analyze household surveillance data from the 2009 influenza A(H1N1) pandemic. In an appendix, we show that these methods extend chain-binomial models to continuous time.

stat.ME

Contact intervals, survival analysis of epidemic data, and estimation of R_0

We argue that the time from the onset of infectiousness to infectious contact, which we call the contact interval, is a better basis for inference in epidemic data than the generation or serial interval. Since contact intervals can be right-censored, survival analysis is the natural approach to estimation. Estimates of the contact interval distribution can be used to estimate R_0 in both mass-action and network-based models.

stat.AP

Generation interval contraction and epidemic data analysis

The generation interval is the time between the infection time of an infected person and the infection time of his or her infector. Probability density functions for generation intervals have been an important input for epidemic models and epidemic data analysis. In this paper, we specify a general stochastic SIR epidemic model and prove that the mean generation interval decreases when susceptible persons are at risk of infectious contact from multiple sources. The intuition behind this is that when a susceptible person has multiple potential infectors, there is a ``race'' to infect him or her in which only the first infectious contact leads to infection. In an epidemic, the mean generation interval contracts as the prevalence of infection increases. We call this global competition among potential infectors. When there is rapid transmission within clusters of contacts, generation interval contraction can be caused by a high local prevalence of infection even when the global prevalence is low. We call this local competition among potential infectors. Using simulations, we illustrate both types of competition. Finally, we show that hazards of infectious contact can be used instead of generation intervals to estimate the time course of the effective reproductive number in an epidemic. This approach leads naturally to partial likelihoods for epidemic data that are very similar to those that arise in survival analysis, opening a promising avenue of methodological research in infectious disease epidemiology.

q-bio.QM

Network-based analysis of stochastic SIR epidemic models with random and proportionate mixing

In this paper, we outline the theory of epidemic percolation networks and their use in the analysis of stochastic SIR epidemic models on undirected contact networks. We then show how the same theory can be used to analyze stochastic SIR models with random and proportionate mixing. The epidemic percolation networks for these models are purely directed because undirected edges disappear in the limit of a large population. In a series of simulations, we show that epidemic percolation networks accurately predict the mean outbreak size and probability and final size of an epidemic for a variety of epidemic models in homogeneous and heterogeneous populations. Finally, we show that epidemic percolation networks can be used to re-derive classical results from several different areas of infectious disease epidemiology. In an appendix, we show that an epidemic percolation network can be defined for any time-homogeneous stochastic SIR model in a closed population and prove that the distribution of outbreak sizes given the infection of any given node in the SIR model is identical to the distribution of its out-component sizes in the corresponding probability space of epidemic percolation networks. We conclude that the theory of percolation on semi-directed networks provides a very general framework for the analysis of stochastic SIR models in closed populations.

q-bio.QM