SearcharxivSearch

arXiv subjects

Patrick M. Schnell

Publications and source records attributed to Patrick M. Schnell.

4 recordsLinked to original sources

Kalman Filtering and Smoothing for Improving Precision in Horvitz--Thompson Estimation of Infectious Disease Prevalence

Horvitz--Thompson (HT) estimators can provide unbiased daily estimates of infectious disease prevalence under repeated surveillance by correcting for nonrandom testing induced by scheduled, symptom-based, and contact-tracing components. However, because each HT estimate is based on the testing data available for that day and may involve highly variable inverse probability weights, it can be noisy, have precision that varies over time, and become unavailable during temporary interruptions in testing. The daily HT estimator is modeled as a noisy observation of an underlying prevalence process, with day-specific observation variances estimated using a delete-a-group jackknife. Our primary specification is a joint local linear trend state-space model that extends the standard level-only random walk by adding a latent slope. The Kalman filter improves precision by borrowing information from past estimates. At each time $t$, the process variances are estimated or carried forward using only observations available through time $t$, so the resulting filtered estimate is available in real time. We also describe the corresponding Kalman smoother as a retrospective extension based on the full observed series. When daily HT estimates are missing, the Kalman filter proceeds through prediction-only updates, whereas the corresponding smoother retrospectively reconstructs those periods using later observations. In simulations, the joint Kalman filter substantially improves precision relative to the raw daily HT estimator while preserving the main temporal pattern, and the smoother provides a more stable retrospective summary. In The Ohio State University's fall 2020 SARS-CoV-2 surveillance data, the filter and smoother provide estimates on no-testing days, when the HT estimator provides neither point nor interval estimates, and yield narrower confidence intervals than HT intervals on observed days.

stat.ME

A Counterfactual Framework for Estimating Infectious Disease Prevalence under Repeated Testing with Symptomatic and Contact-Tracing Components

This paper addresses the problem of estimating infectious disease prevalence under longitudinal testing programs that include scheduled, symptomatic, and contact-tracing testing. Our study is motivated by data from The Ohio State University, where a mandatory once-per-week COVID-19 testing and isolation program was implemented during the Fall 2020 semester, supplemented by additional testing for symptomatic individuals and identified contacts. In this setting, the probability of being tested depends on symptoms or contact-tracing status, creating a complex observation process. We develop a counterfactual framework that links the observation process to a hypothetical process in which infection is prevented. This formulation enables unbiased estimation of disease prevalence by modeling the testing process, possibly nonparametrically, without requiring explicit modeling of transmission dynamics, even though the testing and infection processes are jointly dependent.

stat.ME

Graphical tools for detection and control of selection bias with multiple exposures and samples

Among recent developments in definitions and analysis of selection bias is the potential outcomes approach of Kenah (Epidemiology, 2023), which allows non-parametric analysis using single-world intervention graphs, linking selection of study participants to identification of causal effects. Mohan & Pearl (JASA, 2021) provide a framework for missing data via directed acyclic graphs augmented with nodes indicating missingness for each sometimes-missing variable, which allows for analysis of more general missing data problems but cannot easily encode scenarios in which different groups of variables are observed in specific subsamples. We give an alternative formulation of the potential outcomes framework based on conditional separable effects and indicators for selection into subsamples. This is practical for problems between the single-sample scenarios considered by Kenah and the variable-wise missingness considered by Mohan & Pearl. This simplifies identification conditions and admits generalizations to scenarios with multiple, potentially nested or overlapping study samples, as well as multiple or time-dependent exposures. We give examples of identifiability arguments for case-cohort studies, multiple or time-dependent exposures, and direct effects of selection.

stat.ME

Overcoming Repeated Testing Schedule Bias in Estimates of Disease Prevalence

During the COVID-19 pandemic, many institutions such as universities and workplaces implemented testing regimens with every member of some population tested longitudinally, and those testing positive isolated for some time. Although the primary purpose of such regimens was to suppress disease spread by identifying and isolating infectious individuals, testing results were often also used to obtain prevalence and incidence estimates. Such estimates are helpful in risk assessment and institutional planning and various estimation procedures have been implemented, ranging from simple test-positive rates to complex dynamical modeling. Unfortunately, the popular test-positive rate is a biased estimator of prevalence under many seemingly innocuous longitudinal testing regimens with isolation. We illustrate how such bias arises and identify conditions under which the test-positive rate is unbiased. Further, we identify weaker conditions under which prevalence is identifiable and propose a new estimator of prevalence under longitudinal testing. We evaluate the proposed estimation procedure via simulation study and illustrate its use on a dataset derived by anonymizing testing data from The Ohio State University.

stat.ME