SearcharxivSearch

arXiv subjects

Tyler McCormick

Publications and source records attributed to Tyler McCormick.

16 recordsLinked to original sources

The Epidemiology of Artificial Intelligence

Artificial intelligence (AI) systems increasingly shape how people access health information, make medical decisions, and receive care -- yet epidemiology lacks frameworks for measuring AI exposure or studying its health effects at the population level. Here we argue that AI now functions as a determinant of health and propose a conceptual framework, borrowed from environmental epidemiology, for studying it. We distinguish ambient AI exposure -- algorithmic curation and AI-mediated institutional decisions that affect populations regardless of individual choice -- from personal AI exposure -- direct, volitional use of AI tools. We characterize AI's possible causal roles in epidemiological models, show that existing experimental approaches are inadequate for capturing chronic, population-level effects, and illustrate these ideas with nationally representative US survey data. We discuss implications for study design, health equity, and AI governance.

stat.OT

Spatially Robust Inference with Predicted and Missing at Random Labels

When outcome data are expensive or onerous to collect, scientists increasingly substitute predictions from machine learning and AI models for unlabeled cases, a process which has consequences for downstream statistical inference. While recent methods provide valid uncertainty quantification under independent sampling, real-world applications involve missing at random (MAR) labeling and spatial dependence. For inference in this setting, we propose a doubly robust estimator with cross-fit nuisances. We show that cross-fitting induces fold-level correlation that distorts spatial variance estimators, producing unstable or overly conservative confidence intervals. To address this, we propose a jackknife spatial heteroscedasticity and autocorrelation consistent (HAC) variance correction that separates spatial dependence from fold-induced noise. Under standard identification and dependence conditions, the resulting intervals are asymptotically valid. Simulations and benchmark datasets show substantial improvement in finite-sample calibration, particularly under MAR labeling and clustered sampling.

stat.ML

A Unified Framework for Inference with General Missingness Patterns and Machine Learning Imputation

Pre-trained machine learning (ML) predictions have been increasingly used to complement incomplete data to enable downstream scientific inquiries, but their naive integration risks biased inferences. Recently, multiple methods have been developed to provide valid inference with ML imputations regardless of prediction quality and to enhance efficiency relative to complete-case analyses. However, existing approaches are often limited to missing outcomes under a missing-completely-at-random (MCAR) assumption, failing to handle general missingness patterns (missing in both the outcome and exposures) under the more realistic missing-at-random (MAR) assumption. This paper develops a novel method that delivers a valid statistical inference framework for general Z-estimation problems using ML imputations under the MAR assumption and for general missingness patterns. The core technical idea is to stratify observations by distinct missingness patterns and construct an estimator by appropriately weighting and aggregating pattern-specific information through a masking-and-imputation procedure on the complete cases. We provide theoretical guarantees of asymptotic normality of the proposed estimator and efficiency dominance over weighted complete-case analyses. Practically, the method affords simple implementations by leveraging existing weighted complete-case analysis software. Extensive simulations are carried out to validate theoretical results. A real data example is provided to further illustrate the practical utility of the proposed method. The paper concludes with a brief discussion on practical implications, limitations, and potential future directions.

stat.ME

Unique Rashomon Sets for Robust Active Learning

Collecting labeled data for machine learning models is often expensive and time-consuming. Active learning addresses this challenge by selectively labeling the most informative observations, but when initial labeled data is limited, it becomes difficult to distinguish genuinely informative points from those appearing uncertain primarily due to noise. Ensemble methods like random forests are a powerful approach to quantifying this uncertainty but do so by aggregating all models indiscriminately. This includes poor performing models and redundant models, a problem that worsens in the presence of noisy data. We introduce UNique Rashomon Ensembled Active Learning (UNREAL), which selectively ensembles only distinct models from the Rashomon set, which is the set of nearly optimal models. Restricting ensemble membership to high-performing models with different explanations helps distinguish genuine uncertainty from noise-induced variation. We show that UNREAL achieves faster theoretical convergence rates than traditional active learning approaches and demonstrates empirical improvements of up to 20% in predictive accuracy across five benchmark datasets, while simultaneously enhancing model interpretability.

stat.ML

Data-adaptive exposure thresholds for the Horvitz-Thompson estimator of the Average Treatment Effect in experiments with network interference

Randomized controlled trials often suffer from interference, a violation of the Stable Unit Treatment Values Assumption (SUTVA) in which a unit's treatment assignment affects the outcomes of its neighbors. This interference causes bias in naive estimators of the average treatment effect (ATE). A popular method to achieve unbiasedness is to pair the Horvitz-Thompson estimator of the ATE with a known exposure mapping: a function that identifies which units in a given randomization are not subject to interference. For example, an exposure mapping can specify that any unit with at least $h$-fraction of its neighbors having the same treatment status does not experience interference. However, this threshold $h$ is difficult to elicit from domain experts, and a misspecified threshold can induce bias. In this work, we propose a data-adaptive method to select the "$h$"-fraction threshold that minimizes the mean squared error of the Hortvitz-Thompson estimator. Our method estimates the bias and variance of the Horvitz-Thompson estimator under different thresholds using a linear dose-response model of the potential outcomes. We present simulations illustrating that our method improves upon non-adaptive choices of the threshold. We further illustrate the performance of our estimator by running experiments on a publicly-available Amazon product similarity graph. Furthermore, we demonstrate that our method is robust to deviations from the linear potential outcomes model.

stat.ME

Some models are useful, but for how long?: A decision theoretic approach to choosing when to refit large-scale prediction models

Large-scale prediction models using tools from artificial intelligence (AI) or machine learning (ML) are increasingly common across a variety of industries and scientific domains. Despite their effectiveness, training AI and ML tools at scale can cost tens or hundreds of thousands of dollars (or more); and even after a model is trained, substantial resources must be invested to keep models up-to-date. This paper presents a decision-theoretic framework for deciding when to refit an AI/ML model when the goal is to perform unbiased statistical inference using partially AI/ML-generated data. Drawing on portfolio optimization theory, we treat the decision of {\it recalibrating} a model or statistical inference versus {\it refitting} the model as a choice between ``investing'' in one of two ``assets.'' One asset, recalibrating the model based on another model, is quick and relatively inexpensive but bears uncertainty from sampling and may not be robust to model drift. The other asset, {\it refitting} the model, is costly but removes the drift concern (though not statistical uncertainty from sampling). We present a framework for balancing these two potential investments while preserving statistical validity. We evaluate the framework using simulation and data on electricity usage and predicting flu trends.

stat.ME

Dempster-Shafer P-values: Thoughts on an Alternative Approach for Multinomial Inference

In this paper, we demonstrate that a new measure of evidence we developed called the Dempster-Shafer p-value which allow for insights and interpretations which retain most of the structure of the p-value while covering for some of the disadvantages that traditional p- values face. Moreover, we show through classical large-sample bounds and simulations that there exists a close connection between our form of DS hypothesis testing and the classical frequentist testing paradigm. We also demonstrate how our approach gives unique insights into the dimensionality of a hypothesis test, as well as models the effects of adversarial attacks on multinomial data. Finally, we demonstrate how these insights can be used to analyze text data for public health through an analysis of the Population Health Metrics Research Consortium dataset for verbal autopsies.

stat.ME

Bayesian Active Questionnaire Design for Cause-of-Death Assignment Using Verbal Autopsies

Only about one-third of the deaths worldwide are assigned a medically-certified cause, and understanding the causes of deaths occurring outside of medical facilities is logistically and financially challenging. Verbal autopsy (VA) is a routinely used tool to collect information on cause of death in such settings. VA is a survey-based method where a structured questionnaire is conducted to family members or caregivers of a recently deceased person, and the collected information is used to infer the cause of death. As VA becomes an increasingly routine tool for cause-of-death data collection, the lengthy questionnaire has become a major challenge to the implementation and scale-up of VAs. In this paper, we propose a novel active questionnaire design approach that optimizes the order of the questions dynamically to achieve accurate cause-of-death assignment with the smallest number of questions. We propose a fully Bayesian strategy for adaptive question selection that is compatible with any existing probabilistic cause-of-death assignment methods. We also develop an early stopping criterion that fully accounts for the uncertainty in the model parameters. We also propose a penalized score to account for constraints and preferences of existing question structures. We evaluate the performance of our active designs using both synthetic and real data, demonstrating that the proposed strategy achieves accurate cause-of-death assignment using considerably fewer questions than the traditional static VA survey instruments.

stat.AP

Asymptotically Normal Estimation of Local Latent Network Curvature

Network data, commonly used throughout the physical, social, and biological sciences, consist of nodes (individuals) and the edges (interactions) between them. One way to represent network data's complex, high-dimensional structure is to embed the graph into a low-dimensional geometric space. The curvature of this space, in particular, provides insights about the structure in the graph, such as the propensity to form triangles or present tree-like structures. We derive an estimating function for curvature based on triangle side lengths and the length of the midpoint of a side to the opposing corner. We construct an estimator where the only input is a distance matrix and also establish asymptotic normality. We next introduce a novel latent distance matrix estimator for networks and an efficient algorithm to compute the estimate via solving iterative quadratic programs. We apply this method to the Los Alamos National Laboratory Unified Network and Host dataset and show how curvature estimates can be used to detect a red-team attack faster than naive methods, as well as discover non-constant latent curvature in co-authorship networks in physics. The code for this paper is available at https://github.com/SteveJWR/netcurve, and the methods are implemented in the R package https://github.com/SteveJWR/lolaR.

stat.ME

Sequential Estimation of Temporally Evolving Latent Space Network Models

In this article we focus on dynamic network data which describe interactions among a fixed population through time. We model this data using the latent space framework, in which the probability of a connection forming is expressed as a function of low-dimensional latent coordinates associated with the nodes, and consider sequential estimation of model parameters via Sequential Monte Carlo (SMC) methods. In this setting, SMC is a natural candidate for estimation which offers greater scalability than existing approaches commonly considered in the literature, allows for estimates to be conveniently updated given additional observations and facilitates both online and offline inference. We present a novel approach to sequentially infer parameters of dynamic latent space network models by building on techniques from the high-dimensional SMC literature. Furthermore, we examine the scalability and performance of our approach via simulation, demonstrate the flexibility of our approach to model variants and analyse a real-world dataset describing classroom contacts.

stat.ME

The "given data" paradigm undermines both cultures

Breiman organizes "Statistical modeling: The two cultures" around a simple visual. Data, to the far right, are compelled into a "black box" with an arrow and then catapulted left by a second arrow, having been transformed into an output. Breiman then posits two interpretations of this visual as encapsulating a distinction between two cultures in statistics. The divide, he argues is about what happens in the "black box." In this comment, I argue for a broader perspective on statistics and, in doing so, elevate questions from "before" and "after" the box as fruitful areas for statistical innovation and practice.

stat.ML

Verbal Autopsy in Civil Registration and Vital Statistics: The Symptom-Cause Information Archive

The burden of disease is fundamental to understanding, prioritizing, and monitoring public health interventions. Cause of death is required to calculate the burden of disease, but in many parts of the developing world deaths are neither detected nor given a cause. Verbal autopsy is a feasible way of assigning a cause to a death based on an interview with those who cared for the person who died. As civil registration and vital statistics systems improve in the developing world, verbal autopsy is playing and increasingly important role in providing information about cause of death and the burden of disease. This note motivates the creation of a global symptom-cause archive containing reference deaths with both a verbal autopsy and a cause assigned through an independent mechanism. This archive could provide training and validation data for refining, developing, and testing machine-based algorithms to automate cause assignment from verbal autopsy data. This, in turn, would improve the comparability of machine-assigned causes and provide a means to fine-tune individual cause assignment within specific contexts.

stat.AP

Introducing Bayesian Analysis with $\text{m&m's}^\circledR$: an active-learning exercise for undergraduates

We present an active-learning strategy for undergraduates that applies Bayesian analysis to candy-covered chocolate $\text{m&m's}^\circledR$. The exercise is best suited for small class sizes and tutorial settings, after students have been introduced to the concepts of Bayesian statistics. The exercise takes advantage of the non-uniform distribution of $\text{m&m's}^\circledR~$ colours, and the difference in distributions made at two different factories. In this paper, we provide the intended learning outcomes, lesson plan and step-by-step guide for instruction, and open-source teaching materials. We also suggest an extension to the exercise for the graduate-level, which incorporates hierarchical Bayesian analysis.

stat.OT

Redrawing the 'Color Line': Examining Racial Segregation in Associative Networks on Twitter

Online social spaces are increasingly salient contexts for associative tie formation. However, the racial composition of associative networks within most of these spaces has yet to be examined. In this paper, we use data from the social media platform Twitter to examine racial segregation patterns in online associative networks. Acknowledging past work on the role that social structure and agency play in influencing the racial composition of individuals' networks, we argue that Twitter blurs the influence of these forces and may invite users to generate networks that are both more or less segregated than what has been observed offline, depending on use. While we expect to find some level of racial segregation within this space, this paper unpacks the extent to which we observe same-race connectedness for black and white users, assesses whether these patterns are likely generated by opportunity or by choice, and contextualizes results by comparing them with patterns of same-race connectedness observed offline.

cs.SI

Hyak Mortality Monitoring System: Innovative Sampling and Estimation Methods - Proof of Concept by Simulation

Traditionally health statistics are derived from civil and/or vital registration. Civil registration in low-income countries varies from partial coverage to essentially nothing at all. Consequently the state of the art for public health information in low-income countries is efforts to combine or triangulate data from different sources to produce a more complete picture across both time and space - data amalgamation. Data sources amenable to this approach include sample surveys, sample registration systems, health and demographic surveillance systems, administrative records, census records, health facility records and others. We propose a new statistical framework for gathering health and population data - Hyak - that leverages the benefits of sampling and longitudinal, prospective surveillance to create a cheap, accurate, sustainable monitoring platform. Hyak has three fundamental components: 1) Data Amalgamation: a sampling and surveillance component that organizes two or more data collection systems to work together: a) data from HDSS with frequent, intense, linked, prospective follow-up and b) data from sample surveys conducted in large areas surrounding the Health and Demographic Surveillance System sites using informed sampling so as to capture as many events as possible; 2) Cause of Death: verbal autopsy to characterize the distribution of deaths by cause at the population level; and 3) SES: measurement of socioeconomic status in order to characterize poverty and wealth. We conduct a simulation study of the informed sampling component of Hyak based on the Agincourt HDSS site in South Africa. Compared to traditional cluster sampling, Hyak's informed sampling captures more deaths, and when combined with an estimation model that includes spatial smoothing, produces estimates mortality that have lower variance and small bias.

stat.OT

InSilicoVA: A Method to Automate Cause of Death Assignment for Verbal Autopsy

Verbal autopsies (VA) are widely used to provide cause-specific mortality estimates in developing world settings where vital registration does not function well. VAs assign cause(s) to a death by using information describing the events leading up to the death, provided by care givers. Typically physicians read VA interviews and assign causes using their expert knowledge. Physician coding is often slow, and individual physicians bring bias to the coding process that results in non-comparable cause assignments. These problems significantly limit the utility of physician-coded VAs. A solution to both is to use an algorithmic approach that formalizes the cause-assignment process. This ensures that assigned causes are comparable and requires many fewer person-hours so that cause assignment can be conducted quickly without disrupting the normal work of physicians. Peter Byass' InterVA method is the most widely used algorithmic approach to VA coding and is aligned with the WHO 2012 standard VA questionnaire. The statistical model underpinning InterVA can be improved; uncertainty needs to be quantified, and the link between the population-level CSMFs and the individual-level cause assignments needs to be statistically rigorous. Addressing these theoretical concerns provides an opportunity to create new software using modern languages that can run on multiple platforms and will be widely shared. Building on the overall framework pioneered by InterVA, our work creates a statistical model for automated VA cause assignment.

stat.OT