SearcharxivSearch

arXiv subjects

Jeremy Goldwasser

Publications and source records attributed to Jeremy Goldwasser.

12 recordsLinked to original sources

Estimating Epidemic Rate Parameters: Adaptivity, Bias, and Convolution

Metrics like the case-fatality rate and reproduction number are key descriptors of epidemics from the COVID-19 pandemic to the seasonal flu. In retrospect, these quantities enrich our understanding of infectious disease outbreaks; in real-time, they are absolutely critical to informing public health response. Thus, an important question in epidemiology is how best to estimate such metrics, especially in real-time. This question is complicated by practical considerations like data availability, as well as the fact that the metrics themselves may change as the epidemic unfolds.

stat.AP

Fast, Frequentist Estimation of Epidemic Reproduction Numbers

The effective reproduction number $R_t$ is one of the most important indicators of epidemic dynamics. Estimating $R_t$, typically from case reports or hospitalization counts, poses a challenging inverse problem. One key issue is lag: $R_t$ acts at the moment of transmission, while the data it generates surface days later. To handle this delay and infer recent infections in real time, popular methods take a Bayesian approach, which can be slow and sensitive to prior specification. As an alternative, we propose ConvRt, a frequentist method for retrospective and real-time estimation. ConvRt deconvolves latent infections and then estimates $R_t$ with successive penalized-likelihood steps, using a spline basis to model smooth curves. Across both stylized and data-driven simulations, we demonstrate favorable performance in point estimation, uncertainty quantification, and runtime. Moreover, by untangling smoothness from future projections, ConvRt enables researchers to assess which qualitative narratives about $R_t$ the data support.

stat.AP

On the Equivalence of Instantaneous and Mechanistic Reproduction Numbers

The effective reproduction number ($R_t$) is widely used to track epidemic dynamics in real time. The standard estimation framework uses "instantaneous $R_t$," defined via the renewal equation, which relates new infections to past infections through a generation interval distribution. Compartmental models like SEIR yield a seemingly distinct quantity, "mechanistic $R_t$," based on the effective contact rate and duration of infectiousness. We prove these two definitions are equivalent under homogeneous mixing, the standard assumption in compartmental modeling. We also derive the generation interval distribution implied by SEIR dynamics. A practical consequence is that generation intervals, often treated as assumption-light inputs to renewal equation estimators, in fact encode specific compartmental structure.

q-bio.PE

Estimating Time-Varying Epidemic Severity Rates with Adaptive Deconvolution

Several key metrics in public health convey the probability that a primary event will lead to a more serious secondary event in the future.Several key metrics in public health convey the probability that a primary event will lead to a more serious secondary event in the future. These "severity rates" can change over the course of an epidemic in response to shifting conditions like new therapeutics, variants, or public health interventions. In practice, time-varying parameters such as the case-fatality rate are typically estimated from aggregate count data. Prior work has demonstrated that commonly-used ratio-based estimators can be highly biased, motivating the development of new methods. In this paper, we develop an adaptive deconvolution approach based on approximating a Poisson-binomial model for secondary events, and we regularize the maximum likelihood solution in this model with a trend filtering penalty to produce smooth but locally adaptive estimates of severity rates over time. This enables us to compute severity rates both retrospectively and in real time. Experiments based on COVID-19 death and hospitalization data show that our deconvolution estimator is generally more accurate than the standard ratio-based methods, and displays reasonable robustness to model misspecification.

stat.AP

Unifying Image Counterfactuals and Feature Attributions with Latent-Space Adversarial Attacks

Counterfactuals are a popular framework for interpreting machine learning predictions. These what if explanations are notoriously challenging to create for computer vision models: standard gradient-based methods are prone to produce adversarial examples, in which imperceptible modifications to image pixels provoke large changes in predictions. We introduce a new, easy-to-implement framework for counterfactual images that can flexibly adapt to contemporary advances in generative modeling. Our method, Counterfactual Attacks, resembles an adversarial attack on the representation of the image along a low-dimensional manifold. In addition, given an auxiliary dataset of image descriptors, we show how to accompany counterfactuals with feature attribution that quantify the changes between the original and counterfactual images. These importance scores can be aggregated into global counterfactual explanations that highlight the overall features driving model predictions. While this unification is possible for any counterfactual method, it has particular computational efficiency for ours. We demonstrate the efficacy of our approach with the MNIST and CelebA datasets.

cs.LG

Gaussian Rank Verification

Statistical experiments often seek to identify random variables with the largest population means. This inferential task, known as rank verification, has been well-studied on Gaussian data with equal variances. This work provides the first treatment of the unequal variances case, utilizing ideas from the selective inference literature. We design a hypothesis test that verifies the rank of the largest observed value without losing power due to multiple testing corrections. This test is subsequently extended for two procedures: Identifying some number of correctly-ordered Gaussian means, and validating the top-K set. The testing procedures are validated on NHANES survey data.

stat.ME

Stabilizing Estimates of Shapley Values with Control Variates

Shapley values are among the most popular tools for explaining predictions of blackbox machine learning models. However, their high computational cost motivates the use of sampling approximations, inducing a considerable degree of uncertainty. To stabilize these model explanations, we propose ControlSHAP, an approach based on the Monte Carlo technique of control variates. Our methodology is applicable to any machine learning model and requires virtually no extra computation or modeling effort. On several high-dimensional datasets, we find it can produce dramatic reductions in the Monte Carlo variability of Shapley estimates.

stat.ML

Statistical Significance of Feature Importance Rankings

Feature importance scores are ubiquitous tools for understanding the predictions of machine learning models. However, many popular attribution methods suffer from high instability due to random sampling. Leveraging novel ideas from hypothesis testing, we devise techniques that ensure the most important features are correct with high-probability guarantees. These assess the set of $K$ top-ranked features, as well as the order of its elements. Given a set of local or global importance scores, we demonstrate how to retrospectively verify the stability of the highest ranks. We then introduce two efficient sampling algorithms that identify the $K$ most important features, perhaps in order, with probability exceeding $1-\alpha$. The theoretical justification for these procedures is validated empirically on SHAP and LIME.

stat.ML

Ascle: A Python Natural Language Processing Toolkit for Medical Text Generation

This study introduces Ascle, a pioneering natural language processing (NLP) toolkit designed for medical text generation. Ascle is tailored for biomedical researchers and healthcare professionals with an easy-to-use, all-in-one solution that requires minimal programming expertise. For the first time, Ascle evaluates and provides interfaces for the latest pre-trained language models, encompassing four advanced and challenging generative functions: question-answering, text summarization, text simplification, and machine translation. In addition, Ascle integrates 12 essential NLP functions, along with query and search capabilities for clinical databases. The toolkit, its models, and associated data are publicly available via https://github.com/Yale-LILY/MedGen.

cs.CL

EHRKit: A Python Natural Language Processing Toolkit for Electronic Health Record Texts

The Electronic Health Record (EHR) is an essential part of the modern medical system and impacts healthcare delivery, operations, and research. Unstructured text is attracting much attention despite structured information in the EHRs and has become an exciting research field. The success of the recent neural Natural Language Processing (NLP) method has led to a new direction for processing unstructured clinical notes. In this work, we create a python library for clinical texts, EHRKit. This library contains two main parts: MIMIC-III-specific functions and tasks specific functions. The first part introduces a list of interfaces for accessing MIMIC-III NOTEEVENTS data, including basic search, information retrieval, and information extraction. The second part integrates many third-party libraries for up to 12 off-shelf NLP tasks such as named entity recognition, summarization, machine translation, etc.

cs.CL

Forest Fire Clustering for Single-cell Sequencing with Iterative Label Propagation and Parallelized Monte Carlo Simulation

In the era of single-cell sequencing, there is a growing need to extract insights from data with clustering methods. Here, we introduce Forest Fire Clustering, an efficient and interpretable method for cell-type discovery from single-cell data. Forest Fire Clustering makes minimal prior assumptions and, different from current approaches, calculates a non-parametric posterior probability that each cell is assigned a cell-type label. These posterior distributions allow for the evaluation of a label confidence for each cell and enable the computation of "label entropies," highlighting transitions along developmental trajectories. Furthermore, we show that Forest Fire Clustering can make robust, inductive inferences in an online-learning context and can readily scale to millions of cells. Finally, we demonstrate that our method outperforms state-of-the-art clustering approaches on diverse benchmarks of simulated and experimental data. Overall, Forest Fire Clustering is a useful tool for rare cell type discovery in large-scale single-cell analysis.

cs.LG

Neural Natural Language Processing for Unstructured Data in Electronic Health Records: a Review

Electronic health records (EHRs), digital collections of patient healthcare events and observations, are ubiquitous in medicine and critical to healthcare delivery, operations, and research. Despite this central role, EHRs are notoriously difficult to process automatically. Well over half of the information stored within EHRs is in the form of unstructured text (e.g. provider notes, operation reports) and remains largely untapped for secondary use. Recently, however, newer neural network and deep learning approaches to Natural Language Processing (NLP) have made considerable advances, outperforming traditional statistical and rule-based systems on a variety of tasks. In this survey paper, we summarize current neural NLP methods for EHR applications. We focus on a broad scope of tasks, namely, classification and prediction, word embeddings, extraction, generation, and other topics such as question answering, phenotyping, knowledge graphs, medical dialogue, multilinguality, interpretability, etc.

cs.CL