SearcharxivSearch

arXiv subjects

Andee Kaplan

Publications and source records attributed to Andee Kaplan.

15 recordsLinked to original sources

A Comprehensive Bayesian Approach to Entity Resolution for Data with Multiple Truths

In many applications, from government to ecology, integrating data from diverse and noisy sources is critical for downstream inference. However, a unique identifier to link records cleanly from the same entity may not exist. Entity resolution (also referred to as de-duplication or record linkage) merges such databases to identify duplicates, allowing for more complete data for inference. A multitude of methods from statistics and computer science in recent years assume a single, immutable true value for each variable used in linkage. We argue this restrictive assumption is often violated in practice, leading researchers to discard useful data and biasing downstream inference. For example, a respondent's education status may truly change between two different surveys, yet existing methods would treat at least one observation as a distorted version of the truth, even though both are correct. In this paper, we propose a novel entity resolution model that comprehensively accommodates "multiple truths" by introducing a mixture of exponential family distributions that further handles multiple data types. We provide options to fit the model with Markov chain Monte Carlo routines and variational inference for massive datasets. We demonstrate the method's value via simulation and linking a longitudinal survey of Italian household wealth.

stat.ME

ldmppr: Location Dependent Marked Point Processes in R

In this article, we present $\textbf{ldmppr}$, an R package for estimating, evaluating, simulating from, and visualizing location-dependent marked spatial point processes. To date, it has commonly been assumed that the marks associated with a point process are independent of the locations. However, when dealing with many point processes, such as those arising in forestry applications, the independence assumption proves unreasonable. We introduce a practical framework for generating marked point processes with dependence between the marks and locations. We provide a brief discussion of the theory underpinning our modeling approach and outline the use of the package in a typical scenario involving real data. We highlight the functionality of the package for both generating from and assessing the goodness-of-fit of a given model, enabling users to generate realistic point patterns given a reference pattern or parameter values of interest.

stat.CO

The Fast and the Furious: Tracking the Effect of the Tomoa Skip on Speed Climbing

Sport climbing is an athletic discipline comprised of three sub-disciplines -- lead climbing, bouldering, and speed climbing. These three sub-disciplines have distinct goals, resulting in specialization of athletes into one of the three events. The year 2020 marked the first inclusion of sport climbing in the Olympic Games. While this decision was met with excitement from the climbing community, it was not without controversy. The International Olympic Committee had allocated one set of medals for the entire sport, necessitating the combination of sub-disciplines into one competition. As a result, athletes who specialized in lead and bouldering were forced to train and compete in speed for the first time in their careers. One such athlete was Tomoa Narasaki, a World Champion boulderer, who introduced a new method of approaching the speed event. This approach, deemed the Tomoa Skip (TS), was subsequently adopted by many of the top speed climbers. Concurrently, speed records fell rapidly (from 5.48s in 2017 to 4.90s in 2023). Speed climbing involves ascending a 15m wall containing the same pattern of obstacles. Thus, records can be compared across time. In this paper we investigate the effect of the TS on speed climbing by answering two questions: (1) Did the TS result in a decrease in speed times? and (2) Do climbers who utilize the TS show less consistency? The success of the TS highlights the potential of collaboration between different disciplines of sport, showing athletes of diverse backgrounds may contribute to the evolution of competition.

stat.AP

A Unified Bayesian Framework for Modeling Measurement Error in Multinomial Data

Measurement error in multinomial data is a well-known and well-studied inferential problem that is encountered in many fields, including engineering, biomedical and omics research, ecology, finance, official statistics, and social sciences. Methods developed to accommodate measurement error in multinomial data are typically equipped to handle false negatives or false positives, but not both. We provide a unified framework for accommodating both forms of measurement error using a Bayesian hierarchical approach. We demonstrate the proposed method's performance on simulated data and apply it to acoustic bat monitoring and official crime data.

stat.ME

Generative Filtering for Recursive Bayesian Inference with Streaming Data

In the streaming data setting, where data arrive continuously or in frequent batches and there is no pre-determined amount of total data, Bayesian models can employ recursive updates, incorporating each new batch of data into the model parameters' posterior distribution. Filtering methods are currently used to perform these updates efficiently, however, they suffer from eventual degradation as the number of unique values within the filtered samples decreases. We propose Generative Filtering, a method for efficiently performing recursive Bayesian updates in the streaming setting. Generative Filtering retains the speed of a filtering method while using parallel updates to avoid degenerate distributions after repeated applications. We derive rates of convergence for Generative Filtering and conditions for the use of sufficient statistics instead of fully storing all past data. We investigate the alleviation of filtering degradation through simulation and Ecological species count data.

stat.CO

Fast Bayesian Record Linkage for Streaming Data Contexts

Record linkage is the task of combining records from multiple files which refer to overlapping sets of entities when there is no unique identifying field. In streaming record linkage, files arrive sequentially in time and estimates of links are updated after the arrival of each file. This problem arises in settings such as longitudinal surveys, electronic health records, and online events databases, among others. The challenge in streaming record linkage is to efficiently update parameter estimates as new data arrive. We approach the problem from a Bayesian perspective with estimates calculated from posterior samples of parameters and present methods for updating link estimates after the arrival of a new file that are faster than fitting a joint model with each new data file. In this paper, we generalize a two-file Bayesian Fellegi-Sunter model to the multi-file case and propose two methods to perform streaming updates. We examine the effect of prior distribution on the resulting linkage accuracy as well as the computational trade-offs between the methods when compared to a Gibbs sampler through simulated and real-world survey panel data. We achieve near-equivalent posterior inference at a small fraction of the compute time. Supplemental materials for this article are available online.

stat.CO

Interactive Exploration of Large Dendrograms with Prototypes

Hierarchical clustering is one of the standard methods taught for identifying and exploring the underlying structures that may be present within a data set. Students are shown examples in which the dendrogram, a visual representation of the hierarchical clustering, reveals a clear clustering structure. However, in practice, data analysts today frequently encounter data sets whose large scale undermines the usefulness of the dendrogram as a visualization tool. Densely packed branches obscure structure, and overlapping labels are impossible to read. In this paper we present a new workflow for performing hierarchical clustering via the R package called protoshiny that aims to restore hierarchical clustering to its former role of being an effective and versatile visualization tool. Our proposal leverages interactivity combined with the ability to label internal nodes in a dendrogram with a representative data point (called a prototype). After presenting the workflow, we provide three case studies to demonstrate its utility.

stat.OT

Tracking the Transmission Dynamics of COVID-19 with a Time-Varying Coefficient State-Space Model

The spread of COVID-19 has been greatly impacted by regulatory policies and behavior patterns that vary across counties, states, and countries. Population-level dynamics of COVID-19 can generally be described using a set of ordinary differential equations, but these deterministic equations are insufficient for modeling the observed case rates, which can vary due to local testing and case reporting policies and non-homogeneous behavior among individuals. To assess the impact of population mobility on the spread of COVID-19, we have developed a novel Bayesian time-varying coefficient state-space model for infectious disease transmission. The foundation of this model is a time-varying coefficient compartment model to recapitulate the dynamics among susceptible, exposed, undetected infectious, detected infectious, undetected removed, detected non-infectious, detected recovered, and detected deceased individuals. The infectiousness and detection parameters are modeled to vary by time, and the infectiousness component in the model incorporates information on multiple sources of population mobility. Along with this compartment model, a multiplicative process model is introduced to allow for deviation from the deterministic dynamics. We apply this model to observed COVID-19 cases and deaths in several US states and Colorado counties. We find that population mobility measures are highly correlated with transmission rates and can explain complicated temporal variation in infectiousness in these regions. Additionally, the inferred connections between mobility and epidemiological parameters, varying across locations, have revealed the heterogeneous effects of different policies on the dynamics of COVID-19.

stat.AP

d-blink: Distributed End-to-End Bayesian Entity Resolution

Entity resolution (ER; also known as record linkage or de-duplication) is the process of merging noisy databases, often in the absence of unique identifiers. A major advancement in ER methodology has been the application of Bayesian generative models, which provide a natural framework for inferring latent entities with rigorous quantification of uncertainty. Despite these advantages, existing models are severely limited in practice, as standard inference algorithms scale quadratically in the number of records. While scaling can be managed by fitting the model on separate blocks of the data, such a naïve approach may induce significant error in the posterior. In this paper, we propose a principled model for scalable Bayesian ER, called "distributed Bayesian linkage" or d-blink, which jointly performs blocking and ER without compromising posterior correctness. Our approach relies on several key ideas, including: (i) an auxiliary variable representation that induces a partition of the entities and records into blocks; (ii) a method for constructing well-balanced blocks based on k-d trees; (iii) a distributed partially-collapsed Gibbs sampler with improved mixing; and (iv) fast algorithms for performing Gibbs updates. Empirical studies on six data sets---including a case study on the 2010 Decennial Census---demonstrate the scalability and effectiveness of our approach.

stat.CO

Simulating Markov random fields with a conclique-based Gibbs sampler

For spatial and network data, we consider models formed from a Markov random field (MRF) structure and the specification of a conditional distribution for each observation. Fast simulation from such MRF models is often an important consideration, particularly when repeated generation of large numbers of data sets is required. However, a standard Gibbs strategy for simulating from MRF models involves single-site updates, performed with the conditional univariate distribution of each observation in a sequential manner, whereby a complete Gibbs iteration may become computationally involved even for moderate samples. As an alternative, we describe a general way to simulate from MRF models using Gibbs sampling with "concliques" (i.e., groups of non-neighboring observations). Compared to standard Gibbs sampling, this simulation scheme can be much faster by reducing Gibbs steps and independently updating all observations per conclique at once. The speed improvement depends on the number of concliques relative to the sample size for simulation, and order-of-magnitude speed increases are possible with many MRF models (e.g., having appropriately bounded neighborhoods). We detail the simulation method, establish its validity, and assess its computational performance through numerical studies, where speed advantages are shown for several spatial and network examples.

stat.CO

On the instability and degeneracy of deep learning models

A probability model exhibits instability if small changes in a data outcome result in large, and often unanticipated, changes in probability. This instability is a property of the probability model, given by a distributional form and a given configuration of parameters. For correlated data structures found in several application areas, there is increasing interest in identifying such sensitivity in model probability structure. We consider the problem of quantifying instability for general probability models defined on sequences of observations, where each sequence of length N has a finite number of possible values that can be taken at each point. A sequence of probability models results, indexed by N, and an associated parameter sequence, that accommodates data of expanding dimension. Model instability is formally shown to occur when a certain log-probability ratio under such models grows faster than N. In this case, a one component change in the data sequence can shift probability by orders of magnitude. Also, as instability becomes more extreme, the resulting probability models are shown to tend to degeneracy, placing all their probability on potentially small portions of the sample space. These results on instability apply to large classes of models commonly used in random graphs, network analysis, and machine learning contexts.

math.ST

A Practical Approach to Proper Inference with Linked Data

Entity resolution (ER), comprising record linkage and de-duplication, is the process of merging noisy databases in the absence of unique identifiers to remove duplicate entities. One major challenge of analysis with linked data is identifying a representative record among determined matches to pass to an inferential or predictive task, referred to as the \emph{downstream task}. Additionally, incorporating uncertainty from ER in the downstream task is critical to ensure proper inference. To bridge the gap between ER and the downstream task in an analysis pipeline, we propose five methods to choose a representative (or canonical) record from linked data, referred to as canonicalization. Our methods are scalable in the number of records, appropriate in general data scenarios, and provide natural error propagation via a Bayesian canonicalization stage. The proposed methodology is evaluated on three simulated data sets and one application -- determining the relationship between demographic information and party affiliation in voter registration data from the North Carolina State Board of Elections. We first perform Bayesian ER and evaluate our proposed methods for canonicalization before considering the downstream tasks of linear and logistic regression. Bayesian canonicalization methods are empirically shown to improve downstream inference in both settings through prediction and coverage.

stat.ME

Properties and Bayesian fitting of restricted Boltzmann machines

A restricted Boltzmann machine (RBM) is an undirected graphical model constructed for discrete or continuous random variables, with two layers, one hidden and one visible, and no conditional dependency within a layer. In recent years, RBMs have risen to prominence due to their connection to deep learning. By treating a hidden layer of one RBM as the visible layer in a second RBM, a deep architecture can be created. RBMs are thought to thereby have the ability to encode very complex and rich structures in data, making them attractive for supervised learning. However, the generative behavior of RBMs is largely unexplored and typical fitting methodology does not easily allow for uncertainty quantification in addition to point estimates. In this paper, we discuss the relationship between RBM parameter specification in the binary case and model properties such as degeneracy, instability and uninterpretability. We also describe the associated difficulties that can arise with likelihood-based inference and further discuss the potential Bayes fitting of such (highly flexible) models, especially as Gibbs sampling (quasi-Bayes) methods are often advocated for the RBM model structure.

stat.ML

Designing Modular Software: A Case Study in Introductory Statistics

Modular programming is a development paradigm that emphasizes self-contained, flexible, and independent pieces of functionality. This practice allows new features to be seamlessly added when desired, and unwanted features to be removed, thus simplifying the user-facing view of the software. The recent rise of web-based software applications has presented new challenges for designing an extensible, modular software system. In this paper, we outline a framework for designing such a system, with a focus on reproducibility of the results. We present as a case study a Shiny-based web application called intRo, that allows the user to perform basic data analyses and statistical routines. Finally, we highlight some challenges we encountered, and how to address them, when combining modular programming concepts with reactive programming as used by Shiny.

stat.OT

Putting Down Roots: A Graphical Exploration of Community Attachment

In this paper, we explore the relationships that individuals have with their communities. This work was prepared as part of the ASA Data Expo '13 sponsored by the Graphics Section and the Computing Section, using data provided by the Knight Foundation Soul of the Community survey. The Knight Foundation in cooperation with Gallup surveyed 43,000 people over three years in 26 communities across the United States with the intention of understanding the association between community attributes and the degree of attachment people feel towards their community. These include the different facets of both urban and rural communities, the impact of quality education, and the trend in the perceived economic conditions of a community over time. The goal of our work is to facilitate understanding of why people feel attachment to their communities through the use of an interactive and web-based visualization. We will explain the development and use of web-based interactive graphics, including an overview of the R package Shiny and the JavaScript library D3, focusing on the choices made in producing the visualizations and technical aspects of how they were created. Then we describe the stories about community attachment that unfolded from our analysis.

stat.OT