SearcharxivSearch

arXiv subjects

Adrian Dobra

Publications and source records attributed to Adrian Dobra.

At least 19 recordsLinked to original sources

Bayesian Gaussian Mixture Modeling for Symmetric Matrix Variate Data

Statistical inference on individual activity networks has been a historically difficult task due to the lack of available data at the appropriate granularity and the complexity of modeling individual mobility patterns. The recent availability of GPS data from individual devices, combined with highly detailed demographic information, suggests that one of these challenges can now be addressed. We introduce a new model which we call the Symmetric Matrix-Variate Normal Mixture Model (STRUCTURED) to estimate how demographic traits influence changes in human activity networks, using sociomatrices that capture the probabilistic spatial overlap between individuals over time. We exploit the commutativity constraint inherent in the symmetric matrix-variate normal distribution to parameterize the column precision matrix as a polynomial of the row precision matrix, reducing the effective parameter space by an order of magnitude. We develop two variants of STRUCTURED: STRUCTURED-FP, which estimates the full polynomial, and STRUCTURED-RJ, which uses reversible-jump MCMC to select a reduced-order parameterization. Simulation studies demonstrate that STRUCTURED-RJ outperforms existing methods in sparse-data regimes, whereas STRUCTURED-FP is preferred when sample sizes are large. We apply the model to GPS-derived sociomatrices of 293 individuals in King County, WA, finding that local crime environments and youth employment density are the dominant demographic factors explaining variation in weekly activity overlap patterns.

stat.ME

An Object-Oriented Spatial Statistics Approach for Human Activity Space Estimation

Human activity spaces are shaped by individual mobility and the built environment, motivating statistical methods that integrate GPS observations with GIS representations of places and routes. We propose a novel methodology to estimate activity spaces in built environments from GPS data within the Object Oriented Spatial Statistics framework. We characterize daily mobility through the distribution of time across spatial polygons and road segments, aiming to capture entity-specific time-use fractions and level-$\gamma$ activity spaces. We develop a time-weighted estimator to handle irregularly sampled GPS observations. We derive an error bound that quantifies the effects of measurement error, nearest-entity misclassification, temporal gaps, boundary crossings, and day-to-day variability. We also develop a map-augmented representation of daily activity patterns, a dwell-time-weighted distance for clustering daily trajectories, and polygon- and road-based stability summaries. Simulation studies and a real-data application demonstrate that the proposed framework recovers concentrated stationary anchors, interpretable travel corridors, and distinct stabilization behavior for dwelling and movement components, supporting the benefits of weighting under irregular sampling. KEYWORDS: GPS data, GIS, human mobility, space-time geography.

stat.ME

Estimating Population Viral Load Contextual Exposure Using GPS-Derived Activity Spaces in Rural South Africa

This article introduces novel methodologies for estimating contextual exposure to HIV population viral load using GPS data. We propose a comprehensive analytical framework comprising (i) local (grid-cell level) estimation of HIV population viral load, (ii) derivation of individual activity spaces from GPS trajectories, and (iii) quantification of contextual exposure to HIV within these activity spaces. We integrate HIV surveillance and sociodemographic survey data with GPS-based mobility data collected in rural KwaZulu-Natal, South Africa, to characterize mobility patterns among young adults aged 20-30 years. Using derived measures of mobility and contextual exposure, we assess whether participants' sex and age systematically influence the magnitude, configuration, and heterogeneity of their mobility patterns. Furthermore, we describe analytical approaches to examine how contextual exposure to HIV evolves as activity spaces extend beyond static residential locations, outlining procedures to identify GPS-tracked participants at elevated risk of HIV acquisition. KEYWORDS: Population viral load exposure; GPS-based mobility analysis; Activity space

stat.AP

Estimation of Contextual Exposure to HIV from GPS Data

We present a comprehensive statistical methodological framework for estimating contextual exposure to HIV that includes local (grid-cell level) estimation of HIV prevalence and human activity space estimation based on GPS data. The development of our framework was necessary to analyze HIV surveillance and sociodemographic survey data in conjunction with GPS data collected in rural KwaZulu-Natal, South Africa, to study the mobility patterns of young people. Based on mobility and contextual exposure measures, we examine whether the sex and age of study participants systematically influence the extent and structure of their mobility patterns. We discuss techniques for investigating how the study participants' contextual exposure to HIV changes as their activity spaces expand beyond residential locations, as well as methods for identifying study participants who may be at increased risk of acquiring HIV. KEYWORDS: Contextual HIV exposure; GPS-based mobility analysis; Activity space; HIV prevalence mapping

stat.ME

Bayesian Inference for Single-factor Graphical Models

We introduce efficient MCMC algorithms for Bayesian inference for single-factor models with correlated residuals where the residuals' distribution is a Gaussian graphical model. We call this family of models single-factor graphical models. We extend single-factor graphical models to datasets that also involve binary and ordinal categorical variables and to the modeling of multiple datasets that are spatially or temporally related. Our models are able to capture multivariate associations through latent factors across time and space, as well as residual conditional dependence structures at each spatial location or time point through Gaussian graphical models. We illustrate the application of single-factor graphical models in simulated and real-world examples.

stat.ME

Modeling Human Spatial Mobility Patterns with the L\'evy Flight Cluster Model

Despite the extensive collection of individual mobility data over the past decade, fueled by the widespread use of GPS-enabled personal devices, the existing statistical literature on estimating human spatial mobility patterns from temporally irregular location data remains limited. In this paper, we introduce the L\'{e}vy Flight Cluster Model (LFCM), a hierarchical Bayesian mixture model designed to analyze an individual's activity distribution. The LFCM can be utilized to determine probabilistic overlaps between individuals' activity patterns and serves as an anonymization tool to generate synthetic location data. We present our methodology using real-world human location data, demonstrating its ability to accurately capture the key characteristics of human movement.

stat.ME

A statistical framework for analyzing activity pattern from GPS data

We introduce a novel statistical framework for analyzing the GPS data of a single individual. Our approach models daily GPS observations as noisy measurements of an underlying random trajectory, enabling the definition of meaningful concepts such as the average GPS density function. We propose estimators for this density function and establish their asymptotic properties. To study human activity patterns using GPS data, we develop a simple movement model based on mixture models for generating random trajectories. Building on this framework, we introduce several analytical tools to explore activity spaces and mobility patterns. We demonstrate the effectiveness of our approach through applications to both simulated and real-world GPS data, uncovering insightful mobility trends.

stat.ME

Bayesian Finite Mixtures of Ising Models

We introduce finite mixtures of Ising models as a novel approach to study multivariate patterns of associations of binary variables. Our proposed models combine the strengths of Ising models and multivariate Bernoulli mixture models. We examine conditions required for the identifiability of Ising mixture models, and develop a Bayesian framework for fitting them. Through simulation experiments and real data examples, we show that Ising mixture models lead to meaningful results for sparse binary contingency tables.

stat.ME

A statistical framework for measuring the temporal stability of human mobility patterns

Despite the growing popularity of human mobility studies that collect GPS location data, the problem of determining the minimum required length of GPS monitoring has not been addressed in the current statistical literature. In this paper we tackle this problem by laying out a theoretical framework for assessing the temporal stability of human mobility based on GPS location data. We define several measures of the temporal dynamics of human spatiotemporal trajectories based on the average velocity process, and on activity distributions in a spatial observation window. We demonstrate the use of our methods with data that comprise the GPS locations of 185 individuals over the course of 18 months. Our empirical results suggest that GPS monitoring should be performed over periods of time that are significantly longer than what has been previously suggested. Furthermore, we argue that GPS study designs should take into account demographic groups. KEYWORDS: Density estimation; global positioning systems (GPS); human mobility; spatiotemporal trajectories; temporal dynamics

stat.OT

Identifying mediating variables with graphical models: an application to the study of causal pathways in people living with HIV

We empirically demonstrate that graphical models can be a valuable tool in the identification of mediating variables in causal pathways. We make use of graphical models to elucidate the causal pathway through which the treatment influences the levels of fatigue and weakness in people living with HIV (PLHIV) based on a secondary analysis of a categorical dataset collected in a behavioral clinical trial: is weakness a mediator for the treatment and fatigue, or is fatigue a mediator for the treatment and weakness? Causal mediation analysis could not offer any definite answers to these questions.\\ KEYWORDS: Contingency tables; graphical models; loglinear models; HIV; mediation

stat.AP

Measuring Human Activity Spaces from GPS Data with Density Ranking and Summary Curves

Activity spaces are fundamental to the assessment of individuals' dynamic exposure to social and environmental risk factors associated with multiple spatial contexts that are visited during activities of daily living. In this paper we survey existing approaches for measuring the geometry, size and structure of activity spaces based on GPS data, and explain their limitations. We propose addressing these shortcomings through a nonparametric approach called density ranking, and also through three summary curves: the mass-volume curve, the Betti number curve, and the persistence curve. We introduce a novel mixture model for human activity spaces, and study its asymptotic properties. We prove that the kernel density estimator which, at the present time, is one of the most widespread methods for measuring activity spaces is not a stable estimator of their structure. We illustrate the practical value of our methods with a simulation study, and with a recently collected GPS dataset that comprises the locations visited by ten individuals over a six months period.

stat.AP

Modeling association in microbial communities with clique loglinear models

There is a growing awareness of the important roles that microbial communities play in complex biological processes. Modern investigation of these often uses next generation sequencing of metagenomic samples to determine community composition. We propose a statistical technique based on clique loglinear models and Bayes model averaging to identify microbial components in a metagenomic sample at various taxonomic levels that have significant associations. We describe the model class, a stochastic search technique for model selection, and the calculation of estimates of posterior probabilities of interest. We demonstrate our approach using data from the Human Microbiome Project and from a study of the skin microbiome in chronic wound healing. Our technique also identifies significant dependencies among microbial components as evidence of possible microbial syntrophy. KEYWORDS: contingency tables, graphical models, model selection, microbiome, next generation sequencing

stat.AP

Loglinear model selection and human mobility

Methods for selecting loglinear models were among Steve Fienberg's research interests since the start of his long and fruitful career. After we dwell upon the string of papers focusing on loglinear models that can be partly attributed to Steve's contributions and influential ideas, we develop a new algorithm for selecting graphical loglinear models that is suitable for analyzing hyper-sparse contingency tables. We show how multi-way contingency tables can be used to represent patterns of human mobility. We analyze a dataset of geolocated tweets from South Africa that comprises 46 million latitude/longitude locations of 476,601 Twitter users that is summarized as a contingency table with 214 variables. KEYWORDS: contingency tables, model selection, human mobility, graphical models, Bayesian structural learning, birth-death processes, pseudo-likelihood

stat.ME

Analyzing Genome-wide Association Study Data with the R Package genMOSS

The R package (R Core Team (2016)) genMOSS is specifically designed for the Bayesian analysis of genome-wide association study data. The package implements the mode oriented stochastic search (MOSS) procedure as well as a simple moving window approach to identify combinations of single nucleotide polymorphisms associated with a response. The prior used in Bayesian computations is the generalized hyper Dirichlet.

stat.CO

Restricted Covariance Priors with Applications in Spatial Statistics

We present a Bayesian model for area-level count data that uses Gaussian random effects with a novel type of G-Wishart prior on the inverse variance--covariance matrix. Specifically, we introduce a new distribution called the truncated G-Wishart distribution that has support over precision matrices that lead to positive associations between the random effects of neighboring regions while preserving conditional independence of non-neighboring regions. We describe Markov chain Monte Carlo sampling algorithms for the truncated G-Wishart prior in a disease mapping context and compare our results to Bayesian hierarchical models based on intrinsic autoregression priors. A simulation study illustrates that using the truncated G-Wishart prior improves over the intrinsic autoregressive priors when there are discontinuities in the disease risk surface. The new model is applied to an analysis of cancer incidence data in Washington State.

stat.ME

Graphical Modeling of Spatial Health Data

The literature on Gaussian graphical models (GGMs) contains two equally rich and equally significant domains of research efforts and interests. The first research domain relates to the problem of graph determination. That is, the underlying graph is unknown and needs to be inferred from the data. The second research domain dominates the applications in spatial epidemiology. In this context GGMs are typically referred to as Gaussian Markov random fields (GMRFs). Here the underlying graph is assumed to be known: the vertices correspond to geographical areas, while the edges are associated with areas that are considered to be neighbors of each other (e.g., if they share a border). We introduce multi-way Gaussian graphical models that unify the statistical approaches to inference for spatiotemporal epidemiology with the literature on general GGMs. The novelty of the proposed work consists of the addition of the G-Wishart distribution to the substantial collection of statistical tools used to model multivariate areal data. As opposed to fixed graphs that describe geography, there is an inherent uncertainty related to graph determination across the other dimensions of the data. Our new class of methods for spatial epidemiology allow the simultaneous use of GGMs to represent known spatial dependencies and to determine unknown dependencies in the other dimensions of the data. KEYWORDS: Gaussian graphical models, Gaussian Markov random fields, spatiotemporal multivariate models

stat.ME

Spatiotemporal Detection of Unusual Human Population Behavior Using Mobile Phone Data

With the aim to contribute to humanitarian response to disasters and violent events, scientists have proposed the development of analytical tools that could identify emergency events in real-time, using mobile phone data. The assumption is that dramatic and discrete changes in behavior, measured with mobile phone data, will indicate extreme events. In this study, we propose an efficient system for spatiotemporal detection of behavioral anomalies from mobile phone data and compare sites with behavioral anomalies to an extensive database of emergency and non-emergency events in Rwanda. Our methodology successfully captures anomalous behavioral patterns associated with a broad range of events, from religious and official holidays to earthquakes, floods, violence against civilians and protests. Our results suggest that human behavioral responses to extreme events are complex and multi-dimensional, including extreme increases and decreases in both calling and movement behaviors. We also find significant temporal and spatial variance in responses to extreme events. Our behavioral anomaly detection system and extensive discussion of results are a significant contribution to the long-term project of creating an effective real-time event detection system with mobile phone data and we discuss the implications of our findings for future research to this end. KEYWORDS: Big data, call detail record, emergency events, human mobility

physics.soc-ph

Measures of Human Mobility Using Mobile Phone Records Enhanced with GIS Data

In the past decade, large scale mobile phone data have become available for the study of human movement patterns. These data hold an immense promise for understanding human behavior on a vast scale, and with a precision and accuracy never before possible with censuses, surveys or other existing data collection techniques. There is already a significant body of literature that has made key inroads into understanding human mobility using this exciting new data source, and there have been several different measures of mobility used. However, existing mobile phone based mobility measures are inconsistent, inaccurate, and confounded with social characteristics of local context. New measures would best be developed immediately as they will influence future studies of mobility using mobile phone data. In this article, we do exactly this. We discuss problems with existing mobile phone based measures of mobility and describe new methods for measuring mobility that address these concerns. Our measures of mobility, which incorporate both mobile phone records and detailed GIS data, are designed to address the spatial nature of human mobility, to remain independent of social characteristics of context, and to be comparable across geographic regions and time. We also contribute a discussion of the variety of uses for these new measures in developing a better understanding of how human mobility influences micro-level human behaviors and well-being, and macro-level social organization and change.

physics.soc-ph