SearcharxivSearch

arXiv subjects

Abel Rodriguez

Publications and source records attributed to Abel Rodriguez.

At least 19 recordsLinked to original sources

Modeling Ordinal Survey Data with Unfolding Models

Surveys that rely on ordinal polychotomous (Likert-like) items are widely employed to capture individual preferences because they allow respondents to express both the direction and strength of their preferences. Latent factor models traditionally used in this context implicitly assume that the response functions (the cumulative distribution of the ordinal outcome) are monotonic on the latent trait. This assumption can be too restrictive in several application areas, including in political science and marketing. In this work, we propose a novel ordinal probit unfolding model that can accommodate both monotonic and non-monotonic response functions. The advantages of the model are illustrated by analyzing an immigration attitude survey conducted in the United States.

stat.ME

"Rebuilding" Statistics in the Age of AI: A Town Hall Discussion on Culture, Infrastructure, and Training

This article presents the full, original record of the 2024 Joint Statistical Meetings (JSM) town hall, "Statistics in the Age of AI," which convened leading statisticians to discuss how the field is evolving in response to advances in artificial intelligence, foundation models, large-scale empirical modeling, and data-intensive infrastructures. The town hall was structured around open panel discussion and extensive audience Q&A, with the aim of eliciting candid, experience-driven perspectives rather than formal presentations or prepared statements. This document preserves the extended exchanges among panelists and audience members, with minimal editorial intervention, and organizes the conversation around five recurring questions concerning disciplinary culture and practices, data curation and "data work," engagement with modern empirical modeling, training for large-scale AI applications, and partnerships with key AI stakeholders. By providing an archival record of this discussion, the preprint aims to support transparency, community reflection, and ongoing dialogue about the evolving role of statistics in the data- and AI-centric future.

stat.ML

pumBayes: Bayesian Estimation of Probit Unfolding Models for Binary Preference Data in R

Probit unfolding models (PUMs) are a novel class of scaling models that allow for items with both monotonic and non-monotonic response functions and have shown great promise in the estimation of preferences from voting data in various deliberative bodies. This paper presents the R package pumBayes, which enables Bayesian inference for both static and dynamic PUMs using Markov chain Monte Carlo algorithms that require minimal or no tuning. In addition to functions that carry out the sampling from the posterior distribution of the models, the package also includes various support functions that can be used to pre-process data, select hyperparameters, summarize output, and compute metrics of model fit. We demonstrate the use of the package through an analysis of two datasets, one corresponding to roll-call voting data from the 116th U.S. House of Representatives, and a second one corresponding to voting records in the U.S. Supreme Court between 1937 and 2021.

stat.CO

Semiparametric estimation for multivariate Hawkes processes using dependent Dirichlet processes: An application to order flow data in financial markets

The order flow in high-frequency financial markets has been of particular research interest in recent years, as it provides insights into trading and order execution strategies and leads to better understanding of the supply-demand interplay and price formation. In this work, we propose a semiparametric multivariate Hawkes process model that relies on (mixtures of) dependent Dirichlet processes to analyze order flow data. Such a formulation avoids the kind of strong parametric assumptions about the excitation functions of the Hawkes process that often accompany traditional models and which, as we show, are not justified in the case of order flow data. It also allows us to borrow information across dimensions, improving estimation of the individual excitation functions. To fit the model, we develop two algorithms, one using Markov chain Monte Carlo methods and one using a stochastic variational approximation. In the context of simulation studies, we show that our model outperforms benchmark methods in terms of lower estimation error for both algorithms. In the context of real order flow data, we show that our model can capture features of the excitation functions such as non-monotonicity that cannot be accommodated by standard parametric models.

stat.ME

Dirichlet process mixtures of block $g$ priors for model selection and prediction in linear models

This paper introduces Dirichlet process mixtures of block $g$ priors for model selection and prediction in linear models. These priors are extensions of traditional mixtures of $g$ priors that allow for differential shrinkage for various (data-selected) blocks of parameters while fully accounting for the predictors' correlation structure, providing a bridge between the literatures on model selection and continuous shrinkage priors. We show that Dirichlet process mixtures of block $g$ priors are consistent in various senses and, in particular, that they avoid the conditional Lindley ``paradox'' highlighted by Som et al. (2016). Further, we develop a Markov chain Monte Carlo algorithm for posterior inference that requires only minimal ad-hoc tuning. Finally, we investigate the empirical performance of the prior in various real and simulated datasets. In the presence of a small number of very large effects, Dirichlet process mixtures of block $g$ priors lead to higher power for detecting smaller but significant effects without only a minimal increase in the number of false discoveries.

stat.ME

Logit unfolding choice models for binary data

Discrete choice models with non-monotonic response functions are important in many areas of application, especially political sciences and marketing. This paper describes a novel unfolding model for binary data that allows for heavy-tailed shocks to the underlying utilities. One of our key contributions is a Markov chain Monte Carlo algorithm that requires little or no parameter tuning, fully explores the support of the posterior distribution, and can be used to fit various extensions of our core model that involve (Bayesian) hypothesis testing on the latent construct. Our empirical evaluations of the model and the associated algorithm suggest that they provide better complexity-adjusted fit to voting data from the United States House of Representatives.

stat.ME

On Data Analysis Pipelines and Modular Bayesian Modeling

The most common approach to implementing data analysis pipelines involves obtaining point estimates from the upstream modules and then treating these as known quantities when working with the downstream ones. This approach is straightforward, but it is likely to underestimate the overall uncertainty associated with any final estimates. An alternative approach involves estimating parameters from the modules jointly using a Bayesian hierarchical model, which has the advantage of propagating upstream uncertainty into the downstream estimates. However, when modules are misspecified, such a joint model can behave in unexpected ways. Furthermore, hierarchical models require the development of ad-hoc computational implementations that can be laborious and computationally expensive. Cut inference modifies the posterior distribution to prevent information flow between certain parameters and provides a third alternative for statistical inference in data analysis pipelines. This paper presents a unified framework that encompasses two-step, cut, and joint inference in the context of data analysis pipelines with two modules and uses two examples to illustrate the tradeoffs associated with these approaches. Our work shows that cut inference provides both some level of robustness and ease of implementation for data analysis pipelines at a lower cost in terms of statistical inference.

stat.ME

Explaining Differences in Voting Patterns Across Voting Domains Using Hierarchical Bayesian Models

Spatial voting models of legislators' preferences are used in political science to test theories about their voting behavior. These models posit that legislators' ideologies as well as the ideologies reflected in votes for and against a bill or measure exist as points in some low dimensional space, and that legislators vote for positions that are close to their own ideologies. Bayesian spatial voting models have been developed to test sharp hypotheses about whether a legislator's revealed ideal point differs for two distinct sets of bills. This project extends such a model to identify covariates that explain whether legislators exhibit such differences in ideal points. We use our method to examine voting behavior on procedural versus final passage votes in the U.S. house of representatives for the 93rd through 113th congresses. The analysis provides evidence that legislators in the minority party as well as legislators with a moderate constituency are more likely to have different ideal points for procedural versus final passage votes.

stat.AP

A Novel Class of Unfolding Models for Binary Preference Data

We develop a new class of spatial voting models for binary preference data that can accommodate both monotonic and non-monotonic response functions, and are more flexible than alternative "unfolding" models previously introduced in the literature. We then use these models to estimate revealed preferences for legislators in the U.S. House of Representatives and justices on the U.S. Supreme Court. The results from these applications indicate that the new models provide superior complexity-adjusted performance to various alternatives and also that the additional flexibility leads to preferences' estimates that are closer matches to the perceived ideological positions of legislators and justices.

stat.AP

Dynamic Factor Models for Binary Data in Circular Spaces: An Application to the U.S. Supreme Court

Latent factor models are widely used in the social and behavioral science as scaling tools to map discrete multivariate outcomes into low dimensional, continuous scales. In political science, dynamic versions of classical factor models have been widely used to study the evolution of justices' preferences in multi-judge courts. In this paper, we discuss a new dynamic factor model that relies on a latent circular space that can accommodate voting behaviors in which justices commonly understood to be on opposite ends of the ideological spectrum vote together on a substantial number of otherwise closely-divided opinions. We apply this model to data on non-unanimous decisions made by the U.S. Supreme Court between 1937 and 2021, and show that, for most of this period, voting patterns can be better described by a circular latent space.

stat.AP

Laplace Power-expected-posterior priors for generalized linear models with applications to logistic regression

Power-expected-posterior (PEP) methodology, which borrows ideas from the literature on power priors, expected-posterior priors and unit information priors, provides a systematic way to construct objective priors. The basic idea is to use imaginary training samples to update a noninformative prior into a minimally-informative prior. In this work, we develop a novel definition of PEP priors for generalized linear models that relies on a Laplace expansion of the likelihood of the imaginary training sample. This approach has various computational, practical and theoretical advantages over previous proposals for non-informative priors for generalized linear models. We place a special emphasis on logistic regression models, where sample separation presents particular challenges to alternative methodologies. We investigate both asymptotic and finite-sample properties of the procedures, showing that is both asymptotic and intrinsic consistent, and that its performance is at least competitive and, in some settings, superior to that of alternative approaches in the literature.

stat.ME

Computational strategies and estimation performance with Bayesian semiparametric Item Response Theory models

Item response theory (IRT) models typically rely on a normality assumption for subject-specific latent traits, which is often unrealistic in practice. Semiparametric extensions based on Dirichlet process mixtures offer a more flexible representation of the unknown distribution of the latent trait. However, the use of such models in the IRT literature has been extremely limited, in good part because of the lack of comprehensive studies and accessible software tools. This paper provides guidance for practitioners on semiparametric IRT models and their implementation. In particular, we rely on NIMBLE, a flexible software system for hierarchical models that enables the use of Dirichlet process mixtures. We highlight efficient sampling strategies for model estimation and compare inferential results under parametric and semiparametric models.

stat.ME

High Dimensional Bayesian Network Classification with Network Global-Local Shrinkage Priors

This article proposes a novel Bayesian classification framework for networks with labeled nodes. While literature on statistical modeling of network data typically involves analysis of a single network, the recent emergence of complex data in several biological applications, including brain imaging studies, presents a need to devise a network classifier for subjects. This article considers an application from a brain connectome study, where the overarching goal is to classify subjects into two separate groups based on their brain network data, along with identifying influential regions of interest (ROIs) (referred to as nodes). Existing approaches either treat all edge weights as a long vector or summarize the network information with a few summary measures. Both these approaches ignore the full network structure, may lead to less desirable inference in small samples and are not designed to identify significant network nodes. We propose a novel binary logistic regression framework with the network as the predictor and a binary response, the network predictor coefficient being modeled using a novel class global-local shrinkage priors. The framework is able to accurately detect nodes and edges in the network influencing the classification. Our framework is implemented using an efficient Markov Chain Monte Carlo algorithm. Theoretically, we show asymptotically optimal classification for the proposed framework when the number of network edges grows faster than the sample size. The framework is empirically validated by extensive simulation studies and analysis of a brain connectome data.

stat.ME

A Bayesian Approach to Spherical Factor Analysis for Binary Data

Factor models are widely used across diverse areas of application for purposes that include dimensionality reduction, covariance estimation, and feature engineering. Traditional factor models can be seen as an instance of linear embedding methods that project multivariate observations onto a lower dimensional Euclidean latent space. This paper discusses a new class of geometric embedding models for multivariate binary data in which the embedding space correspond to a spherical manifold, with potentially unknown dimension. The resulting models include traditional factor models as a special case, but provide additional flexibility. Furthermore, unlike other techniques for geometric embedding, the models are easy to interpret, and the uncertainty associated with the latent features can be properly quantified. These advantages are illustrated using both simulation studies and real data on voting records from the U.S. Senate.

stat.ME

A Bayesian Approach for De-duplication in the Presence of Relational Data

In this paper, we study the impact of combining profile and network data in a de-duplication setting. We also assess the influence of a range of prior distributions on the linkage structure. Furthermore, we explore stochastic gradient Hamiltonian Monte Carlo methods as a faster alternative to obtain samples from the posterior distribution for network parameters. Our methodology is evaluated using the RLdata500 data, which is a popular dataset in the record linkage literature.

stat.ME

A Record Linkage Model Incorporating Relational Data

In this paper we introduce a novel Bayesian approach for linking multiple social networks in order to discover the same real world person having different accounts across networks. In particular, we develop a latent model that allow us to jointly characterize the network and linkage structures relying in both relational and profile data. In contrast to other existing approaches in the machine learning literature, our Bayesian implementation naturally provides uncertainty quantification via posterior probabilities for the linkage structure itself or any function of it. Our findings clearly suggest that our methodology can produce accurate point estimates of the linkage structure even in the absence of profile information, and also, in an identity resolution setting, our results confirm that including relational data into the matching process improves the linkage accuracy. We illustrate our methodology using real data from popular social networks such as Twitter, Facebook, and YouTube.

stat.AP

Bayesian Regression with Undirected Network Predictors with an Application to Brain Connectome Data

This article proposes a Bayesian approach to regression with a continuous scalar response and an undirected network predictor. Undirected network predictors are often expressed in terms of symmetric adjacency matrices, with rows and columns of the matrix representing the nodes, and zero entries signifying no association between two corresponding nodes. Network predictor matrices are typically vectorized prior to any analysis, thus failing to account for the important structural information in the network. This results in poor inferential and predictive performance in presence of small sample sizes. We propose a novel class of network shrinkage priors for the coefficient corresponding to the undirected network predictor. The proposed framework is devised to detect both nodes and edges in the network predictive of the response. Our framework is implemented using an efficient Markov Chain Monte Carlo algorithm. Empirical results in simulation studies illustrate strikingly superior inferential and predictive gains of the proposed framework in comparison with the ordinary high dimensional Bayesian shrinkage priors and penalized optimization schemes. We apply our method to a brain connectome dataset that contains information on brain networks along with a measure of creativity for multiple individuals. Here, interest lies in building a regression model of the creativity measure on the network predictor to identify important regions and connections in the brain strongly associated with creativity. To the best of our knowledge, our approach is the first principled Bayesian method that is able to detect scientifically interpretable regions and connections in the brain actively impacting the continuous response (creativity) in the presence of a small sample size.

stat.ME

A Latent Space Model for Cognitive Social Structures Data

This paper introduces a novel approach for modeling a set of directed, binary networks in the context of cognitive social structures (CSSs) data. We adopt a relativist approach in which no assumption is made about the existence of an underlying true network. More specifically, we rely on a generalized linear model that incorporates a bilinear structure to model transitivity effects within networks, and a hierarchical specification on the bilinear effects to borrow information across networks. This is a spatial model, in which the perception of each individual about the strength of the relationships can be explained by the perceived position of the actors (themselves and others) on a latent social space. A key goal of the model is to provide a mechanism to formally assess the agreement between each actors' perception of their own social roles with that of the rest of the group. Our experiments with both real and simulated data show that the capabilities of our model are comparable with or, even superior to, other models for CSS data reported in the literature.

stat.ME