SearcharxivSearch

arXiv subjects

Giuseppe Arena

Publications and source records attributed to Giuseppe Arena.

4 recordsLinked to original sources

What is your Prior Worth? Effective Sample Size and Sample Size Planning for Gaussian Graphical Models

In Bayesian analysis, the prior effective sample size (ESS) expresses the information carried by a prior distribution in units of observations, quantifying how much independent information the prospective data must provide to outweigh an informative prior elicited from a previous study. For network models such as Gaussian graphical models (GGMs), the prior ESS is not straightforward to compute. The Wishart and G-Wishart priors induce dependence among the entries of the precision matrix, and their informativeness has never been expressed in an interpretable, observation-equivalent unit. As a result, researchers eliciting an informative prior for a GGM have had no principled basis for sample size planning. In this paper, we close this gap by formalizing a pre-data ESS for GGMs under the Wishart and G-Wishart priors. We adapt five ESS estimators to the GGM setting and compute each through two aggregation schemes: a global ESS measure based on a determinant ratio, and a parameterwise version based on a Cholesky decomposition. Building on these measures, we introduce two complementary planning strategies: the Data-to-Prior Information Ratio (DPIR), which determines the sample size at which the data dominate the prior, and a GGM extension of Bayes Factor Design Analysis (BFDA), which determines the sample size required for conclusive edge-based evidence. Simulation studies show that the two procedures target complementary design goals and that the ESS estimators differ systematically in their sensitivity to network structure and geometry. We conclude by outlining extensions to other graphical models, including time-dependent variants, as well as to matrix-variate mixture priors.

stat.ME

Bayesian Inference for Discrete Markov Random Fields Through Coordinate Rescaling

Discrete Markov random fields are undirected graphical models that capture complex conditional dependencies between discrete variables. Conducting exact posterior inference in these models is often computationally challenging because evaluating their normalizing constant requires summation over all possible state configurations, and the size of this state space grows exponentially with the number of variables and their possible states. As a result, exact likelihood-based inference is infeasible in many practical settings, and existing methods, such as Double Metropolis-Hastings or pseudo-likelihood approximations, either scale poorly to large systems or underestimate posterior variability. To address these limitations, we propose a new class of coordinate-rescaling sampling methods that transform pseudo-likelihood-based posteriors toward the target posterior while preserving computational efficiency. The resulting samplers retain scalability while improving uncertainty quantification. In simulation studies, we compare the proposed methods to existing approaches and demonstrate that coordinate-rescaling sampling yields more accurate estimates of posterior variability, providing a scalable and reliable approach to Bayesian inference in discrete MRFs.

stat.ME

Comparing Variable Selection and Model Averaging Methods for Logistic Regression

Model uncertainty is a central challenge in statistical models for binary outcomes such as logistic regression, arising when it is unclear which predictors should be included in the model. Many methods have been proposed to address this issue for logistic regression, but their relative performance under realistic conditions remains poorly understood. We therefore conducted a preregistered, simulation-based comparison of 28 established methods for variable selection and inference under model uncertainty, using 11 empirical datasets spanning a range of sample sizes and number of predictors, in cases both with and without separation. We found that Bayesian model averaging (BMA) methods based on g-priors, particularly g = max(n, p^2), show the strongest overall performance when separation is absent. When separation occurs, penalized likelihood approaches, especially the LASSO, provide the most stable results, while BMA with the local empirical Bayes (EB-local) prior is competitive in both situations. These findings offer practical guidance for applied researchers on how to effectively address model uncertainty in logistic regression in modern empirical and machine learning research.

stat.ME

A Bayesian semi-parametric approach for modeling memory decay in dynamic social networks

In relational event networks, the tendency for actors to interact with each other depends greatly on the past interactions between the actors in a social network. Both the quantity of past interactions and the time that elapsed since the past interactions occurred affect the actors' decision-making to interact with other actors in the network. Recently occurred events generally have a stronger influence on current interaction behavior than past events that occurred a long time ago--a phenomenon known as "memory decay". Previous studies either predefined a short-run and long-run memory or fixed a parametric exponential memory using a predefined half-life period. In real-life relational event networks however it is generally unknown how the memory of actors about the past events fades as time goes by. For this reason it is not recommendable to fix this in an ad hoc manner, but instead we should learn the shape of memory decay from the observed data. In this paper, a novel semi-parametric approach based on Bayesian Model Averaging is proposed for learning the shape of the memory decay without requiring any parametric assumptions. The method is applied to relational event history data among socio-political actors in India.

stat.ME