SearcharxivSearch

arXiv subjects

Maarten Marsman

Publications and source records attributed to Maarten Marsman.

At least 19 recordsLinked to original sources

Accelerating Bayesian Variable Selection using Piecewise Deterministic Markov Processes

Bayesian variable selection becomes computationally challenging when models contain many dependent parameters. We study Piecewise Deterministic Markov Process (PDMP) samplers as a continuous-time alternative to conventional Markov chain Monte Carlo for spike-and-slab variable selection. In sticky PDMP samplers, active parameters evolve continuously until they reach zero, where they may remain for a random duration before re-entering the model. While one parameter enters or leaves the model, the remaining parameters continue to evolve along the deterministic flow, offering a potentially advantageous mechanism for exploring posteriors with strongly dependent parameters. We make two methodological contributions. First, we extend existing sticky PDMP methods beyond independent spike-and-slab priors to dependent model priors and dependent slab distributions. Second, we investigate the use of unbiased stochastic gradients to reduce the computational cost of variable selection when the likelihood decomposes into many factors while retaining the same target distribution. We study these extensions to two psychometric models: a Gaussian random intercept cross-lagged panel model and an ordinal Markov random field. For the former, marginalization yields a fixed-dimensional sufficient-statistic representation that permits efficient model evaluation. For the latter, the model factorizes, which enables subsampling over person-node contributions. In simulation studies, we compare ZigZag, Bouncy Particle, and Boomerang dynamics with reversible-jump MCMC, and examine the effects of prior dependence and stochastic-gradient subsampling on sampling efficiency. We illustrate the methodology using data from an empirical study on mental well-being. Finally, we discuss the advantages and challenges when using PDMP samplers for Bayesian variable selection.

stat.ME

What is your Prior Worth? Effective Sample Size and Sample Size Planning for Gaussian Graphical Models

In Bayesian analysis, the prior effective sample size (ESS) expresses the information carried by a prior distribution in units of observations, quantifying how much independent information the prospective data must provide to outweigh an informative prior elicited from a previous study. For network models such as Gaussian graphical models (GGMs), the prior ESS is not straightforward to compute. The Wishart and G-Wishart priors induce dependence among the entries of the precision matrix, and their informativeness has never been expressed in an interpretable, observation-equivalent unit. As a result, researchers eliciting an informative prior for a GGM have had no principled basis for sample size planning. In this paper, we close this gap by formalizing a pre-data ESS for GGMs under the Wishart and G-Wishart priors. We adapt five ESS estimators to the GGM setting and compute each through two aggregation schemes: a global ESS measure based on a determinant ratio, and a parameterwise version based on a Cholesky decomposition. Building on these measures, we introduce two complementary planning strategies: the Data-to-Prior Information Ratio (DPIR), which determines the sample size at which the data dominate the prior, and a GGM extension of Bayes Factor Design Analysis (BFDA), which determines the sample size required for conclusive edge-based evidence. Simulation studies show that the two procedures target complementary design goals and that the ESS estimators differ systematically in their sensitivity to network structure and geometry. We conclude by outlining extensions to other graphical models, including time-dependent variants, as well as to matrix-variate mixture priors.

stat.ME

Reversible Jump MCMC With No Regrets: Bayesian Variable Selection Using Mixtures of Mutually Singular Distributions

Bayesian variable selection requires sampling from a posterior distribution that combines discrete model indicators with continuously varying parameters, a challenge often addressed through reversible jump Markov chain Monte Carlo (RJMCMC). Despite its generality, RJMCMC is widely regarded as difficult to design and implement correctly. We present mixtures of mutually singular (MoMS) distributions as a transparent alternative in which competing models are represented within a single fixed-dimensional parameter space partitioned into mutually singular subspaces. We show that this formulation reproduces the exact spike-and-slab interpretation of Bayesian variable selection and that, under appropriate constructions, MoMS and RJMCMC share the same Metropolis--Hastings acceptance probability. On a benchmark dataset with ten predictors, both methods recover posterior inclusion probabilities that match full enumeration, while MoMS achieves comparable or superior effective sample size per second relative to a carefully engineered RJMCMC scheme. We further illustrate the approach in a mixed-effects logistic regression for a sleep-and-memory experiment and in factor-loading selection for a multidimensional generalized partial credit model. Together, these results show that Bayesian variable selection can be carried out within standard fixed-dimensional Markov chain Monte Carlo methodology -- without regret.

stat.ME

Efficient Bayes Factor Sensitivity Analysis via Posterior Density Ratios

Bayes factor sensitivity analysis examines how the evidence for one hypothesis over another depends on the prior distribution. In complex models, the standard approach refits the model at each hyper-parameter value, and the total computational cost scales linearly in the grid size. We propose a method that recovers the entire sensitivity curve from a single additional model fit. The key identity decomposes the Bayes factor at any hyper-parameter value $\gamma_x$ into an ``anchor'' Bayes factor at a fixed reference $\gamma_0$ and a Savage--Dickey density ratio in an extended model that places a hyper-prior on $\gamma$. Once this extended model is fit, the Bayes factor at any $\gamma_x$ follows from the anchor value and a ratio of two posterior density ordinates. To approximate this ratio, we employ the importance-weighted marginal density estimator (IWMDE). Because the sensitivity parameter enters the model only through the prior distribution on the model parameters, the data likelihood cancels in the IWMDE, reducing it to a simple ratio of prior density evaluations on the MCMC draws, without any additional likelihood computation. The resulting estimator is fast, remains accurate even with small MCMC samples, and substantially outperforms kernel density estimation across the full sensitivity range. The method extends naturally to simultaneous sensitivity over multiple hyper-parameters and to Bayesian model averaging. We illustrate it on a univariate Bayesian $t$-test with exact Bayes factors for validation, a bivariate informed $t$-test, and a Bayesian model-averaged meta-analysis, obtaining accurate sensitivity curves at a fraction of the brute-force cost.

stat.ME

Blume-Capel model: Estimation of a three stable state network for $-\bf 1$, $\bf 0$ and $\bf +1$ data

An extension of the Ising model is proposed as a viable alternative for data with values $-1$, $0$ and $+1$ in the inverse problem, i.e., estimation of the parameters. This model is called the Blume-Capel (BC) model, adapted from physics for small networks. The advantage of the BC model is not only the fact that it is possible to have a neutral (centrist) position on the response scale, but also that this model allows for three stable states. We illustrate magnetisation properties of the BC model using simulations and mean field results. For estimation of the BC parameters, we show that the BC model is part of the exponential family of distributions and show that the model is identified, except for the (inverse) temperature. We then show that combining pseudo-likelihood with lasso yields accurate parameter recovery for the BC model, even in small networks. Moreover, confidence intervals with good coverage properties can be obtained using the desparsified lasso together with sandwich and shrinkage techniques. We apply the methods to data obtained from the online platform \textit{Stemwijzer}, intended to aid people in deciding for whom to vote.

stat.AP

Bayesian Inference for Discrete Markov Random Fields Through Coordinate Rescaling

Discrete Markov random fields are undirected graphical models that capture complex conditional dependencies between discrete variables. Conducting exact posterior inference in these models is often computationally challenging because evaluating their normalizing constant requires summation over all possible state configurations, and the size of this state space grows exponentially with the number of variables and their possible states. As a result, exact likelihood-based inference is infeasible in many practical settings, and existing methods, such as Double Metropolis-Hastings or pseudo-likelihood approximations, either scale poorly to large systems or underestimate posterior variability. To address these limitations, we propose a new class of coordinate-rescaling sampling methods that transform pseudo-likelihood-based posteriors toward the target posterior while preserving computational efficiency. The resulting samplers retain scalability while improving uncertainty quantification. In simulation studies, we compare the proposed methods to existing approaches and demonstrate that coordinate-rescaling sampling yields more accurate estimates of posterior variability, providing a scalable and reliable approach to Bayesian inference in discrete MRFs.

stat.ME

Comparing Variable Selection and Model Averaging Methods for Logistic Regression

Model uncertainty is a central challenge in statistical models for binary outcomes such as logistic regression, arising when it is unclear which predictors should be included in the model. Many methods have been proposed to address this issue for logistic regression, but their relative performance under realistic conditions remains poorly understood. We therefore conducted a preregistered, simulation-based comparison of 28 established methods for variable selection and inference under model uncertainty, using 11 empirical datasets spanning a range of sample sizes and number of predictors, in cases both with and without separation. We found that Bayesian model averaging (BMA) methods based on g-priors, particularly g = max(n, p^2), show the strongest overall performance when separation is absent. When separation occurs, penalized likelihood approaches, especially the LASSO, provide the most stable results, while BMA with the local empirical Bayes (EB-local) prior is competitive in both situations. These findings offer practical guidance for applied researchers on how to effectively address model uncertainty in logistic regression in modern empirical and machine learning research.

stat.ME

The Principle of Redundant Reflection

The fact that redundant information does not update a rational belief implies that rational beliefs are updated using Bayes rule. In the framework of Hild (1998a), this is true under mild conditions for discrete, continuous, and arbitrary measure spaces. We prove this result and illustrate it with two examples.

stat.ME

Relations between networks, regression, partial correlation, and latent variable model

The Gaussian graphical model (GGM) has become a popular tool for analyzing networks of psychological variables. In a recent paper in this journal, Forbes, Wright, Markon, and Krueger (FWMK) voiced the concern that GGMs that are estimated from partial correlations wrongfully remove the variance that is shared by its constituents. If true, this concern has grave consequences for the application of GGMs. Indeed, if partial correlations only capture the unique covariances, then the data that come from a unidimensional latent variable model ULVM should be associated with an empty network (no edges), as there are no unique covariances in a ULVM. We know that this cannot be true, which suggests that FWMK are missing something with their claim. We introduce a connection between the ULVM and the GGM and use that connection to prove that we find a fully-connected and not an empty network associated with a ULVM. We then use the relation between GGMs and linear regression to show that the partial correlation indeed does not remove the common variance.

stat.ME

Measurement error and reliability from a different perspective: Reliability of decision functions and the sum score

We propose an alternative framework for measurement error that allows us to determine reliability at the level of the decision, such as the decision to pass or fail. This framework makes it possible to answer the question which properties of a decision function are relevant to make decisions. The mechanism is relatively simple for dichotomous items: The response to any item has probability $\tfrac{1}{2}(1-\rho)$, with $\rho\in [-1,1]$, to be different from the original response. We show that this framework satisfies the classical axioms (those of Lord and Novick) and can therefore be considered as being aligned with classical test theory. We also show some relations to modern test theory and its connections to graphs. We then connect the idea of reliability at the level of the decision to the idea of how many of the (two-response) items are required to be flipped to change the decision; this is referred to as stability. Finally, we argue that from this perspective the best decision function (with respect to a set of properties including stability) is a weighted sum score with threshold value.

stat.ME

Interpreting the Ising Model: The Input Matters

The Ising model is a model for pairwise interactions between binary variables that has become popular in the psychological sciences. It has been first introduced as a theoretical model for the alignment between positive (+1) and negative (-1) atom spins. In many psychological applications, however, the Ising model is defined on the domain $\{0,1\}$ instead of the classical domain $\{-1,1\}$. While it is possible to transform the parameters of a given Ising model in one domain to obtain a statistically equivalent model in the other domain, the parameters in the two versions of the Ising model lend themselves to different interpretations and imply different dynamics, when studying the Ising model as a dynamical system. In this tutorial paper, we provide an accessible discussion of the interpretation of threshold and interaction parameters in the two domains and show how the dynamics of the Ising model depends on the choice of domain. Finally, we provide a transformation that allows to transform the parameters in an Ising model in one domain into a statistically equivalent Ising model in the other domain.

stat.ME

Intervention in undirected Ising graphs and the partition function

Undirected graphical models have many applications in such areas as machine learning, image processing, and, recently, psychology. Psychopathology in particular has received a lot of attention, where symptoms of disorders are assumed to influence each other. One of the most relevant questions practically is on which symptom (node) to intervene to have the most impact. Interventions in undirected graphical models is equal to conditioning, and so we have available the machinery with the Ising model to determine the best strategy to intervene. In order to perform such calculations the partition function is required, which is computationally difficult. Here we use a Curie-Weiss approach to approximate the partition function in applications of interventions. We show that when the connection weights in the graph are equal within each clique then we obtain exactly the correct partition function. And if the weights vary according to a sub-Gaussian distribution, then the approximation is exponentially close to the correct one. We confirm these results with simulations.

stat.ME

Bayesian Rank-Based Hypothesis Testing for the Rank Sum Test, the Signed Rank Test, and Spearman's $ρ$

Bayesian inference for rank-order problems is frustrated by the absence of an explicit likelihood function. This hurdle can be overcome by assuming a latent normal representation that is consistent with the ordinal information in the data: the observed ranks are conceptualized as an impoverished reflection of an underlying continuous scale, and inference concerns the parameters that govern the latent representation. We apply this generic data-augmentation method to obtain Bayes factors for three popular rank-based tests: the rank sum test, the signed rank test, and Spearman's $ρ_s$.

stat.ME

The physics of (ir)rational choice

Even though classic theories and models of discrete choice pose man as a rational being, it has been shown extensively that people persistently violate rationality in their actual choices. Recent models of decision-making take these violations often (partially) into account, however, a unified framework has not been established. Here we propose such a framework, inspired by the Ising model from statistical physics, and show that representing choice problems as a graph, together with a simple choice process, allows us to explain both rational decisions as well as violations of rationality.

physics.soc-ph

Combine Statistical Thinking With Open Scientific Practice: A Protocol of a Bayesian Research Project

Current developments in the statistics community suggest that modern statistics education should be structured holistically, that is, by allowing students to work with real data and to answer concrete statistical questions, but also by educating them about alternative frameworks, such as Bayesian inference. In this article, we describe how we incorporated such a holistic structure in a Bayesian research project on ordered binomial probabilities. The project was conducted with a group of three undergraduate psychology students who had basic knowledge of Bayesian statistics and programming, but lacked formal mathematical training. The research project aimed to (1) convey the basic mathematical concepts of Bayesian inference; (2) have students experience the entire empirical cycle including collection, analysis, and interpretation of data and (3) teach students open science practices.

stat.AP

Logistic regression and Ising networks: prediction and estimation when violating lasso assumptions

The Ising model was originally developed to model magnetisation of solids in statistical physics. As a network of binary variables with the probability of becoming 'active' depending only on direct neighbours, the Ising model appears appropriate for many other processes. For instance, it was recently applied in psychology to model co-occurrences of mental disorders. It has been shown that the connections between the variables (nodes) in the Ising network can be estimated with a series of logistic regressions. This naturally leads to questions of how well such a model predicts new observations and how well parameters of the Ising model can be estimated using logistic regressions. Here we focus on the high-dimensional setting with more parameters than observations and consider violations of assumptions of the lasso. In particular, we determine the consequences for both prediction and estimation when the sparsity and restricted eigenvalue assumptions are not satisfied. We explain by using the idea of connected copies (extreme multicollinearity) the fact that prediction becomes better when either sparsity or multicollinearity is not satisfied. We illustrate these results with simulations.

math.ST

Bayesian Estimation of Kendall's tau Using a Latent Normal Approach

The rank-based association between two variables can be modeled by introducing a latent normal level to ordinal data. We demonstrate how this approach yields Bayesian inference for Kendall's rank correlation coefficient, improving on a recent Bayesian solution from asymptotic properties of the test statistic.

stat.ME