Searcharxiv⌕ Search

arXiv · 2609.34546

Bayesian Nonparametric Factor Analysis via Marginalized Dirichlet Process Column Clustering with Spike-and-Slab Sparsity

Abstract

We propose a Bayesian nonparametric factor model that infers the number of factors, induces row-wise sparsity, and merges redundant dictionary elements via exact clustering. A Dirichlet process prior is placed on the columns of an overcomplete loading matrix and fully marginalized to an exact Pólya urn, avoiding stick-breaking and auxiliary variables. A spike-and-slab base measure allows entire factors to be exactly zero. The model uniquely combines exact zeros, exchangeability over columns, and exact merging within a single marginalized Dirichlet process, unlike CUSP, MGP, or the beta process. An exact Gibbs sampler with canonical relabeling and parallel C/MPI implementation is developed. We prove posterior contraction at rate $\sqrt{M s_0 \log n / n}$ for the covariance matrix, and in the fixed-dictionary setting obtain the minimax optimal rate $\sqrt{s_0 \log n / n}$ plus underfitting consistency; the overfitting direction is an open conjecture. The spike-and-slab is essential: without it the effective dimension scales as $pM$, yielding a slower rate. Simulations show the method is the only fully adaptive approach to recover the true rank, achieving the smallest covariance, loading, and signal-reconstruction errors, beating an oracle baseline. On van 't Veer breast cancer data ($n=97$, $p=1213$), the posterior concentrates on eight interpretable programmes; seven pass coherence and two pass Bonferroni-corrected Hallmark enrichment.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Durba Bhattacharya, Sourabh Bhattacharya. 2026-09-28. Bayesian Nonparametric Factor Analysis via Marginalized Dirichlet Process Column Clustering with Spike-and-Slab Sparsity. https://arxiv.org/abs/2609.34546

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Univariate-Guided Interaction Modeling

We propose a procedure for sparse regression with pairwise interactions, by generalizing the Univariate Guided Sparse Regression (UniLasso) methodology. A central contribution is our introduction of TripletScan, which screens a pair $(j,k)$ using the coefficient of $X_jX_k$ in the local regression of the response on $1$, $X_j$, $X_k$, and $X_jX_k$. The retained products are incorporated either jointly with the main effects through UniLasso, yielding uniPairs, or after a first-stage main-effects fit, yielding uniPairs-2stage. For the UniLasso components of the procedures, we prove false-positive exclusion and uniform coefficient-error bounds. In simulations and an HIV drug-resistance application, the proposed procedures produce substantially smaller fitted models than competing interaction methods while retaining competitive predictive performance.

stat.ME↗

Factoring A-Optimality into D-Optimality and Sphericity

The D criterion measures the volume of the joint confidence ellipsoid for the linear model coefficients and ignores its shape, so designs with the same D value can estimate individual effects with different variances. Meanwhile, A-optimality minimizes average coefficient variance. With both criteria expressed as information values, A equals D multiplied by a sphericity index for the same ellipsoid. In a fixed coefficient basis, sphericity further factors into coefficient-variance balance and a determinant-based correlation component. In five published screening comparisons, the A-optimal design has a larger correlation component despite slightly poorer variance balance; in three it also has a smaller D value. In a seven-run family of designs that all tie under D, the two with equal coefficient variances have the lowest A value. Both sphericity components can be calculated directly from standard errors and estimate correlations available in statistical software. After whitening by a prediction moment matrix, the same determinant-sphericity factorization applies to the I-criterion.

stat.ME↗

Testing the equality of parameters in fixed and increasing dimension

This paper proposes a general and unified framework for testing the equality of a broad class of parameters, defined as a smooth function of expectations of symmetric kernels, across multiple independent populations. We consider two test statistics, a Wald-type statistic and an ANOVA-type statistic. The asymptotic distribution of the first one is derived under a fixed-dimension regime, whereas the second one is studied under both fixed and increasing-dimension regimes, where the parameter dimension diverges with the sample size. Based on these limiting distributions, we construct test procedures enabling asymptotically exact inference without parametric assumptions. Additionally, an alternative null distribution estimator based on a weighted bootstrap approximation is studied, which is applicable to the ANOVA-type statistic under a fixed-dimension regime. The finite-sample performance and computational efficiency of the proposed procedures are evaluated through an extensive simulation study and a real dataset application.

stat.ME↗