Searcharxiv⌕ Search

arXiv subjects

Nicola Bariletto

Publications and source records attributed to Nicola Bariletto.

11 recordsLinked to original sources

Partial Differential Equation Barriers to Identifiability in Infinite Mixture Models

We study identifiability of mixing measures in infinite mixture models. We show that, in many common cases, lack of identifiability can be characterized in terms of certain differential structures of the kernel family with respect to its parameters. In our main results, we prove that when the kernel is annihilated by a non-trivial differential or difference-differential operator over the parameter space, there exist infinitely many distinct mixing measures yielding the same mixture density. We give verifiable conditions for such operators to exist, covering many common cases, including the location-scale Gaussian, location-scale Student-t, Gamma, Beta, Dirichlet, negative binomial and non-central Chi-squared families. Furthermore, our conditions apply to any exponential family whose parameter dimension exceeds the dimension of its sufficient statistic and, more generally, to kernels with polynomially-growing score functions. We complement our results with a minimax lower bound on the estimation error for the mixing measure in the Wasserstein distance under non-identifiability. On the flips side, we describe three classes of kernels for which identifiability is preserved and nonparametric statistical inference remains possible.

math.ST↗

On the geometry of weak convergence without total variation convergence

We study some geometric consequences of the discrepancy between weak and total variation convergence of probability measures. We consider a sequence of probability measures on $\mathbb R^d$, admitting densities with respect to the Lebesgue measure, that converge weakly to a limiting measure but stay bounded away from it in total variation distance. We show that the sets on which the sequence passes from below to above the limiting density must grow unboundedly in perimeter, as measured by the $(d-1)$-dimensional Hausdorff measure. Moreover, this growth persists within a fixed compact set, so that it must reflect an increase in the geometric complexity of these sets rather than only an unbounded expansion in ambient space. We further provide a sufficient condition under which the number of connected components of the sets diverges, recovering a behavior that is closely reminiscent of the one-dimensional case, in which the number of oscillations of the sequence of densities around the limit grows without bound. We also show that this condition cannot be dispensed with in general, by means of an explicit sequence of measures in the plane whose passing sets remain connected in a single component at every stage while growing in length and complexity. Another sequence, built from cosine oscillations, illustrates the complementary behavior, in which the number of components diverges.

math.PR↗

Convergence Rates for Latent Mixing Measures in Infinite Homoscedastic Location-Scale Mixture Models

We study posterior contraction rates for mixing measures in homoscedastic location-scale mixture models with infinitely many components. While posterior convergence at the level of densities is well understood, ensuring convergence of the latent mixing measure is more challenging and has remained an open problem in settings where both location and scale parameters are unknown. We address this by deriving novel lower-bounds that connect the $L^1$ distance between mixture densities to discrepancies, based on the Wasserstein distances and the operator norm, between the underlying mixing measures and scale matrices. Our approach combines the dual formulation of the $W_1$ distance with functional-analytic approximation techniques. This leads to general inequalities, whose strength is determined (i) by the smoothness of the mixture kernel via the rate of decay of its characteristic function, and (ii) by a key lower-bound on the $L^1$ metric involving the operator norm discrepancy between scale parameters. Moreover, a novel PDE inversion condition yields a sharper inequality for important ordinary-smooth cases. We specialize these bounds to popular mixtures based on multivariate Gaussian, Cauchy, and Laplace kernels. As a consequence, we obtain first-of-their-kind contraction rates in the context of Dirichlet process mixtures with an unknown scale parameter shared across components. As a byproduct of our inequalities, we can distinguish the convergence behavior of the location mixing measure from that of the scale parameter across a range of kernel choices, leading to nuanced insights into their respective rates.

math.ST↗

On Bayesian Softmax-Gated Mixture-of-Experts Models

Mixture-of-experts models provide a flexible framework for learning complex probabilistic input-output relationships by combining multiple expert models through an input-dependent gating mechanism. These models have become increasingly prominent in modern machine learning, yet their theoretical properties in the Bayesian framework remain largely unexplored. In this paper, we study Bayesian mixture-of-experts models, focusing on the ubiquitous softmax-based gating mechanism. Specifically, we investigate the asymptotic behavior of the posterior distribution for three fundamental statistical tasks: density estimation, parameter estimation, and model selection. First, we establish posterior contraction rates for density estimation, both in the regimes with a fixed, known number of experts and with a random learnable number of experts. We then analyze parameter estimation and derive convergence guarantees based on tailored Voronoi-type losses, which account for the complex identifiability structure of mixture-of-experts models. Finally, we propose and analyze two complementary strategies for selecting the number of experts. Taken together, these results provide one of the first systematic theoretical analyses of Bayesian mixture-of-experts models with softmax gating, and yield several theory-grounded insights for practical model design.

stat.ML↗

Scalable Posterior Uncertainty for Flexible Density-Based Clustering

We introduce a novel framework for uncertainty quantification in clustering that combines martingale posterior distributions with density-based clustering. Unlike classical model-based approaches, which define clusters at the latent level of a mixture model, we treat clusters as explicit functionals of the data-generating density, without assuming any specific parametric form. To characterize density uncertainty, we obtain martingale posterior samples via a predictive resampling scheme driven by model score evaluations. This allows us to leverage state-of-the-art differentiable density estimators, such as normalizing flows, making density resampling efficient in large-scale settings and fully parallelizable on modern GPU hardware. Martingale posterior samples of the clustering structure are then obtained by applying density-based clustering to the density draws, enabling principled inference on any clustering-related quantity. Casting the inference target as a density functional further enables a rigorous theoretical analysis of the procedure's convergence properties. We apply our methodology to image and single-cell RNA sequencing data, demonstrating the computational efficiency afforded by its GPU compatibility as well as its ability to recover meaningful clustering structures, with associated uncertainty, across diverse domains.

stat.ML↗

Conformalized Bayesian Inference, with Applications to Random Partition Models

Bayesian posterior distributions naturally represent parameter uncertainty informed by data. However, when the parameter space is complex, as in many nonparametric settings where it is infinite-dimensional or combinatorially large, standard summaries such as posterior means, credible intervals, or simple notions of multimodality are often unavailable, hindering interpretable posterior uncertainty quantification. We introduce Conformalized Bayesian Inference (CBI), a broadly applicable and computationally efficient framework for posterior inference on nonstandard parameter spaces. CBI yields a point estimate, a credible region with assumption-free posterior coverage guarantees, and a principled analysis of posterior multimodality, requiring only Monte Carlo samples from the posterior and a notion of discrepancy between parameters. The method builds a pseudo-density score for each parameter value, yielding a MAP-like point estimate and a credible region derived from conformal prediction principles. The key conceptual step underlying this construction is the reinterpretation of posterior inference as prediction on the parameter space. A final density-based clustering step identifies representative posterior modes. We investigate a number of theoretical and methodological properties of CBI and demonstrate its practicality, scalability, and versatility in simulated and real data clustering applications with random partition models. An accompanying Python library, cbi_partitions, is available at github.com/nbariletto/cbi_partitions_repo.

stat.ME↗

On A Necessary Condition For Posterior Inconsistency: New Insights From A Classic Counterexample

The consistency of posterior distributions in density estimation is at the core of Bayesian statistical theory. Classical work established sufficient conditions, typically combining KL support with complexity bounds on sieves of high prior mass, to guarantee consistency with respect to the Hellinger distance. Yet no systematic theory explains a widely held belief: under KL support, Hellinger consistency is exceptionally hard to violate. This suggests that existing sufficient conditions, while useful in practice, may overlook some key aspects of posterior behavior. We address this gap by directly investigating what must fail for inconsistency to arise, aiming to identify a substantive necessary condition for Hellinger inconsistency. Our starting point is Andrew Barron's classical counterexample, the only known violation of Hellinger consistency under KL support, which relies on a contrived family of oscillatory densities and a prior with atoms. We show that, within a broad class of models including Barron's, inconsistency requires persistent posterior concentration on densities with exponentially high likelihood ratios. In turn, such behavior demands a prior encoding implausibly precise knowledge of the true, yet unknown data-generating distribution, making inconsistency essentially unattainable in any realistic inference problem. Our results confirm the long-standing intuition that posterior inconsistency in density estimation is not a natural phenomenon, but rather an artifact of pathological prior constructions.

math.ST↗

Posterior Consistency in Parametric Models via a Tighter Notion of Identifiability

We study Bayesian posterior consistency in parametric density models with proper priors, challenging the perception that the problem is settled. Classical results established consistency via MLE convergence under regularity and identifiability assumptions, with the latter taken for granted and rarely examined. We refocus attention on identifiability, showing that inconsistency arises only when the true distribution coincides with a weak limit of model densities in a way that violates identifiability. While such failures occur naturally in nonparametric settings, they are implausible and effectively self-inflicted in parametric models. Our analysis shows that classical regularity conditions are unnecessary: a mild strengthening of identifiability suffices to ensure consistency in parametric models, even when the MLE is inconsistent. We also demonstrate that parametric inconsistency requires carefully engineered, oscillatory model features aligned with the true distribution, which is unlikely to occur without adversarial design. Our findings also clarify the distinct mechanisms behind Bayesian and frequentist inconsistency and advocate for separate theoretical treatments.

math.ST↗

Data-Driven DRO and Economic Decision Theory: An Analytical Synthesis With Bayesian Nonparametric Advancements

We develop an analytical synthesis that bridges data-driven Distributionally Robust Optimization (DRO) and Economic Decision Theory under Ambiguity (DTA). By reinterpreting standard regularization and DRO techniques as data-driven counterparts of ambiguity-averse decision models, we provide a unified framework that clarifies their intrinsic connections. Building on this synthesis, we propose a novel DRO approach that leverages a popular DTA model of smooth ambiguity-averse preferences together with tools from Bayesian nonparametric statistics. Our baseline framework employs Dirichlet Process (DP) posteriors, which naturally extend to heterogeneous data sources via Hierarchical Dirichlet Processes (HDPs), and can be further refined to induce outlier robustness through a procedure that selectively filters poorly-fitting observations during training. Theoretical performance guarantees and convergence results, together with extensive simulations and real-data experiments, illustrate the method's favorable performance in terms of prediction accuracy and stability.

stat.ML↗

Bayesian Nonparametrics Meets Data-Driven Distributionally Robust Optimization

Training machine learning and statistical models often involves optimizing a data-driven risk criterion. The risk is usually computed with respect to the empirical data distribution, but this may result in poor and unstable out-of-sample performance due to distributional uncertainty. In the spirit of distributionally robust optimization, we propose a novel robust criterion by combining insights from Bayesian nonparametric (i.e., Dirichlet process) theory and a recent decision-theoretic model of smooth ambiguity-averse preferences. First, we highlight novel connections with standard regularized empirical risk minimization techniques, among which Ridge and LASSO regressions. Then, we theoretically demonstrate the existence of favorable finite-sample and asymptotic statistical guarantees on the performance of the robust optimization procedure. For practical implementation, we propose and study tractable approximations of the criterion based on well-known Dirichlet process representations. We also show that the smoothness of the criterion naturally leads to standard gradient-based numerical optimization. Finally, we provide insights into the workings of our method by applying it to a variety of tasks based on simulated and real datasets.

stat.ML↗

Quasi-Monte Carlo for 3D Sliced Wasserstein

Monte Carlo (MC) integration has been employed as the standard approximation method for the Sliced Wasserstein (SW) distance, whose analytical expression involves an intractable expectation. However, MC integration is not optimal in terms of absolute approximation error. To provide a better class of empirical SW, we propose quasi-sliced Wasserstein (QSW) approximations that rely on Quasi-Monte Carlo (QMC) methods. For a comprehensive investigation of QMC for SW, we focus on the 3D setting, specifically computing the SW between probability measures in three dimensions. In greater detail, we empirically evaluate various methods to construct QMC point sets on the 3D unit-hypersphere, including the Gaussian-based and equal area mappings, generalized spiral points, and optimizing discrepancy energies. Furthermore, to obtain an unbiased estimator for stochastic optimization, we extend QSW to Randomized Quasi-Sliced Wasserstein (RQSW) by introducing randomness in the discussed point sets. Theoretically, we prove the asymptotic convergence of QSW and the unbiasedness of RQSW. Finally, we conduct experiments on various 3D tasks, such as point-cloud comparison, point-cloud interpolation, image style transfer, and training deep point-cloud autoencoders, to demonstrate the favorable performance of the proposed QSW and RQSW variants.

stat.ML↗