SearcharxivSearch

arXiv subjects

Hien Duy Nguyen

Publications and source records attributed to Hien Duy Nguyen.

At least 19 recordsLinked to original sources

A variational framework for modal estimation

Multivariate mode estimation arises in many statistical problems such as inverse problems, multimodal sampling, and density-based clustering, but becomes challenging in moderate to high dimensions, especially when the underlying density is not directly evaluable. We introduce GERVE (Gibbs-measure Entropy-Regularized Variational Estimation), a sample-based method for estimating multivariate modes by approximating Gibbs distributions directly from samples, without estimating or evaluating the density. GERVE uses Gaussian-mixture variational annealing and natural-gradient optimization, producing a mixture concentrated in high-density regions whose component responsibilities also provide a clustering of the observations. We prove theoretical guarantees in two regimes: as the Gibbs temperature goes to zero, the optimal variational mixture concentrates around the global modes of the population density; at fixed positive temperature, we prove existence, consistency, and asymptotic normality of empirical maximizers and propose a bootstrap procedure for uncertainty quantification. Simulations and a real-data experiment show that GERVE accurately recovers modes and produces meaningful clusters.

stat.ME

On the minimax-rate optimality of approximate Bayesian computation in nonparametric problems

Approximate Bayesian computation (ABC) replaces likelihood evaluation by simulation and comparison of observed and synthetic data. We establish minimax-rate guarantees for nonparametric ABC under random-series priors with simulable finite-dimensional coordinates. The contraction theorem uses local prior mass, bounds on ABC acceptance probabilities, and control of prior mass outside a sieve. In fixed-design orthogonal-series regression with centered $g$-and-$k$ errors, an infinite Gaussian series prior with a compact scale hyperprior yields minimax-rate contraction and a minimax-rate clipped posterior mean. In compound Poisson decompounding, only random sums are observed and the target is the underlying jump density. With an unknown count intensity in a fixed compact subinterval of $(0,π/2)$, we prove stability of the zero-count-augmented trigonometric population summaries and use a square-root Gaussian series prior on the space of probability density functions. Over bounded periodic Sobolev classes of smoothness $α>d/2$, a polynomially enlarged synthetic sample yields ABC contraction at rate $n^{-α/(2α+d)}$ and posterior mean squared risk of order $n^{-2α/(2α+d)}$, matching a lower bound for the aggregate-observation model. Rejection-ABC Monte Carlo approximations inherit these rates under sufficient sampling budgets.

math.ST

Closed-form solutions to some generalized variational inference problems

The Donsker--Varadhan formula characterizes the ordinary Bayesian posterior as the solution of an unrestricted $\mathsf{KL}$-regularized variational problem. Generalized variational inference replaces this regularizer by other divergences, but the resulting measure-valued optimization problem is often studied only after restriction to a parametric variational family. This paper studies the unrestricted measure-level problem. Given a measurable space $(\mathcal{Z},\mathfrak{Z})$, a prior probability measure $P$, a measurable loss $\ell:\mathcal{Z}\to(-\infty,\infty]$, a regularization strength $α>0$, and a divergence $\mathsf{D}(Q\Vert P)$, we seek probability measures in \[ \underset{Q\in\mathcal{P}(\mathcal{Z})}{\mathrm{arg\,min}}\left\{\int_{\mathcal{Z}} \ell\,\mathrm{d}Q+α\mathsf{D}(Q\Vert P)\right\}. \] For $f$-divergence penalties we derive a scalar inverse-gradient density formula and a one-dimensional dual identity; the Kullback--Leibler, Cressie--Read, and squared-Hellinger problems are treated as examples. Reverse $f$-divergences and mixed forward/reverse Kullback--Leibler penalties follow from the same separable integral principle. For Bregman divergences between densities we obtain a density-space solution with a scalar mass multiplier, including least-squares, density-power, and Burg/Itakura--Saito examples. For Rényi penalties of order $r>1$ we derive a normalized truncated-power characterization and a threshold equation for every global optimizer. Finite model-weight formulas and simple conjugate Bayesian model illustrations show how these closed forms are realized in practice and differ from the traditional solutions.

math.ST

Consistency of variational approximations under bounded Kullback--Leibler divergence

Variational methods are widely used to approximate posterior distributions in Bayesian inference when exact computation is infeasible. We study when such approximations inherit posterior consistency. Our first result shows that, on a general metric space, a uniform bound on the Kullback--Leibler divergence from the approximating measures to a tight sequence of target measures forces the approximating sequence to be tight. It follows that if the target posteriors converge weakly to a Dirac mass at the true parameter, then any variational sequence with bounded Kullback--Leibler divergence to the targets is also consistent. We also give simple logarithmic-moment conditions that verify this boundedness condition, and illustrate them for smooth generalised posterior distributions.

math.ST

TopoGeoScore: A Self-Supervised Source-Only Geometric Framework for OOD Checkpoint Selection

Out-of-distribution (OOD) robustness is difficult to diagnose when target-domain labels are unavailable. We consider a more restrictive source-only variant of unsupervised accuracy estimation: selecting robust checkpoints using only source-domain representations, with no target samples or target labels. We propose \textbf{TopoGeoScore}, a source-only geometric scorer for label-free OOD checkpoint selection. Given a trained checkpoint, we construct class-conditional mutual $k$-nearest-neighbour graphs from source embeddings and extract three interpretable signals: a torsion-inspired reduced Laplacian log-determinant for global class-manifold complexity, Ollivier--Ricci curvature for local neighbourhood regularity, and higher-order topological summaries for fragmented connectivity, loops, and global--local inconsistency. Instead of fixing their weights by hand, TopoGeoScore learns a non-negative linear score through a self-supervised objective that enforces invariance under approximately geometry-preserving embedding views and separation from structure-breaking views. The score remains interpretable and uses no target-domain samples or labels. Results across CIFAR-based corruption and distribution-shift benchmarks, ImageNet-C, MNLI$\to$HANS transfer, and OGBN-Arxiv suggest that source representations contain measurable global--local--topological evidence of robustness, supporting practical checkpoint selection before deployment under distribution shift.

cs.LG

Bounds on the Number of Modes of a Gaussian Mixture Density

We derive explicit upper bounds for the number of nondegenerate critical points of a $k$-component Gaussian mixture density in $\mathbb{R}^d$, and the number of modes when the modal set is finite, together with lower bounds. By normalizing the critical-point equations by a reference component, for $k\ge2$ we get the direct Pfaffian bound \[ U_{\mathrm{het}}(d,k)=2^{\,d+\binom{k-1}{2}}\left(d+2\min(d,k-1)+1\right)^{k-1}. \] For the same parameter range, an exact elimination augmented by an algebraic reciprocal variable gives the alternative bound \[ U_{\mathrm{aug}}(d,k)= 2^{\binom{k-1}{2}}(d+1)\left((2k-1)d+2k-1\right)^{k-1}. \] Thus, for $k\ge2$, the best critical-point bound is their minimum. A Morse-theoretic argument improves the corresponding finite-mode upper bound to \[ \left\lfloor \frac{\min\{U_{\mathrm{het}}(d,k),U_{\mathrm{aug}}(d,k)\}+1}{2}\right\rfloor. \] In the homoscedastic case, for $k\ge2$, the direct bound improves to \[ U_{\mathrm{hom}}(d,k)=2^{\,d+\binom{k-1}{2}}\left(d+\min(d,k-1)+1\right)^{k-1}, \] an affine-rank reduction replaces $d$ by the affine rank of the component means, and an augmented homoscedastic reduction gives the dimension-free bound \[ U_{\mathrm{aug,hom}}(k)=2^{\binom{k-1}{2}+1}(2k)^{k-1}. \] On the lower-bound side, for $d,k\ge 2$ we obtain \[ L_{\mathrm{bin}}(d,k)=k+\max_{2\le r\le \min(d,k)}\binom{k}{r}, \] together with a padding-product family that in particular implies the linear lower bound $d+k-1$, and a seed-closure principle that packages product and padding constructions. We further give explicit bounds for the number of connected components of the critical set.

math.ST

Bayesian inference with sources of uncertainty: from confidence modelling to sparse estimation

We introduce a general framework that extends Bayesian inference by allowing the researcher to explicitly encode confidence in each source of uncertainty within the model. This mechanism provides a new handle for model design and regularisation control. Building on this framework, we develop a general approach for inducing sparsity in statistical models and illustrate its use in linear and logistic regression, as well as in Bayesian neural networks.

stat.ME

Shifted asymmetric Laplace mixtures of experts

Mixtures of experts (MoE) models provide a flexible framework for modelling heterogeneity in data for regression and model-based clustering and classification. MoE models for regression are typically based on the Gaussian assumption for the expert distributions. To robustify the MoE framework with respect to data exhibiting skewness, heavy tails and outliers, we propose a robust non-normal MoE model using the shifted asymmetric Laplace (SAL) distribution. The proposed SALMoE model overcomes the limitations of the Gaussian MoE model when the observed data are asymmetric and heavy-tailed. Through a combination of the minorization-maximization (MM) algorithm with the classical Expectation-Maximization (EM), we develop a dedicated hybrid EM-MM algorithm to estimate the parameters of the SALMoE model. The EM-MM algorithm is shown to yield a nondecreasing observed log-likelihood. A simulation study demonstrates the robustness and practical utility of the proposed model. Finally, the SALMoE model is applied to two real-world economic datasets.

stat.ME

On rates of convergence for sample average approximations without smoothness

Sample average approximation (SAA) replaces an intractable expected objective by an empirical average and is a basic device of modern stochastic optimization. We develop a rate theory for optimal values and empirical $\varepsilon$-minimizers that does not assume continuity, lower semicontinuity, or smooth perturbation structure of the sample objectives. Working on $\ell^{\infty}(X)$ with the Hoffmann--Jørgensen outer-probability formalism, we show that uniform control of the empirical objective process transfers deterministically to convergence rates for optimal values, excess risks of empirical $\varepsilon$-minimizers, and, under a sharp-growth condition, distances to the expected objective solution set. Combined with the directional differentiability of the infimum functional, this yields weak limits for empirical optimal values at the $n^{-1/2}$ scale. Combined with LILs and maximal inequalities, it yields outer almost-sure and outer-mean rates. The definability, envelope, and VC-subgraph hypotheses are verified for definable discontinuous or non-Lipschitz classes arising in direct $0$--$1$ classification, fixed-architecture neural networks, threshold regression, and non-Lipschitz $\ell_{p}$-type objectives with rational $0<p<1$. Practical sufficient conditions for measurability hypotheses are discussed. Together, the framework extends continuity-based SAA theory to a tame-topological setting.

math.OC

Approximation rates for finite mixtures of location-scale models and fast least-squares estimators

Finite mixture models provide a flexible framework for approximating and estimating multivariate probability densities. We study mixtures formed from translated and rescaled copies of a fixed density kernel and obtain explicit results for both approximation and least-squares estimation. Our main deterministic result is a quantisation theorem showing that, after smoothing the target density at a fixed resolution, the resulting convolution can be compressed into a finite location mixture with controlled error. Combining this with the smoothing bias yields approximation rates in $\mathcal{L}_{p}$ over Sobolev classes. For estimation, we analyse least-squares $\varepsilon$-minimisers over suitably tuned mixture sieves. Under exponential decay of the Fourier transform of the kernel, a matching moment condition, and bounded Sobolev targets, the estimator attains a squared $\mathcal{L}_{2}$ risk bound whose rate matches the Sobolev minimax benchmark up to a logarithmic factor. If, in addition, the kernel is bandlimited, then the same theorem recovers the Sobolev rate $n^{-2s/\left(2s+d\right)}$. We further report a slower convergence rate under weaker VC-type assumptions. At fixed scale, the Fourier-based approach also gives a nearly parametric risk bound for the associated location-mixture class, and the same bandlimited simplification removes the logarithmic correction. In the Gaussian case, this recovers the known Gaussian location-mixture rate. We also prove matching lower bounds on Gaussian convolution submodels, including strict submodels of the Gaussian location-mixture class, and on the tensor-product odd-degree Student-$t$ location-mixture family.

math.ST

Characterisations of Kullback--Leibler approximation by finite Gaussian mixtures

We study the Kullback--Leibler (KL) divergence approximation theory of Gaussian mixture models (GMMs) by isolating an abstract mechanism behind several necessary-and-sufficient statements. The necessity direction is universal: if a density is approximable in KL divergence by finite GMMs, then it must have finite second moment. The sufficient direction is reduced to the construction of approximating GMMs whose likelihood ratios converge pointwise and whose finite log-ratios form a uniformly integrable family. We verify this mechanism on a finite log-moment class of continuous strictly positive target densities, from which bounded, $\mathcal L^p$ $(p>1)$, and Orlicz-dominated subfamilies follow immediately. We also show that a countable-scale support-aware target density class, which allows zero density regions, satisfies the same equivalence. Finally, we give counterexamples showing that the countable-scale class strictly extends the fixed-scale class, that the finite log-moment and countable-scale support-aware classes do not contain one another, and that their union is not exhaustive.

math.ST

Consistency of the Bayesian Information Criterion for Model Selection in Exploratory Factor Analysis

We study model selection by the Bayesian information criterion (BIC) in fixed-dimensional exploratory factor analysis over a fixed finite family of compact covariance classes. Our main result shows that the BIC is strongly consistent for the pseudo-true factor order under misspecification, provided that all globally optimal models share a common pseudo-true covariance set, the population Gaussian criterion has a local quadratic margin away from that set, and the BIC complexity counts are order-separating at the pseudo-true order. The candidate models may have an unknown mean vector, exact-zero restrictions in the loading matrix, and either diagonal or spherical error covariance structures, and the selection target is the smallest candidate factor order that yields the best Gaussian approximation, in Kullback--Leibler divergence, to the data-generating covariance structure. The proof works directly in covariance space, so it does not require a regular loading parametrization and accommodates the familiar singularities caused by rotations and redundant factors. Under correct specification, the assumptions reduce to familiar properties of the true covariance matrix. More generally, the same argument applies to other information criteria whose penalties satisfy the same gap conditions, including several BIC-type modifications.

math.ST

On the identifiability of Dirichlet mixture models

We study identifiability of finite mixtures of Dirichlet distributions on the interior of the simplex. We first prove a shift identity showing that every Dirichlet density can be written as a mixture of $J$ shifted Dirichlet densities, where $J-1$ is the dimension of the simplex support, which yields non-identifiability on the full parameter space. We then show that identifiability is recovered on a fixed-total parameter slice and on restricted box-type regions. On the full parameter space, we prove that any nontrivial linear relation among Dirichlet kernels must involve at least $J$ coefficients sharing a common sign, and deduce that mixtures with fewer than $J$ atoms are identifiable. We further report direct non-identifiability implications for unrestricted finite mixtures of generalized Dirichlet, Dirichlet-multinomial, fixed-topic-matrix latent Dirichlet allocation, Beta-Liouville, and inverted Beta-Liouville models.

math.ST

Approximation by mixtures of multivariate Erlang distributions

We prove that finite multivariate Erlang mixture densities with a common rate parameter are dense in the class of probability densities on $\mathbb{R}_{+}^{d}$ that belong to $L^{p}$, for every dimension $d\in\mathbb{N}$ and every $1\le p<\infty$. The argument is constructive: the one-dimensional Szász--Mirakjan--Kantorovich operator yields Erlang mixture approximations, and its tensor product yields multivariate approximants with a common scale. We then obtain several quantitative consequences. These include compact-set uniform approximation bounds and, under local Hölder conditions of order $α\in(0,1]$, rates of order $n^{-α/2}$ as the common scale $1/n$ tends to zero, whole-domain convergence in weighted sup norms, weighted and unweighted $L^{p}$ rates, and explicit rates for finite mixtures indexed by the number of mixture components. In particular, if the approximating density is required to have at most $K$ mixture components, then on fixed compact cubes we obtain an algebraic rate of order $K^{-α/(2d)}$; in global weighted sup norms we obtain the explicit algebraic component-count rate $K^{-α/[2d(2d+α)]}$; and for $1<p<\infty$ we obtain corresponding weighted $L^{p}$ component-count rates. The results strengthen the weak-approximation theory for multivariate Erlang mixture distributions and yield immediate corollaries for broader classes such as product-gamma mixtures. \noindent\textbf{Keywords:} multivariate Erlang mixtures; Erlang distributions; Szász--Mirakjan--Kantorovich operator; density approximation; weighted $L^{p}$ approximation; approximation rates.

math.ST

Modifications of the BIC for order selection in finite mixture models

Finite mixture models are ubiquitous in modern statistical modeling, and a recurring practical issue is choosing the model order. In \citet[Sankhyā Series A, \textbf62, pp. 49--66]{keribin2000consistent}, the Bayesian information criterion (BIC) was proved consistent in mixtures, but under strong regularity, including high moments and high-order derivatives of the component density. We introduce the $ν$-BIC and $ε$-BIC, which weight the BIC penalty by negligibly small logarithmic factors immaterial in practice. This minor modification yields consistency under substantially weaker conditions, without differentiability and with mild moment assumptions, and we also give a misspecification result: when the truth lies outside the candidate family, any vanishing-penalty IC eventually selects a Kullback--Leibler optimal order among candidates. Finally, we clarify two limitations of consistent IC-based selection in mixtures: there is no universally minimal BIC-scale penalty within our sufficient conditions, and order consistency can conflict with minimax optimality in Hellinger risk. We illustrate the theory for Gaussian mixtures, non-differentiable Laplace mixtures, heavy-tailed $t$-mixtures, and mixtures of regression models.

math.ST

Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts

Processing high-volume, streaming data is increasingly common in modern statistics and machine learning, where batch-mode algorithms are often impractical because they require repeated passes over the full dataset. This has motivated incremental stochastic estimation methods, including the incremental stochastic Expectation-Maximization (EM) algorithm formulated via stochastic approximation. In this work, we revisit and analyze an incremental stochastic variant of the Majorization-Minimization (MM) algorithm, which generalizes incremental stochastic EM as a special case. Our approach relaxes key EM requirements, such as explicit latent-variable representations, enabling broader applicability and greater algorithmic flexibility. We establish theoretical guarantees for the incremental stochastic MM algorithm, proving consistency in the sense that the iterates converge to a stationary point characterized by a vanishing gradient of the objective. We demonstrate these advantages on a softmax-gated mixture of experts (MoE) regression problem, for which no stochastic EM algorithm is available. Empirically, our method consistently outperforms widely used stochastic optimizers, including stochastic gradient descent, root mean square propagation, adaptive moment estimation, and second-order clipped stochastic optimization. These results support the development of new incremental stochastic algorithms, given the central role of softmax-gated MoE architectures in contemporary deep neural networks for heterogeneous data modeling. Beyond synthetic experiments, we also validate practical effectiveness on two real-world datasets, including a bioinformatics study of dent maize genotypes under drought stress that integrates high-dimensional proteomics with ecophysiological traits, where incremental stochastic MM yields stable gains in predictive performance.

stat.ML

Statistical process control via $p$-values

We study statistical process control (SPC) through charting of $p$-values. When in control (IC), any valid sequence $(P_{t})_{t}$ is super-uniform, a requirement that can hold in nonparametric and two-phase designs without parametric modelling of the monitored process. Within this framework, we analyse the Shewhart rule that signals when $P_{t}\leα$. Under super-uniformity alone, and with no assumptions on temporal dependence, we derive universal IC lower bounds for the average run length (ARL) and for the expected time to the $k$th false alarm ($k$-ARL). When conditional super-uniformity holds, these bounds sharpen to the familiar $α^{-1}$ and $kα^{-1}$ rates, giving simple, distribution-free calibration for $p$-value charts. Beyond thresholding, we use merging functions for dependent $p$-values to build EWMA-like schemes that output, at each time $t$, a valid $p$-value for the hypothesis that the process has remained IC up to $t$, enabling smoothing without ad hoc control limits. We also study uniform EWMA processes, giving explicit distribution formulas and left-tail guarantees. Finally, we propose a modular approach to directional and coordinate localisation in multivariate SPC via closed testing, controlling the family-wise error rate at the time of alarm. Numerical examples illustrate the utility and variety of our approach.

stat.ME

On the large-sample limits of some Bayesian model evaluation statistics

Model selection and order selection problems frequently arise in statistical practice. A popular approach to addressing these problems in the frequentist setting involves information criteria based on penalised maxima of log-likelihoods for competing models. In the Bayesian context, similar criteria are employed, replacing the maximised log-likelihoods with posterior expectations of the log-likelihood. Despite their popularity in applications, the large-sample behaviour of these criteria -- such as the deviance information criterion (DIC), Bayesian predictive information criterion (BPIC), and widely applicable Bayesian information criterion (WBIC) -- has received relatively little attention. In this work, we investigate the almost-sure limits of these criteria and establish novel results on posterior and generalised posterior consistency, which are of independent interest. The utility of our theoretical findings is demonstrated via illustrative technical and numerical examples.

math.ST