SearcharxivSearch

arXiv subjects

Jacob Westerhout

Publications and source records attributed to Jacob Westerhout.

6 recordsLinked to original sources

Closed-form solutions to some generalized variational inference problems

The Donsker--Varadhan formula characterizes the ordinary Bayesian posterior as the solution of an unrestricted $\mathsf{KL}$-regularized variational problem. Generalized variational inference replaces this regularizer by other divergences, but the resulting measure-valued optimization problem is often studied only after restriction to a parametric variational family. This paper studies the unrestricted measure-level problem. Given a measurable space $(\mathcal{Z},\mathfrak{Z})$, a prior probability measure $P$, a measurable loss $\ell:\mathcal{Z}\to(-\infty,\infty]$, a regularization strength $\alpha>0$, and a divergence $\mathsf{D}(Q\Vert P)$, we seek probability measures in \[ \underset{Q\in\mathcal{P}(\mathcal{Z})}{\mathrm{arg\,min}}\left\{\int_{\mathcal{Z}} \ell\,\mathrm{d}Q+\alpha\mathsf{D}(Q\Vert P)\right\}. \] For $f$-divergence penalties we derive a scalar inverse-gradient density formula and a one-dimensional dual identity; the Kullback--Leibler, Cressie--Read, and squared-Hellinger problems are treated as examples. Reverse $f$-divergences and mixed forward/reverse Kullback--Leibler penalties follow from the same separable integral principle. For Bregman divergences between densities we obtain a density-space solution with a scalar mass multiplier, including least-squares, density-power, and Burg/Itakura--Saito examples. For R\'enyi penalties of order $r>1$ we derive a normalized truncated-power characterization and a threshold equation for every global optimizer. Finite model-weight formulas and simple conjugate Bayesian model illustrations show how these closed forms are realized in practice and differ from the traditional solutions.

math.ST

Consistency of variational approximations under bounded Kullback--Leibler divergence

Variational methods are widely used to approximate posterior distributions in Bayesian inference when exact computation is infeasible. We study when such approximations inherit posterior consistency. Our first result shows that, on a general metric space, a uniform bound on the Kullback--Leibler divergence from the approximating measures to a tight sequence of target measures forces the approximating sequence to be tight. It follows that if the target posteriors converge weakly to a Dirac mass at the true parameter, then any variational sequence with bounded Kullback--Leibler divergence to the targets is also consistent. We also give simple logarithmic-moment conditions that verify this boundedness condition, and illustrate them for smooth generalised posterior distributions.

math.ST

On rates of convergence for sample average approximations without smoothness

Sample average approximation (SAA) replaces an intractable expected objective by an empirical average and is a basic device of modern stochastic optimization. We develop a rate theory for optimal values and empirical $\varepsilon$-minimizers that does not assume continuity, lower semicontinuity, or smooth perturbation structure of the sample objectives. Working on $\ell^{\infty}(X)$ with the Hoffmann--J{\o}rgensen outer-probability formalism, we show that uniform control of the empirical objective process transfers deterministically to convergence rates for optimal values, excess risks of empirical $\varepsilon$-minimizers, and, under a sharp-growth condition, distances to the expected objective solution set. Combined with the directional differentiability of the infimum functional, this yields weak limits for empirical optimal values at the $n^{-1/2}$ scale. Combined with LILs and maximal inequalities, it yields outer almost-sure and outer-mean rates. The definability, envelope, and VC-subgraph hypotheses are verified for definable discontinuous or non-Lipschitz classes arising in direct $0$--$1$ classification, fixed-architecture neural networks, threshold regression, and non-Lipschitz $\ell_{p}$-type objectives with rational $0<p<1$. Practical sufficient conditions for measurability hypotheses are discussed. Together, the framework extends continuity-based SAA theory to a tame-topological setting.

math.OC

Approximation rates for finite mixtures of location-scale models and fast least-squares estimators

Finite mixture models provide a flexible framework for approximating and estimating multivariate probability densities. We study mixtures formed from translated and rescaled copies of a fixed density kernel and obtain explicit results for both approximation and least-squares estimation. Our main deterministic result is a quantisation theorem showing that, after smoothing the target density at a fixed resolution, the resulting convolution can be compressed into a finite location mixture with controlled error. Combining this with the smoothing bias yields approximation rates in $\mathcal{L}_{p}$ over Sobolev classes. For estimation, we analyse least-squares $\varepsilon$-minimisers over suitably tuned mixture sieves. Under exponential decay of the Fourier transform of the kernel, a matching moment condition, and bounded Sobolev targets, the estimator attains a squared $\mathcal{L}_{2}$ risk bound whose rate matches the Sobolev minimax benchmark up to a logarithmic factor. If, in addition, the kernel is bandlimited, then the same theorem recovers the Sobolev rate $n^{-2s/\left(2s+d\right)}$. We further report a slower convergence rate under weaker VC-type assumptions. At fixed scale, the Fourier-based approach also gives a nearly parametric risk bound for the associated location-mixture class, and the same bandlimited simplification removes the logarithmic correction. In the Gaussian case, this recovers the known Gaussian location-mixture rate. We also prove matching lower bounds on Gaussian convolution submodels, including strict submodels of the Gaussian location-mixture class, and on the tensor-product odd-degree Student-$t$ location-mixture family.

math.ST

Continuity conditions weaker than lower semi-continuity

Lower semi-continuity (\texttt{LSC}) is a critical assumption in many foundational optimisation theory results; however, in many cases, \texttt{LSC} is stronger than necessary. This has led to the introduction of numerous weaker continuity conditions that enable more general theorem statements. In the context of unstructured optimization over topological domains, we collect these continuity conditions from disparate sources and review their applications. As primary outcomes, we prove two comprehensive implication diagrams that establish novel connections between the reviewed conditions. In doing so, we also introduce previously missing continuity conditions and provide new counterexamples.

math.OC

On the large-sample limits of some Bayesian model evaluation statistics

Model selection and order selection problems frequently arise in statistical practice. A popular approach to addressing these problems in the frequentist setting involves information criteria based on penalised maxima of log-likelihoods for competing models. In the Bayesian context, similar criteria are employed, replacing the maximised log-likelihoods with posterior expectations of the log-likelihood. Despite their popularity in applications, the large-sample behaviour of these criteria -- such as the deviance information criterion (DIC), Bayesian predictive information criterion (BPIC), and widely applicable Bayesian information criterion (WBIC) -- has received relatively little attention. In this work, we investigate the almost-sure limits of these criteria and establish novel results on posterior and generalised posterior consistency, which are of independent interest. The utility of our theoretical findings is demonstrated via illustrative technical and numerical examples.

math.ST