SearcharxivSearch

arXiv subjects

Stephen G. Walker

Publications and source records attributed to Stephen G. Walker.

At least 19 recordsLinked to original sources

On the geometry of weak convergence without total variation convergence

We study some geometric consequences of the discrepancy between weak and total variation convergence of probability measures. We consider a sequence of probability measures on $\mathbb R^d$, admitting densities with respect to the Lebesgue measure, that converge weakly to a limiting measure but stay bounded away from it in total variation distance. We show that the sets on which the sequence passes from below to above the limiting density must grow unboundedly in perimeter, as measured by the $(d-1)$-dimensional Hausdorff measure. Moreover, this growth persists within a fixed compact set, so that it must reflect an increase in the geometric complexity of these sets rather than only an unbounded expansion in ambient space. We further provide a sufficient condition under which the number of connected components of the sets diverges, recovering a behavior that is closely reminiscent of the one-dimensional case, in which the number of oscillations of the sequence of densities around the limit grows without bound. We also show that this condition cannot be dispensed with in general, by means of an explicit sequence of measures in the plane whose passing sets remain connected in a single component at every stage while growing in length and complexity. Another sequence, built from cosine oscillations, illustrates the complementary behavior, in which the number of components diverges.

math.PR

Bayesian Nonparametrics: Principles and Practice

This extended preface [to the Book `Bayesian Nonparametrics', Cambridge University Press, 2010, by NL Hjort, CC Holmes, P Mueller, SG Walker] is meant to explain why you are right to be curious about Bayesian nonparametrics -- why you may actually need it and how you can manage to understand it and use it. The preface also serves as an introductory chapter, giving an overview of the aims and contents of the book. We also explain the background for how the book came into existence, delve briefly on the history of the still relatively young field of Bayesian nonparametrics, and offer some concluding remarks, pertaining to various challenges and likely future developments of the area.

stat.ME

A note on kernel density estimators with optimal bandwidths

We show that the cumulative distribution function corresponding to a kernel density estimator with optimal bandwidth lies outside any confidence interval, around the empirical distribution function, with probability tending to 1 as the sample size increases.

math.ST

Posterior inference via Hill's prediction model

This paper is concerned with the construction of prior free posterior distributions which rely on the use of one step ahead predictive distribution functions. These are typically more straightforward to motivate than prior distributions. Recent interest has been with Hill's $A_n$ prediction model through what has become known as conformal prediction. This model predicts the next observation to lie with equal probability in the intervals created by the observed data. The prediction model generates complete data sets which can be used to provide posterior inference on any statistic of interest.

stat.ME

Scalable Posterior Uncertainty for Flexible Density-Based Clustering

We introduce a novel framework for uncertainty quantification in clustering that combines martingale posterior distributions with density-based clustering. Unlike classical model-based approaches, which define clusters at the latent level of a mixture model, we treat clusters as explicit functionals of the data-generating density, without assuming any specific parametric form. To characterize density uncertainty, we obtain martingale posterior samples via a predictive resampling scheme driven by model score evaluations. This allows us to leverage state-of-the-art differentiable density estimators, such as normalizing flows, making density resampling efficient in large-scale settings and fully parallelizable on modern GPU hardware. Martingale posterior samples of the clustering structure are then obtained by applying density-based clustering to the density draws, enabling principled inference on any clustering-related quantity. Casting the inference target as a density functional further enables a rigorous theoretical analysis of the procedure's convergence properties. We apply our methodology to image and single-cell RNA sequencing data, demonstrating the computational efficiency afforded by its GPU compatibility as well as its ability to recover meaningful clustering structures, with associated uncertainty, across diverse domains.

stat.ML

On Stein's Method of Moments and Generalized Score Matching

The Stein class used in method of moments parameter estimation has two functions whichneedtobespecified, forwhichthereisnopersuasiveargumentsforanyparticular choice. We show that by setting one to be the derivative of the density score function with respect to the parameter leads to a generalized score matching estimator with a choice of weight function. However, choosing a suitable weight function for generalized score matching is not straightforward. We show the weight function is equivalent to a transform of the data and using a score estimator, with an optimal transform being to a normal sample, using for example Box-Cox. We compare our proposal with an alternative means by which to handle the weight function, which is to use a generalized method of moment estimator.

stat.ME

On A Necessary Condition For Posterior Inconsistency: New Insights From A Classic Counterexample

The consistency of posterior distributions in density estimation is at the core of Bayesian statistical theory. Classical work established sufficient conditions, typically combining KL support with complexity bounds on sieves of high prior mass, to guarantee consistency with respect to the Hellinger distance. Yet no systematic theory explains a widely held belief: under KL support, Hellinger consistency is exceptionally hard to violate. This suggests that existing sufficient conditions, while useful in practice, may overlook some key aspects of posterior behavior. We address this gap by directly investigating what must fail for inconsistency to arise, aiming to identify a substantive necessary condition for Hellinger inconsistency. Our starting point is Andrew Barron's classical counterexample, the only known violation of Hellinger consistency under KL support, which relies on a contrived family of oscillatory densities and a prior with atoms. We show that, within a broad class of models including Barron's, inconsistency requires persistent posterior concentration on densities with exponentially high likelihood ratios. In turn, such behavior demands a prior encoding implausibly precise knowledge of the true, yet unknown data-generating distribution, making inconsistency essentially unattainable in any realistic inference problem. Our results confirm the long-standing intuition that posterior inconsistency in density estimation is not a natural phenomenon, but rather an artifact of pathological prior constructions.

math.ST

Weighted Support Points from Random Measures: An Interpretable Alternative for Generative Modeling

Support points summarize a large dataset through a smaller set of representative points that can be used for data operations, such as Monte Carlo integration, without requiring access to the full dataset. In this sense, support points offer a compact yet informative representation of the original data. We build on this idea to introduce a generative modeling framework based on random weighted support points, where the randomness arises from a weighting scheme inspired by the Dirichlet process and the Bayesian bootstrap. The proposed method generates diverse and interpretable sample sets from a fixed dataset, without relying on probabilistic modeling assumptions or neural network architectures. We present the theoretical formulation of the method and develop an efficient optimization algorithm based on the Convex--Concave Procedure (CCP). Empirical results on the MNIST and CelebA-HQ datasets show that our approach produces high-quality and diverse outputs at a fraction of the computational cost of black-box alternatives such as Generative Adversarial Networks (GANs) or Denoising Diffusion Probabilistic Models (DDPMs). These results suggest that random weighted support points offer a principled, scalable, and interpretable alternative for generative modeling. A key feature is their ability to produce genuinely interpolative samples that preserve underlying data structure.

stat.ML

Posterior Consistency in Parametric Models via a Tighter Notion of Identifiability

We study Bayesian posterior consistency in parametric density models with proper priors, challenging the perception that the problem is settled. Classical results established consistency via MLE convergence under regularity and identifiability assumptions, with the latter taken for granted and rarely examined. We refocus attention on identifiability, showing that inconsistency arises only when the true distribution coincides with a weak limit of model densities in a way that violates identifiability. While such failures occur naturally in nonparametric settings, they are implausible and effectively self-inflicted in parametric models. Our analysis shows that classical regularity conditions are unnecessary: a mild strengthening of identifiability suffices to ensure consistency in parametric models, even when the MLE is inconsistent. We also demonstrate that parametric inconsistency requires carefully engineered, oscillatory model features aligned with the true distribution, which is unlikely to occur without adversarial design. Our findings also clarify the distinct mechanisms behind Bayesian and frequentist inconsistency and advocate for separate theoretical treatments.

math.ST

Martingale Posteriors from Score Functions

Uncertainty associated with statistical problems arises due to what has not been seen as opposed to what has been seen. Using probability to quantify the uncertainty the task is to construct a probability model for what has not been seen conditional on what has been seen. The traditional Bayesian approach is to use prior distributions for constructing the predictive distributions, though recently a novel approach has used density estimators and the use of martingales to establish convergence of parameter values. In this paper we reply on martingales constructed using score functions. Hence, the method only requires the computing of gradients arising from parametric families of density functions. A key point is that we do not rely on Markov Chain Monte Carlo (MCMC) algorithms, and that the method can be implemented in parallel. We present the theoretical properties of the score driven martingale posterior. Further, we present illustrations under different models and settings.

stat.ME

A general framework for probabilistic model uncertainty

Existing approaches to model uncertainty typically either compare models using a quantitative model selection criterion or evaluate posterior model probabilities having set a prior. In this paper, we propose an alternative strategy which views missing observations as the source of model uncertainty, where the true model would be identified with the complete data. To quantify model uncertainty, it is then necessary to provide a probability distribution for the missing observations conditional on what has been observed. This can be set sequentially using one-step-ahead predictive densities, which recursively sample from the best model according to some consistent model selection criterion. Repeated predictive sampling of the missing data, to give a complete dataset and hence a best model each time, provides our measure of model uncertainty. This approach bypasses the need for subjective prior specification or integration over parameter spaces, addressing issues with standard methods such as the Bayes factor. Predictive resampling also suggests an alternative view of hypothesis testing as a decision problem based on a population statistic, where we directly index the probabilities of competing models. In addition to hypothesis testing, we demonstrate our approach on illustrations from density estimation and variable selection.

stat.ME

Martingale Posterior Distributions for Log-concave Density Functions

The family of log-concave density functions contains various kinds of common probability distributions. Due to the shape restriction, it is possible to find the nonparametric estimate of the density, for example, the nonparametric maximum likelihood estimate (NPMLE). However, the associated uncertainty quantification of the NPMLE is less well developed. The current techniques for uncertainty quantification are Bayesian, using a Dirichlet process prior combined with the use of Markov chain Monte Carlo (MCMC) to sample from the posterior. In this paper, we start with the NPMLE and use a version of the martingale posterior distribution to establish uncertainty about the NPMLE. The algorithm can be implemented in parallel and hence is fast. We prove the convergence of the algorithm by constructing suitable submartingales. We also illustrate results with different models and settings and some real data, and compare our method with that within the literature.

stat.ME

A Bayesian Bootstrap for Mixture Models

This paper proposes a new nonparametric Bayesian bootstrap for a mixture model, by developing the traditional Bayesian bootstrap. We first reinterpret the Bayesian bootstrap, which uses the P\'olya-urn scheme, as a gradient ascent algorithm which associated one-step solver. The key then is to use the same basic mechanism as the Bayesian bootstrap with the switch from a point mass kernel to a continuous kernel. Just as the Bayesian bootstrap works solely from the empirical distribution function, so the new Bayesian bootstrap for mixture models works off the nonparametric maximum likelihood estimator for the mixing distribution. From a theoretical perspective, we prove the convergence and exchangeability of the sample sequences from the algorithm and also illustrate our results with different models and settings and some real data.

stat.ME

A Non-homogeneous Count Process: Marginalizing a Poisson Driven Cox Process

The paper considers a Cox process where the stochastic intensity function for the Poisson data model is itself a non-homogeneous Poisson process. We show that it is possible to obtain the marginal data process, namely a non-homogeneous count process exhibiting over-dispersion. While the intensity function is non-decreasing, it is straightforward to transform the data so that a non-decreasing intensity function is appropriate. We focus on a time series for arrival times of a process and, in particular, we are able to find an exact form for the marginal probability for the observed data, so allowing for an easy to implement estimation algorithm via direct calculations of the likelihood function.

stat.ME

A multidimensional objective prior distribution from a scoring rule

The construction of objective priors is, at best, challenging for multidimensional parameter spaces. A common practice is to assume independence and set up the joint prior as the product of marginal distributions obtained via "standard" objective methods, such as Jeffreys or reference priors. However, the assumption of independence a priori is not always reasonable, and whether it can be viewed as strictly objective is still open to discussion. In this paper, by extending a previously proposed objective approach based on scoring rules for the one dimensional case, we propose a novel objective prior for multidimensional parameter spaces which yields a dependence structure. The proposed prior has the appealing property of being proper and does not depend on the chosen model; only on the parameter space considered.

stat.ME

Bayesian Data Augmentation for Partially Observed Stochastic Compartmental Models

Deterministic compartmental models are predominantly used in the modeling of infectious diseases, though stochastic models are considered more realistic, yet are complicated to estimate due to missing data. In this paper we present a novel algorithm for estimating the stochastic SIR/SEIR epidemic model within a Bayesian framework, which can be readily extended to more complex stochastic compartmental models. Specifically, based on the infinitesimal conditional independence properties of the model, we are able to find a proposal distribution for a Metropolis algorithm which is very close to the correct posterior distribution. As a consequence, rather than perform a Metropolis step updating one missing data point at a time, as in the current benchmark Markov chain Monte Carlo (MCMC) algorithm, we are able to extend our proposal to the entire set of missing observations. This improves the MCMC methods dramatically and makes the stochastic models now a viable modeling option. A number of real data illustrations and the necessary mathematical theory supporting our results are presented.

stat.CO

Uncertainty Quantification and the Marginal MDP Model

The paper presents a new perspective on the mixture of Dirichlet process model which allows the recovery of full and correct uncertainty quantification associated with the full model, even after having integrated out the random distribution function. The implication is that we can run a simple Markov chain Monte Carlo algorithm and subsequently return the original uncertainty which was removed from the integration. This also has the benefit of avoiding more complicated algorithms which do not perform the integration step. Numerous illustrations are presented.

stat.CO

On Integral Theorems and their Statistical Properties

We introduce a class of integral theorems based on cyclic functions and Riemann sums approximating integrals. The Fourier integral theorem, derived as a combination of a transform and inverse transform, arises as a special case. The integral theorems provide natural estimators of density functions via Monte Carlo methods. Assessments of the quality of the density estimators can be used to obtain optimal cyclic functions, alternatives to the sin function, which minimize square integrals. Our proof techniques rely on a variational approach in ordinary differential equations and the Cauchy residue theorem in complex analysis.

stat.CO