SearcharxivSearch

arXiv subjects

Curtis McDonald

Publications and source records attributed to Curtis McDonald.

7 recordsLinked to original sources

Technical Note on Relating Scores of Tilted Distributions

Recent results have shown that for a linear tilt to a reference measure, the scores that would be produced under convolution with a normal variable can be expressed in terms of convolutions of the original density. Here, we extend that result to include constant negative diagonal tilts as well. The relationship follows from relating the denoisers of the two densities, which define the scores via Tweedie formula. A linear tilt results in a location shift to the score operator, while a quadratic tilt results in both a location shift and a time shift. Thus the scores of the tilted density can be understood as the scores of the original convolution process at a different location and noise level. These results are of interest to those in the score based diffusion community, and may lead to better score estimators which take advantage of these tilted score relationships.

math.ST

Rapid Bayesian Computation and Estimation for Neural Networks via Log-Concave Coupling

This paper presents the study of a Bayesian estimation procedure for single-hidden-layer neural networks using $\ell_{1}$ controlled neuron weight vectors. We study the structure of the posterior density and provide a representation that makes it amenable to rapid sampling via Markov Chain Monte Carlo (MCMC). Let the neural network have $K$ neurons with internal weights of dimension $d$ and fix the outer weights. Thus there are $Kd$ parameters overall. With $N$ data observations, use a gain parameter or inverse temperature of $\beta$ in the posterior density for the internal weights. The posterior is intrinsically multi-modal and not naturally suited to rapid mixing of direct MCMC algorithms. For a continuous uniform prior on the $\ell_{1}$ ball, we demonstrate that the posterior density can be written as a mixture density with suitably defined auxiliary random variables, where the mixture components are log-concave. Furthermore, when the total number of model parameters $Kd$ is large enough that $Kd \geq C(\beta N)^{2}$, the mixing distribution of the auxiliary random variables is also log-concave. Thus, neuron parameters can be sampled from the posterior by only sampling log-concave densities. The authors refer to the pairing of weights with such auxiliary random variables as a log-concave coupling.

math.ST

Log-Concave Coupling for Sampling Neural Net Posteriors

In this work, we present a sampling algorithm for single hidden layer neural networks. This algorithm is built upon a recursive series of Bayesian posteriors using a method we call Greedy Bayes. Sampling of the Bayesian posterior for neuron weight vectors $w$ of dimension $d$ is challenging because of its multimodality. Our algorithm to tackle this problem is based on a coupling of the posterior density for $w$ with an auxiliary random variable $\xi$. The resulting reverse conditional $w|\xi$ of neuron weights given auxiliary random variable is shown to be log concave. In the construction of the posterior distributions we provide some freedom in the choice of the prior. In particular, for Gaussian priors on $w$ with suitably small variance, the resulting marginal density of the auxiliary variable $\xi$ is proven to be strictly log concave for all dimensions $d$. For a uniform prior on the unit $\ell_1$ ball, evidence is given that the density of $\xi$ is again strictly log concave for sufficiently large $d$. The score of the marginal density of the auxiliary random variable $\xi$ is determined by an expectation over $w|\xi$ and thus can be computed by various rapidly mixing Markov Chain Monte Carlo methods. Moreover, the computation of the score of $\xi$ permits methods of sampling $\xi$ by a stochastic diffusion (Langevin dynamics) with drift function built from this score. With such dynamics, information-theoretic methods pioneered by Bakry and Emery show that accurate sampling of $\xi$ is obtained rapidly when its density is indeed strictly log-concave. After which, one more draw from $w|\xi$, produces neuron weights $w$ whose marginal distribution is from the desired posterior.

stat.ML

Stochastic Observability and Filter Stability under Several Criteria

Despite being a foundational concept of modern systems theory, there have been few studies on observability of non-linear stochastic systems under partial observations. In this paper, we introduce a definition of observability for stochastic non-linear dynamical systems which involves an explicit functional characterization. To justify its operational use, we establish that this definition implies filter stability under mild continuity conditions: an incorrectly initialized non-linear filter is said to be stable if the filter eventually corrects itself with the arrival of new measurement information. Numerous examples are presented and a detailed comparison with the literature is reported. We also establish implications for various criteria for filter stability under several notions of convergence such as weak convergence, total variation, and relative entropy. These findings are connected to robustness and approximations in partially observed stochastic control.

math.PR

Proposal of a Score Based Approach to Sampling Using Monte Carlo Estimation of Score and Oracle Access to Target Density

Score based approaches to sampling have shown much success as a generative algorithm to produce new samples from a target density given a pool of initial samples. In this work, we consider if we have no initial samples from the target density, but rather $0^{th}$ and $1^{st}$ order oracle access to the log likelihood. Such problems may arise in Bayesian posterior sampling, or in approximate minimization of non-convex functions. Using this knowledge alone, we propose a Monte Carlo method to estimate the score empirically as a particular expectation of a random variable. Using this estimator, we can then run a discrete version of the backward flow SDE to produce samples from the target density. This approach has the benefit of not relying on a pool of initial samples from the target density, and it does not rely on a neural network or other black box model to estimate the score.

stat.ML

Robustness to Incorrect Priors and Controlled Filter Stability in Partially Observed Stochastic Control

We study controlled filter stability and its effects on the robustness properties of optimal control policies designed for systems with incorrect priors applied to a true system. Filter stability refers to the correction of an incorrectly initialized filter for a partially observed stochastic dynamical system (controlled or control-free) with increasing measurements. This problem has been studied extensively in the control-free context, and except for the standard machinery for linear Gaussian systems involving the Kalman Filter, few studies exist for the controlled setup. One of the main differences between control-free and controlled partially observed Markov chains is that the filter is always Markovian under the former, whereas under a controlled model the filter process may not be Markovian since the control policy may depend on past measurements in an arbitrary (measurable) fashion. This complicates the dependency structure and therefore results from the control-free literature do not directly apply to the controlled setup. In this paper, we study the filter stability problem for controlled stochastic dynamical systems, and provide sufficient conditions for when a falsely initialized filter merges with the correctly initialized filter over time. These stability results are applied to robust stochastic control problems: under filter stability, we bound the difference in the expected cost incurred for implementing an incorrectly designed control policy compared to an optimal policy. A conclusion is that filter stability leads to stronger robustness results to incorrect priors (compared with setups without controlled filter stability). Furthermore, if the optimum cost is that same for each prior, the cost of mismatch between the true prior and the assumed prior is zero.

math.OC

Exponential Filter Stability via Dobrushin's Coefficient

Filter stability is a classical problem in the study of partially observed Markov processes (POMP), also known as hidden Markov models (HMM). For a POMP, an incorrectly initialized non-linear filter is said to be (asymptotically) stable if the filter eventually corrects itself as more measurements are collected. Filter stability results in the literature that provide rates of convergence typically rely on very restrictive mixing conditions on the transition kernel and measurement kernel pair, and do not consider their effects independently. In this paper, we introduce an alternative approach using the Dobrushin coefficients associated with both the transition kernel as well as the measurement channel. Such a joint study, which seems to have been unexplored, leads to a concise analysis that can be applied to more general system models under relaxed conditions: in particular, we show that if $(1 - δ(T))(2-δ(Q)) < 1$, where $δ(T)$ and $δ(Q)$ are the Dobrushin coefficients for the transition and the measurement kernels, then the filter is exponentially stable. Our findings are also applicable for controlled models.

math.PR