SearcharxivSearch

arXiv subjects

Jason Beh

Publications and source records attributed to Jason Beh.

3 recordsLinked to original sources

Phase transition for conditional covariance matrices estimated by importance sampling, and implications for cross-entropy schemes in high dimension

Motivated by the estimation of covariance matrices by importance sampling arising in the cross-entropy (CE) algorithm, we study a random matrix model $\hat \Sigma = {\bf X} L {\bf X}^\top$ with two distinct features: $\bf X$ and $L$ are dependent, and $L$ is heavy-tailed. In the high-dimensional regime $d \to \infty$, we prove under suitable assumptions that a phase transition occurs in the polynomial regime $n = d^\kappa$, with $n$ the sample size. Namely, we prove that $\lVert \hat \Sigma - E \hat \Sigma \rVert \Rightarrow 0$ if and only if $\kappa > \kappa_*$ for some threshold $\kappa_*$ determined by the behavior of the maximum likelihood ratios. Moreover, we identify general situations where $\kappa_* = 1/\lambda_1$, with $\lambda_1$ the smallest eigenvalue of the covariance matrix of the auxiliary distribution used to estimate $\hat \Sigma$ by importance sampling. This suggests that importance sampling will work better with covariance matrices having a large smallest eigenvalue. We carry this insight into recent CE schemes proposed to estimate the probability of high-dimensional rare events. Through numerical simulations, we demonstrate that better CE schemes are also the ones with larger smallest eigenvalue, even though these algorithms were not designed to smooth the spectrum. This new spectral interpretation raises stimulating questions and opens research directions for the design of efficient high-dimensional algorithms.

math.ST

Affine invariant interacting Langevin dynamics in Markov chain importance sampling for rare event estimation

This work considers the framework of Markov chain importance sampling~(MCIS), in which one employs a Markov chain Monte Carlo~(MCMC) scheme to sample particles approaching the optimal distribution for importance sampling, prior to estimating the quantity of interest through importance sampling. In rare event estimation, the optimal distribution admits a non-differentiable log-density, thus gradient-based MCMC can only target a smooth approximation of the optimal density. We propose a new gradient-based MCIS scheme for rare event estimation, called affine invariant interacting Langevin dynamics for importance sampling~(ALDI-IS), in which the affine invariant interacting Langevin dynamics~(ALDI) is used to sample particles according to the smoothed zero-variance density. We establish a non-asymptotic error bound when importance sampling is used in conjunction with samples independently and identically distributed according to the smoothed optiaml density to estimate a rare event probability, and an error bound on the sampling bias when a simplified version of ALDI, the unadjusted Langevin algorithm, is used to sample from the smoothed optimal density. We show that the smoothing parameter of the optimal density has a strong influence and exhibits a trade-off between a low importance sampling error and the ease of sampling using ALDI. We perform a numerical study of ALDI-IS and illustrate this trade-off phenomenon on standard rare event estimation test cases.

math.ST

Insight from the Kullback--Leibler divergence into adaptive importance sampling schemes for rare event analysis in high dimension

We study two adaptive importance sampling schemes for estimating the probability of a rare event in the high-dimensional regime $d \to \infty$ with $d$ the dimension. The first scheme is the prominent cross-entropy (CE) method, and the second scheme, motivated by recent results, uses as auxiliary distribution a projection of the optimal auxiliary distribution on a lower dimensional subspace. In these schemes, two samples are used: the first one to learn the auxiliary distribution and the second one, drawn according to the learned distribution, to perform the final probability estimation. Contrary to the common belief that the sample size needs to grow exponentially in the dimension to make the estimator consistent and avoid the weight degeneracy phenomenon, we find that a polynomial sample size in the first learning step is enough. We prove this result assuming that the sought probability is bounded away from 0. For CE, insight is provided on the polynomial growth rate which remains implicit. In contrast, we study the second scheme in a simple computational framework assuming that samples from the conditional distribution are available. This makes it possible to show that the sample size only needs to grow like $rd$ with $r$ the effective dimension of the projection, which highlights the potential benefits of these projection methods.

math.ST