SearcharxivSearch

arXiv subjects

Simon Coste

Publications and source records attributed to Simon Coste.

13 recordsLinked to original sources

The Wasserstein cost of Importance Sampling

Importance sampling (IS) consists in biasing samples from a distribution $f$ towards another distribution $g$. Concretely, given samples $X_i$ from $f$, the IS measure is $$\hat{g}_n = \frac{1}{Z_n}\sum_{i=1}^n \frac{g(X_i)}{f(X_i)} \delta_{X_i},$$ with $Z_n = \sum_{i=1}^n \frac{g(X_i)}{f(X_i)}$. The random measure $\hat{g}_n$ approximates $g$, and is used in many contexts ranging from Monte Carlo integration to Bayesian inference. We show that, in high dimension ($d \geqslant 3$), the Wasserstein cost $W_p^p(\hat{g}_n, g)$ has order $n^{-p/d}$ in expectation, i.e. $$\beta^{\mathrm{low}}_{p,d}\int gf^{-p/d}\leqslant \liminf_{n \to \infty} n^{p/d} \mathbb{E}[W_p^p(\hat{g}_n, g)] \leqslant \limsup_{n \to \infty} n^{p/d} \mathbb{E}[W_p^p(\hat{g}_n, g)] \leqslant\beta_{p,d} \int g f^{-p/d}$$ where $0<\beta^{\mathrm{low}}_{p,d}\leqslant \beta_{p,d}$ are constants depending only on $p$ and $d$, which are equal for $p=2$ and conjectured to be equal for any $p\geqslant 1$. Our results are valid for all $p\geqslant 1$ and $d\geqslant 3$. In the case where $\beta^{\mathrm{low}}_{p,d} = \beta_{p,d}$, we show that the asymptotically optimal sampling distribution $f^*$ for importance sampling is not equal to $g$ but to a tempered version of $g$, namely $f^* \propto g^{d/(p+d)}$, which is reminiscent of Zador's theorem in the domain of measure quantization.

math.PR

The delocalization of eigenvectors of real elliptic matrices

We investigate delocalization phenomena for eigenvectors of real random matrices that are invariant by orthogonal transformations. A specific phenomenon with these ensembles is that an eigenvector is typically more localized when its eigenvalue is closer to the real axis while for unitarily invariant ensembles, all eigenvectors are delocalized at the same level. More precisely, we measure the delocalization level of a vector $x\in \mathbb{C}^N$ using the Inverse Participation Ratio $\mathrm{IPR}(x) = N|x|_4^4 / |x|_2^4 \geqslant 1$. A higher IPR means a more localized vector. Using the exact distribution of the Schur decomposition of some paradigmatic rotation-invariant matrix models, we prove that conditionally on having an eigenvalue $\lambda$ with $|\mathfrak{Im}(\lambda)| = y / \sqrt{N}$, the IPR of the associated eigenvector converges in distribution towards a random variable $\ell_y$ with an explicit density depending only on $y$. We then prove that $\ell_y \to 3$ when $y \to 0$ and $\ell_y \to 2$ when $y\to +\infty$, coherently with the observed phenomenon. This result is explicitly proved for higher-order IPRs and for the real Elliptic Ginibre ensemble at every non-symmetry parameter $\tau \in [0,1[$, including the classical real Ginibre ensemble ($\tau=0$).

math.PR

Random eigenvalues of graphenes and the triangulation of plane

We analyse the numbers of closed paths of length $k\in\mathbb{N}$ on two important regular lattices: the hexagonal lattice (also called $\textit{graphene}$ in chemistry) and its dual triangular lattice. These numbers form a moment sequence of specific random variables connected to the distance of a position of a planar random flight (in three steps) from the origin. Here, we refer to such a random variable as a $\textit{random eigenvalue}$ of the underlying lattice. Explicit formulas for the probability density and characteristic functions of these random eigenvalues are given for both the hexagonal and the triangular lattice. Furthermore, it is proven that both probability distributions can be approximated by a functional of the random variable uniformly distributed on increasing intervals $[0,b]$ as $b\to\infty$. This yields a straightforward method to simulate these random eigenvalues without generating graphene and triangular lattice graphs. To demonstrate this approximation, we first prove a key integral identity for a specific series containing the third powers of the modified Bessel functions $I_n$ of $n$th order, $n\in\mathbb{Z}$. Such series play a crucial role in various contexts, in particular, in analysis, combinatorics, and theoretical physics.

math.SP

Efficient Training of Energy-Based Models Using Jarzynski Equality

Energy-based models (EBMs) are generative models inspired by statistical physics with a wide range of applications in unsupervised learning. Their performance is best measured by the cross-entropy (CE) of the model distribution relative to the data distribution. Using the CE as the objective for training is however challenging because the computation of its gradient with respect to the model parameters requires sampling the model distribution. Here we show how results for nonequilibrium thermodynamics based on Jarzynski equality together with tools from sequential Monte-Carlo sampling can be used to perform this computation efficiently and avoid the uncontrolled approximations made using the standard contrastive divergence algorithm. Specifically, we introduce a modification of the unadjusted Langevin algorithm (ULA) in which each walker acquires a weight that enables the estimation of the gradient of the cross-entropy at any step during GD, thereby bypassing sampling biases induced by slow mixing of ULA. We illustrate these results with numerical experiments on Gaussian mixture distributions as well as the MNIST dataset. We show that the proposed approach outperforms methods based on the contrastive divergence algorithm in all the considered situations.

cs.LG

Wavelet Score-Based Generative Modeling

Score-based generative models (SGMs) synthesize new data samples from Gaussian white noise by running a time-reversed Stochastic Differential Equation (SDE) whose drift coefficient depends on some probabilistic score. The discretization of such SDEs typically requires a large number of time steps and hence a high computational cost. This is because of ill-conditioning properties of the score that we analyze mathematically. We show that SGMs can be considerably accelerated, by factorizing the data distribution into a product of conditional probabilities of wavelet coefficients across scales. The resulting Wavelet Score-based Generative Model (WSGM) synthesizes wavelet coefficients with the same number of time steps at all scales, and its time complexity therefore grows linearly with the image size. This is proved mathematically over Gaussian distributions, and shown numerically over physical processes at phase transition and natural image datasets.

cs.LG

The characteristic polynomial of sums of random permutations and regular digraphs

Let $A_n$ be the sum of $d$ permutation matrices of size $n\times n$, each drawn uniformly at random and independently. We prove that the normalized characteristic polynomial $\frac{1}{\sqrt{d}}\det(I_n - z A_n/\sqrt{d})$ converges when $n\to \infty$ towards a random analytic function on the unit disk. As an application, we obtain an elementary proof of the spectral gap of random regular digraphs. Our results are valid both in the regime where $d$ is fixed and for $d$ slowly growing with $n$.

math.PR

Sparse matrices: convergence of the characteristic polynomial seen from infinity

We prove that the reverse characteristic polynomial $\det(I_n - zA_n)$ of a random $n \times n$ matrix $A_n$ with iid $\mathrm{Bernoulli}(d/n)$ entries converges in distribution towards the random infinite product $\prod_{\ell = 1}^\infty(1-z^\ell)^{Y_\ell}$ where $Y_\ell$ are independent $\mathrm{Poisson}(d^\ell/\ell)$ random variables. We show that this random function is a Poisson analog of more classical Gaussian objects such as the Gaussian holomorphic chaos. As a byproduct, we obtain new simple proofs of previous results on the asymptotic behaviour of extremal eigenvalues of sparse Erd\H{o}s-R\'enyi digraphs: for every $d>1$, the greatest eigenvalue of $A_n$ is close to $d$ and the second greatest is smaller than $\sqrt{d}$, a Ramanujan-like property for irregular digraphs. For $d<1$, the only non-zero eigenvalues of $A_n$ converge to a Poisson multipoint process on the unit circle. Our results also extend to the semi-sparse regime where $d$ is allowed to grow to $\infty$ with $n$, slower than $n^{o(1)}$. We show that the reverse characteristic polynomial converges towards a more classical object written in terms of the exponential of a log-correlated real Gaussian field, as in the dense case studied in a recent paper \cite{bordenave2020convergence}. In the semi-sparse regime, the empirical spectral distribution of $A_n/\sqrt{d_n}$ converges to the circle distribution; as a consequence of our results, the second eigenvalue sticks to the edge of the circle.

math.PR

A simpler spectral approach for clustering in directed networks

We study the task of clustering in directed networks. We show that using the eigenvalue/eigenvector decomposition of the adjacency matrix is simpler than all common methods which are based on a combination of data regularization and SVD truncation, and works well down to the very sparse regime where the edge density has constant order. Our analysis is based on a Master Theorem describing sharp asymptotics for isolated eigenvalues/eigenvectors of sparse, non-symmetric matrices with independent entries. We also describe the limiting distribution of the entries of these eigenvectors; in the task of digraph clustering with spectral embeddings, we provide numerical evidence for the superiority of Gaussian Mixture clustering over the widely used k-means algorithm.

cs.LG

Detection thresholds in very sparse matrix completion

Let $A$ be a rectangular matrix of size $m\times n$ and $A_1$ be the random matrix where each entry of $A$ is multiplied by an independent $\{0,1\}$-Bernoulli random variable with parameter $1/2$. This paper is about when, how and why the non-Hermitian eigen-spectra of the randomly induced asymmetric matrices $A_1 (A - A_1)^*$ and $(A-A_1)^*A_1$ captures more of the relevant information about the principal component structure of $A$ than via its SVD or the eigen-spectra of $A A^*$ and $A^* A$, respectively. Hint: the asymmetry inducing randomness breaks the echo-chamber effect that cripples the SVD. We illustrate the application of this striking phenomenon on the low-rank matrix completion problem for the setting where each entry is observed with probability $d/n$, including the very sparse regime where $d$ is of order $1$, where matrix completion via the SVD of $A$ fails or produces unreliable recovery. We determine an asymptotically exact, matrix-dependent, non-universal detection threshold above which reliable, statistically optimal matrix recovery using a new, universal data-driven matrix-completion algorithm is possible. Averaging the left and right eigenvectors provably improves the recovered matrix but not the detection threshold. We define another variant of this asymmetric procedure that bypasses the randomization step and has a detection threshold that is smaller by a constant factor but with a computational cost that is larger by a polynomial factor of the number of observed entries. Both detection thresholds shatter the seeming barrier due to the well-known information theoretical limit $d \asymp \log n$ for matrix completion found in the literature.

math.PR

Eigenvalues of the non-backtracking operator detached from the bulk

We describe the non-backtracking spectrum of a stochastic block model with connection probabilities $p_{\mathrm{in}}, p_{\mathrm{out}} = \omega(\log n)/n$. In this regime we answer a question posed in Dall'Amico and al. (2019) regarding the existence of a real eigenvalue `inside' the bulk, close to the location $\frac{p_{\mathrm{in}}+ p_{\mathrm{out}}}{p_{\mathrm{in}}- p_{\mathrm{out}}}$. We also introduce a variant of the Bauer-Fike theorem well suited for perturbations of quadratic eigenvalue problems, and which could be of independent interest.

math.PR

Emergence of extended states at zero in the spectrum of sparse random graphs

We confirm the long-standing prediction that $c=e\approx 2.718$ is the threshold for the emergence of a non-vanishing absolutely continuous part (extended states) at zero in the limiting spectrum of the Erd\H{o}s-Renyi random graph with average degree $c$. This is achieved by a detailed second-order analysis of the resolvent $(A-z)^{-1}$ near the singular point $z=0$, where $A$ is the adjacency operator of the Poisson-Galton-Watson tree with mean offspring $c$. More generally, our method applies to arbitrary unimodular Galton-Watson trees, yielding explicit criteria for the presence or absence of extended states at zero in the limiting spectral measure of a variety of random graph models, in terms of the underlying degree distribution.

math.PR

The Spectral Gap of Sparse Random Digraphs

The second largest eigenvalue of a transition matrix $P$ has connections with many properties of the underlying Markov chain, and especially its convergence rate towards the stationary distribution. In this paper, we give an asymptotic upper bound for the second eigenvalue when $P$ is the transition matrix of the simple random walk over a random directed graph with given degree sequence. This is the first result concerning the asymptotic behavior of the spectral gap for sparse non-reversible Markov chains with an unknown stationary distribution. An immediate consequence of our result is a proof of the Alon conjecture for directed regular graphs. Our result is based on a variation of the trace method introduced by Bordenave (2015).

math.PR