SearcharxivSearch

arXiv subjects

Promit Ghosal

Publications and source records attributed to Promit Ghosal.

At least 37 records · Page 2Linked to original sources

Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent

We provide finite-particle convergence rates for the Stein Variational Gradient Descent (SVGD) algorithm in the Kernelized Stein Discrepancy ($\mathsf{KSD}$) and Wasserstein-2 metrics. Our key insight is that the time derivative of the relative entropy between the joint density of $N$ particle locations and the $N$-fold product target measure, starting from a regular initial distribution, splits into a dominant `negative part' proportional to $N$ times the expected $\mathsf{KSD}^2$ and a smaller `positive part'. This observation leads to $\mathsf{KSD}$ rates of order $1/\sqrt{N}$, in both continuous and discrete time, providing a near optimal (in the sense of matching the corresponding i.i.d. rates) double exponential improvement over the recent result by Shi and Mackey (2024). Under mild assumptions on the kernel and potential, these bounds also grow polynomially in the dimension $d$. By adding a bilinear component to the kernel, the above approach is used to further obtain Wasserstein-2 convergence in continuous time. For the case of `bilinear + Mat\'ern' kernels, we derive Wasserstein-2 rates that exhibit a curse-of-dimensionality similar to the i.i.d. setting. We also obtain marginal convergence and long-time propagation of chaos results for the time-averaged particle laws.

math.ST

Proof of Zamolodchikov conjecture for semi-classical conformal blocks on the torus

In 1986, Zamolodchikov conjectured an exponential structure for the semi-classical limit of conformal blocks on a sphere. This paper provides a rigorous proof of the analog of Zamolodchikov conjecture for Liouville conformal blocks on a one-punctured torus, using their probabilistic construction and show the existence of a positive radius of convergence of the semi-classical limit. As a consequence, we obtain a closed form expression for the solution of the Lam\'e equation, and show a relation between its accessory parameter and the classical action of the non-autonomous elliptic Calogero-Moser model evaluated at specific values of the solution.

math-ph

High-dimensional scaling limits and fluctuations of online least-squares SGD with smooth covariance

We derive high-dimensional scaling limits and fluctuations for the online least-squares Stochastic Gradient Descent (SGD) algorithm by taking the properties of the data generating model explicitly into consideration. Our approach treats the SGD iterates as an interacting particle system, where the expected interaction is characterized by the covariance structure of the input. Assuming smoothness conditions on moments of order up to eight orders, and without explicitly assuming Gaussianity, we establish the high-dimensional scaling limits and fluctuations in the form of infinite-dimensional Ordinary Differential Equations (ODEs) or Stochastic Differential Equations (SDEs). Our results reveal a precise three-step phase transition of the iterates; it goes from being ballistic, to diffusive, and finally to purely random behavior, as the noise variance goes from low, to moderate and finally to very-high noise setting. In the low-noise setting, we further characterize the precise fluctuations of the (scaled) iterates as infinite-dimensional SDEs. We also show the existence and uniqueness of solutions to the derived limiting ODEs and SDEs. Our results have several applications, including characterization of the limiting mean-square estimation or prediction errors and their fluctuations, which can be obtained by analytically or numerically solving the limiting equations.

math.PR

Probabilistic conformal blocks for Liouville CFT on the torus

Virasoro conformal blocks are a family of important functions defined as power series via the Virasoro algebra. They are a fundamental input to the conformal bootstrap program for 2D conformal field theory (CFT) and are closely related to four dimensional supersymmetric gauge theory through the Alday-Gaiotto-Tachikawa correspondence. The present work provides a probabilistic construction of the 1-point toric Virasoro conformal block for central change greater than 25. More precisely, we construct an analytic function using a probabilistic tool called Gaussian multiplicative chaos (GMC) and prove that its power series expansion coincides with the 1-point toric Virasoro conformal block. The range $(25,\infty)$ of central charges corresponds to Liouville CFT, an important CFT originating from 2D quantum gravity and bosonic string theory. Our work reveals a new integrable structure underlying GMC and opens the door to the study of non-perturbative properties of Virasoro conformal blocks such as their analytic continuation and modular symmetry. Our proof combines an analysis of GMC with tools from CFT such as Belavin-Polyakov-Zamolodchikov differential equations, operator product expansions, and Dotsenko-Fateev type integrals.

math.PR

Towards Understanding the Dynamics of Gaussian-Stein Variational Gradient Descent

Stein Variational Gradient Descent (SVGD) is a nonparametric particle-based deterministic sampling algorithm. Despite its wide usage, understanding the theoretical properties of SVGD has remained a challenging problem. For sampling from a Gaussian target, the SVGD dynamics with a bilinear kernel will remain Gaussian as long as the initializer is Gaussian. Inspired by this fact, we undertake a detailed theoretical study of the Gaussian-SVGD, i.e., SVGD projected to the family of Gaussian distributions via the bilinear kernel, or equivalently Gaussian variational inference (GVI) with SVGD. We present a complete picture by considering both the mean-field PDE and discrete particle systems. When the target is strongly log-concave, the mean-field Gaussian-SVGD dynamics is proven to converge linearly to the Gaussian distribution closest to the target in KL divergence. In the finite-particle setting, there is both uniform in time convergence to the mean-field limit and linear convergence in time to the equilibrium if the target is Gaussian. In the general case, we propose a density-based and a particle-based implementation of the Gaussian-SVGD, and show that several recent algorithms for GVI, proposed from different perspectives, emerge as special cases of our unified framework. Interestingly, one of the new particle-based instance from this framework empirically outperforms existing approaches. Our results make concrete contributions towards obtaining a deeper understanding of both SVGD and GVI.

math.ST

Randomly Initialized One-Layer Neural Networks Make Data Linearly Separable

Recently, neural networks have demonstrated remarkable capabilities in mapping two arbitrary sets to two linearly separable sets. The prospect of achieving this with randomly initialized neural networks is particularly appealing due to the computational efficiency compared to fully trained networks. This paper contributes by establishing that, given sufficient width, a randomly initialized one-layer neural network can, with high probability, transform two sets into two linearly separable sets without any training. Moreover, we furnish precise bounds on the necessary width of the neural network for this phenomenon to occur. Our initial bound exhibits exponential dependence on the input dimension while maintaining polynomial dependence on all other parameters. In contrast, our second bound is independent of input dimension, effectively surmounting the curse of dimensionality. The main tools used in our proof heavily relies on a fusion of geometric principles and concentration of random matrices.

cs.LG

Linear programming using diagonal linear networks

Linear programming has played a crucial role in shaping decision-making, resource allocation, and cost reduction in various domains. In this paper, we investigate the application of overparametrized neural networks and their implicit bias in solving linear programming problems. Specifically, our findings reveal that training diagonal linear networks with gradient descent, while optimizing the squared $L_2$-norm of the slack variable, leads to solutions for entropically regularized linear programming problems. Remarkably, the strength of this regularization depends on the initialization used in the gradient descent process. We analyze the convergence of both discrete-time and continuous-time dynamics and demonstrate that both exhibit a linear rate of convergence, requiring only mild assumptions on the constraint matrix. For the first time, we introduce a comprehensive framework for solving linear programming problems using diagonal neural networks. We underscore the significance of our discoveries by applying them to address challenges in basis pursuit and optimal transport problems.

math.OC

From Stability to Chaos: Analyzing Gradient Descent Dynamics in Quadratic Regression

We conduct a comprehensive investigation into the dynamics of gradient descent using large-order constant step-sizes in the context of quadratic regression models. Within this framework, we reveal that the dynamics can be encapsulated by a specific cubic map, naturally parameterized by the step-size. Through a fine-grained bifurcation analysis concerning the step-size parameter, we delineate five distinct training phases: (1) monotonic, (2) catapult, (3) periodic, (4) chaotic, and (5) divergent, precisely demarcating the boundaries of each phase. As illustrations, we provide examples involving phase retrieval and two-layer neural networks employing quadratic activation functions and constant outer-layers, utilizing orthogonal training data. Our simulations indicate that these five phases also manifest with generic non-orthogonal data. We also empirically investigate the generalization performance when training in the various non-monotonic (and non-divergent) phases. In particular, we observe that performing an ergodic trajectory averaging stabilizes the test error in non-monotonic (and non-divergent) phases.

cs.LG

Lyapunov exponents of the SHE for general initial data

We consider the $(1+1)$-dimensional stochastic heat equation (SHE) with multiplicative white noise and the Cole-Hopf solution of the Kardar-Parisi-Zhang (KPZ) equation. We show an exact way of computing the Lyapunov exponents of the SHE for a large class of initial data which includes any bounded deterministic positive initial data and the stationary initial data. As a consequence, we derive exact formulas for the upper tail large deviation rate functions of the KPZ equation for general initial data.

math.PR

Moment Intermittency in the PAM with Asymptotically Singular Noise

Let $ξ$ be a singular Gaussian noise on $\mathbb R^d$ that is either white, fractional, or with the Riesz covariance kernel; in particular, there exists a scaling parameter $ω>0$ such that $c^{ω/2}ξ(c\cdot)$ is equal in distribution to $ξ$ for all $c>0$. Let $(ξ_\varepsilon)_{\varepsilon>0}$ be a sequence of smooth mollifications such that $ξ_\varepsilon\toξ$ as $\varepsilon\to0$. We study the asymptotics of the moments of the parabolic Anderson model (PAM) with noise $ξ_\varepsilon$ as $\varepsilon\to0$, both for large (i.e., $t\to\infty$) and fixed times $t$. This approach makes it possible to study the moments of the PAM with regular and singular noises in a unified fashion, as well as interpolate between the two settings. As corollaries of our main results, we obtain the following: $\textbf{(1)}$ When $ξ$ is subcritical (i.e., $0<ω<2$), our results extend the known large-time moment and tail asymptotics for the Stratonovich PAM with noise $ξ$. Our method of proof clarifies the role of the maximizers of the variational problems (known as Hartree ground states) that appear in these moment asymptotics in describing the geometry of intermittency. We take this opportunity to prove the existence and study the properties of the Hartree ground state with a fractional kernel, which we believe is of independent interest. $\textbf{(2)}$ When $ξ$ is critical or supercritical (i.e., $ω=2$ or $ω>2$), our results provide a new interpretation of the moment blowup phenomenon observed in the Stratonovich PAM with noise $ξ$. That is, we uncover that the latter is related to an intermittency effect that occurs in the PAM with noise $ξ_\varepsilon$ as $\varepsilon\to0$ for $\textit{fixed finite times}$ $t>0$.

math.PR

Fractal geometry of the PAM in 2D and 3D with white noise potential

We study the parabolic Anderson model (PAM) \begin{equation} {\partial \over \partial t}u(t,x) =\frac{1}{2}Δu(t,x) + u(t,x)ξ(x), \quad t>0, x\in \mathbb{R}^d, \quad \text{and} \quad u(0,x) \equiv 1, \quad \forall x\in \mathbb{R}^d, \end{equation} where $ξ$ is spatial white noise on $\mathbb{R}^d$ with $d \in\{2,3\}$. We show that the peaks of the PAM are macroscopically multifractal. More precisely, we prove that the spatial peaks of the PAM have infinitely many distinct values and we compute the macroscopic Hausdorff dimension (introduced by Barlow and Taylor) of those peaks. As a byproduct, we obtain the exact spatial asymptotics of the solution of the PAM. We also study the spatio-temporal peaks of the PAM and show their macroscopic multifractality. Some of the major tools used in our proof techniques include paracontrolled calculus and tail probabilities of the largest point in the spectrum of the Anderson Hamiltonian.

math.PR

On Spatial Conditioning of the Spectrum of Discrete Random Schrödinger Operators

Consider a random Schrödinger-type operator of the form $H:=-H_X+V+ξ$ acting on a general graph $\mathscr G=(\mathscr V,\mathscr E)$, where $H_X$ is the generator of a Markov process $X$ on $\mathscr G$, $V$ is a deterministic potential with sufficient growth (so that $H$ has a purely discrete spectrum), and $ξ$ is a random noise with at-most-exponential tails. We prove that $H$'s eigenvalue point process is number rigid in the sense of Ghosh and Peres (Duke Math. J. 166 (2017), no. 10, 1789--1858); that is, the number of eigenvalues in any bounded domain $B\subset\mathbb C$ is determined by the configuration of eigenvalues outside of $B$. Our general setting allows to treat cases where $X$ could be non-symmetric (hence $H$ is non-self-adjoint) and $ξ$ has long-range dependence. Our strategy of proof consists of controlling the variance of the trace of the semigroup $\mathrm e^{-t H}$ using the Feynman-Kac formula.

math-ph

Rigidity of the Stochastic Airy Operator

We prove that the spectrum of the stochastic Airy operator is rigid in the sense of Ghosh and Peres (Duke Math. J., 166(10):1789--1858, 2017) for Dirichlet and Robin boundary conditions. This proves the rigidity of the Airy-$β$ point process and the soft-edge limit of rank-$1$ perturbations of Gaussian $β$-Ensembles for any $β>0$, and solves an open problem mentioned in a previous work of Bufetov, Nikitin, and Qiu (Mosc. Math. J., 19(2):217--274, 2019). Our proof uses a combination of the semigroup theory of the stochastic Airy operator and the techniques for studying insertion and deletion tolerance of point processes developed by Holroyd and Soo (Electron. J. Probab., 18:no. 74, 24, 2013).

math.PR

Long and short time laws of iterated logarithms for the KPZ fixed point

We consider the KPZ fixed point starting from a general class of initial data. In this article, we study the growth of the large peaks of the KPZ fixed point at a spatial point $0$ when time $t$ goes to $\infty$ and when $t$ approaches $1$. We prove that for a very broad class of initial data, as $t\to \infty$, the limsup of the KPZ fixed point height function when scaled by $t^{1/3}(\log\log t)^{2/3}$ almost surely equals a constant. The value of the constant is $(3/4)^{2/3}$ or $(3/2)^{2/3}$ depending on the initial data being non-random or Brownian respectively. Furthermore, we show that the increments of the KPZ fixed point near $t=1$ admits a short time law of iterated logarithm. More precisely, as the time increments $Δt :=t-1$ goes down to $0$, for a large class of initial data including the Brownian data initial data, we show that limsup of the height increments the KPZ fixed point near time $1$ when scaled by $(Δt)^{1/3}(\log\log (Δt)^{-1})^{2/3}$ almost surely equals $(3/2)^{2/3}$.

math.PR

Stability of Entropic Optimal Transport and Schrödinger Bridges

We establish the stability of solutions to the entropically regularized optimal transport problem with respect to the marginals and the cost function. The result is based on the geometric notion of cyclical invariance and inspired by the use of $c$-cyclical monotonicity in classical optimal transport. As a consequence of stability, we obtain the wellposedness of the solution in this geometric sense, even when all transports have infinite cost. More generally, our results apply to a class of static Schrödinger bridge problems including entropic optimal transport.

math.OC

The ASEP speed process

For ASEP with step initial data and a second class particle started at the origin we prove that as time goes to infinity the second class particle almost surely achieves a velocity that is uniformly distributed on $[-1,1]$. This positively resolves Conjecture 1.9 and 1.10 of [Amir, Angel and Valko, "The TASEP speed process", Annals of Probability 39, 1205--1242, 2011] and allows us to construct the ASEP speed process.

math.PR

Universality for multiplicative statistics of Hermitian random matrices and the integro-differential Painlevé II equation

We study multiplicative statistics for the eigenvalues of unitarily-invariant Hermitian random matrix models. We consider one-cut regular polynomial potentials and a large class of multiplicative statistics. We show that in the large matrix limit several associated quantities converge to limits which are universal in both the potential and the family of multiplicative statistics considered. In turn, such universal limits are described by the integro-differential Painlevé II equation, and in particular they connect the random matrix models considered with the narrow wedge solution to the KPZ equation at any finite time.

math-ph

Entropic Optimal Transport: Geometry and Large Deviations

We study the convergence of entropically regularized optimal transport to optimal transport. The main result is concerned with the convergence of the associated optimizers and takes the form of a large deviations principle quantifying the local exponential convergence rate as the regularization parameter vanishes. The exact rate function is determined in a general setting and linked to the Kantorovich potential of optimal transport. Our arguments are based on the geometry of the optimizers and inspired by the use of $c$-cyclical monotonicity in classical transport theory. The results can also be phrased in terms of Schrödinger bridges.

math.OC