SearcharxivSearch

arXiv subjects

Dejan Slepčev

Publications and source records attributed to Dejan Slepčev.

At least 19 recordsLinked to original sources

Wasserstein gradient flows of Maximum Mean Discrepancy with energy kernels

We study the Wasserstein gradient flow of the squared Maximum Mean Discrepancy (MMD) generated by the nonsmooth energy kernels $K(z)=-|z|^q$, $0 0$, we prove global well-posedness on $\mathbb{R}^d$ for probability densities in subcritical $L^p$ spaces, with targets in the same integrability class and with finite moments. We also include the one-dimensional Coulomb endpoint $d=q=1$. For the associated $N$-particle system, we prove global noncollision and fixed-$N$ convergence to the collision-free critical set, a particle-to-continuum criticality principle, and a modulated-energy mean-field estimate that yields convergence of the particle dynamics to the continuum flow as $N\to\infty$ on every finite time interval. We also construct collision-free saddle equilibria, showing that deterministic particle trajectories need not approach global empirical minimizers. For $1\le q<2$, every continuum solution in our class has a narrowly relatively compact orbit, every $ω$-limit point is Lagrangian critical, and the orbit approaches the Lagrangian critical set. For $0<q<1$, the same conclusions hold under uniform-in-time moment and subcritical $L^p$ bounds. We prove that an absolutely continuous Lagrangian critical point equals the target when the source and target have finite moments of order $q$, except when $0<q<1$ and $d\in\{1,3\}$. Under the preceding uniform bounds, rigidity gives convergence of the continuum flow to the target throughout the rigid part of the well-posedness range. Finally, we show that no initial-data-independent multiplicative MMD decay modulus exists on $\mathbb{R}^d$, and that global Polyak--Łojasiewicz inequalities fail in several whole-space and periodic Riesz/Coulomb regimes.

math.AP

Radon--Wasserstein Gradient Flows for Interacting-Particle Sampling in High Dimensions

Gradient flows of the Kullback--Leibler (KL) divergence, such as the Fokker--Planck equation and Stein Variational Gradient Descent, evolve a distribution toward a target density known only up to a normalizing constant. We introduce new gradient flows of the KL divergence with a remarkable combination of properties: they admit accurate interacting-particle approximations in high dimensions, and the per-step cost scales linearly in both the number of particles and the dimension. These gradient flows are based on new transportation-based Riemannian geometries on the space of probability measures: the Radon--Wasserstein geometry and the related Regularized Radon--Wasserstein (RRW) geometry. We define these geometries using the Radon transform so that the gradient-flow velocities depend only on one-dimensional projections. This yields interacting-particle-based algorithms whose per-step cost follows from efficient Fast Fourier Transform-based evaluation of the required 1D convolutions. We additionally provide numerical experiments that study the performance of the proposed algorithms and compare convergence behavior and quantization. Finally, we prove some theoretical results including well-posedness of the flows and long-time convergence guarantees for the RRW flow.

stat.ML

Time-complexity of sampling from a multimodal distribution using sequential Monte Carlo

We study a sequential Monte Carlo algorithm to sample from the Gibbs measure with a non-convex energy function at a low temperature. We use the practical and popular geometric annealing schedule, and use a Langevin diffusion at each temperature level. The Langevin diffusion only needs to run for a time that is long enough to ensure local mixing within energy valleys, which is much shorter than the time required for global mixing. Our main result shows convergence of Monte Carlo estimators with time complexity that, approximately, scales like the fourth power of the inverse temperature, and the square of the inverse allowed error. We also study this algorithm in an illustrative model scenario where more explicit estimates can be given.

math.ST

Convergence of graph Dirichlet energies and graph Laplacians on intersecting manifolds of varying dimensions

We study $Γ$-convergence of graph Dirichlet energies and spectral convergence of graph Laplacians on unions of intersecting manifolds of potentially different dimensions. Our investigation is motivated by problems of machine learning, as real-world data often consist of parts or classes with different intrinsic dimensions. An important challenge is to understand which machine learning methods adapt to such varied dimensionalities. We investigate the standard unnormalized and the normalized graph Dirichlet energies. We show that the unnormalized energy and its associated graph Laplacian asymptotically only sees the variations within the manifold of the highest dimension. On the other hand, we prove that the normalized Dirichlet energy converges to a (tensorized) Dirichlet energy on the union of manifolds that adapts to all dimensions simultaneously. We also establish the related spectral convergence and present a few numerical experiments to illustrate our findings.

math.AP

Geometry and analytic properties of the sliced Wasserstein space

The sliced Wasserstein metric compares probability measures on $\mathbb{R}^d$ by taking averages of the Wasserstein distances between projections of the measures to lines. The distance has found a range of applications in statistics and machine learning, as it is easier to approximate and compute in high dimensions than the Wasserstein distance. While the geometry of the Wasserstein metric is quite well understood, and has led to important advances, very little is known about the geometry and metric properties of the sliced Wasserstein (SW) metric. Here we show that when the measures considered are "nice" (e.g. bounded above and below by positive multiples of the Lebesgue measure) then the SW metric is comparable to the (homogeneous) negative Sobolev norm $\dot H^{-(d+1)/2}$. On the other hand when the measures considered are close in the infinity transportation metric to a discrete measure, then the SW metric between them is close to a multiple of the Wasserstein metric. We characterize the tangent space of the SW space, show that the speed of curves in the space can be described by a quadratic form, but that the SW space is not a length space. We establish a number of properties of the metric given by the minimal length of curves between measures -- the SW length. Finally we highlight the consequences of these properties to the gradient flows in the SW metric.

math.AP

Birth-death dynamics for sampling: Global convergence, approximations and their asymptotics

Motivated by the challenge of sampling Gibbs measures with nonconvex potentials, we study a continuum birth-death dynamics. We improve results in previous works [51,57] and provide weaker hypotheses under which the probability density of the birth-death governed by Kullback-Leibler divergence or by $χ^2$ divergence converge exponentially fast to the Gibbs equilibrium measure, with a universal rate that is independent of the potential barrier. To build a practical numerical sampler based on the pure birth-death dynamics, we consider an interacting particle system, which is inspired by the gradient flow structure and the classical Fokker-Planck equation and relies on kernel-based approximations of the measure. Using the technique of $Γ$-convergence of gradient flows, we show that on the torus, smooth and bounded positive solutions of the kernelized dynamics converge on finite time intervals, to the pure birth-death dynamics as the kernel bandwidth shrinks to zero. Moreover we provide quantitative estimates on the bias of minimizers of the energy corresponding to the kernelized dynamics. Finally we prove the long-time asymptotic results on the convergence of the asymptotic states of the kernelized dynamics towards the Gibbs measure.

math.AP

HV Geometry for Signal Comparison

In order to compare and interpolate signals, we investigate a Riemannian geometry on the space of signals. The metric allows discontinuous signals and measures both horizontal (thus providing many benefits of the Wasserstein metric) and vertical deformations. Moreover, it allows for signed signals, which overcomes the main deficiency of optimal transportation-based metrics in signal processing. We characterize the metric properties of the space of signals and establish the regularity and stability of geodesics. Furthermore, we introduce an efficient numerical scheme to compute the geodesics and present several experiments which highlight the nature of the metric.

math.MG

Nonlocal Wasserstein Distance: Metric and Asymptotic Properties

The seminal result of Benamou and Brenier provides a characterization of the Wasserstein distance as the path of the minimal action in the space of probability measures, where paths are solutions of the continuity equation and the action is the kinetic energy. Here we consider a fundamental modification of the framework where the paths are solutions of nonlocal (jump) continuity equations and the action is a nonlocal kinetic energy. The resulting nonlocal Wasserstein distances are relevant to fractional diffusions and Wasserstein distances on graphs. We characterize the basic properties of the distance and obtain sharp conditions on the (jump) kernel specifying the nonlocal transport that determine whether the topology metrized is the weak or the strong topology. A key result of the paper are the quantitative comparisons between the nonlocal and local Wasserstein distance.

math.AP

Boundary Estimation from Point Clouds: Algorithms, Guarantees and Applications

We investigate identifying the boundary of a domain from sample points in the domain. We introduce new estimators for the normal vector to the boundary, distance of a point to the boundary, and a test for whether a point lies within a boundary strip. The estimators can be efficiently computed and are more accurate than the ones present in the literature. We provide rigorous error estimates for the estimators. Furthermore we use the detected boundary points to solve boundary-value problems for PDE on point clouds. We prove error estimates for the Laplace and eikonal equations on point clouds. Finally we provide a range of numerical experiments illustrating the performance of our boundary estimators, applications to PDE on point clouds, and tests on image data sets.

math.NA

Clustering dynamics on graphs: from spectral clustering to mean shift through Fokker-Planck interpolation

In this work we build a unifying framework to interpolate between density-driven and geometry-based algorithms for data clustering, and specifically, to connect the mean shift algorithm with spectral clustering at discrete and continuum levels. We seek this connection through the introduction of Fokker-Planck equations on data graphs. Besides introducing new forms of mean shift algorithms on graphs, we provide new theoretical insights on the behavior of the family of diffusion maps in the large sample limit as well as provide new connections between diffusion maps and mean shift dynamics on a fixed graph. Several numerical examples illustrate our theoretical findings and highlight the benefits of interpolating density-driven and geometry-based clustering algorithms.

stat.ML

The nonlocal-interaction equation near attracting manifolds

We study the approximation of the nonlocal-interaction equation restricted to a compact manifold $\mathcal{M}$ embedded in $\mathbb{R}^d$, and more generally compact sets with positive reach (i.e. prox-regular sets). We show that the equation on $\mathcal{M}$ can be approximated by the classical nonlocal-interaction equation on $\mathbb{R}^d$ by adding an external potential which strongly attracts to $\mathcal{M}$. The proof relies on the Sandier--Serfaty approach to the $Γ$-convergence of gradient flows. As a by-product, we recover well-posedness for the nonlocal-interaction equation on $\mathcal{M}$. Uniqueness, on the other hand, is established using a stability argument. We also provide an another approximation to the interaction equation on $\mathcal{M}$, based on iterating approximately solving an interaction equation on $\mathbb{R}^d$ and projecting to $\mathcal{M}$. We show convergence of this scheme, together with an estimate on the rate of convergence. Finally, we conduct numerical experiments, for both the attractive-potential-based and the projection-based approaches, that highlight the effects of the geometry on the dynamics.

math.AP

Nonlocal-interaction equation on graphs: gradient flow structure and continuum limit

We consider dynamics driven by interaction energies on graphs. We introduce graph analogues of the continuum nonlocal-interaction equation and interpret them as gradient flows with respect to a graph Wasserstein distance. The particular Wasserstein distance we consider arises from the graph analogue of the Benamou-Brenier formulation where the graph continuity equation uses an upwind interpolation to define the density along the edges. While this approach has both theoretical and computational advantages, the resulting distance is only a quasi-metric. We investigate this quasi-metric both on graphs and on more general structures where the set of "vertices" is an arbitrary positive measure. We call the resulting gradient flow of the nonlocal-interaction energy the nonlocal nonlocal-interaction equation (NL$^2$IE). We develop the existence theory for the solutions of the NL$^2$IE as curves of maximal slope with respect to the upwind Wasserstein quasi-metric. Furthermore, we show that the solutions of the NL$^2$IE on graphs converge as the empirical measures of the set of vertices converge weakly, which establishes a valuable discrete-to-continuum convergence result.

math.AP

Rates of Convergence for Laplacian Semi-Supervised Learning with Low Labeling Rates

We study graph-based Laplacian semi-supervised learning at low labeling rates. Laplacian learning uses harmonic extension on a graph to propagate labels. At very low label rates, Laplacian learning becomes degenerate and the solution is roughly constant with spikes at each labeled data point. Previous work has shown that this degeneracy occurs when the number of labeled data points is finite while the number of unlabeled data points tends to infinity. In this work we allow the number of labeled data points to grow to infinity with the number of labels. Our results show that for a random geometric graph with length scale $\varepsilon>0$ and labeling rate $β>0$, if $β\ll\varepsilon^2$ then the solution becomes degenerate and spikes form, and if $β\gg \varepsilon^2$ then Laplacian learning is well-posed and consistent with a continuum Laplace equation. Furthermore, in the well-posed setting we prove quantitative error estimates of $O(\varepsilonβ^{-1/2})$ for the difference between the solutions of the discrete problem and continuum PDE, up to logarithmic factors. We also study $p$-Laplacian regularization and show the same degeneracy result when $β\ll \varepsilon^p$. The proofs of our well-posedness results use the random walk interpretation of Laplacian learning and PDE arguments, while the proofs of the ill-posedness results use $Γ$-convergence tools from the calculus of variations. We also present numerical results on synthetic and real data to illustrate our results.

math.ST

Mumford-Shah functionals on graphs and their asymptotics

We consider adaptations of the Mumford-Shah functional to graphs. These are based on discretizations of nonlocal approximations to the Mumford-Shah functional. Motivated by applications in machine learning we study the random geometric graphs associated to random samples of a measure. We establish the conditions on the graph constructions under which the minimizers of graph Mumford-Shah functionals converge to a minimizer of a continuum Mumford-Shah functional. Furthermore we explicitly identify the limiting functional. Moreover we describe an efficient algorithm for computing the approximate minimizers of the graph Mumford-Shah functional.

math.AP

Large Data and Zero Noise Limits of Graph-Based Semi-Supervised Learning Algorithms

Scalings in which the graph Laplacian approaches a differential operator in the large graph limit are used to develop understanding of a number of algorithms for semi-supervised learning; in particular the extension, to this graph setting, of the probit algorithm, level set and kriging methods, are studied. Both optimization and Bayesian approaches are considered, based around a regularizing quadratic form found from an affine transformation of the Laplacian, raised to a, possibly fractional, exponent. Conditions on the parameters defining this quadratic form are identified under which well-defined limiting continuum analogues of the optimization and Bayesian semi-supervised learning problems may be found, thereby shedding light on the design of algorithms in the large graph setting. The large graph limits of the optimization formulations are tackled through $Γ-$convergence, using the recently introduced $TL^p$ metric. The small labelling noise limits of the Bayesian formulations are also identified, and contrasted with pre-existing harmonic function approaches to the problem.

stat.ML

Least action principles for incompressible flows and geodesics between shapes

As V. I. Arnold observed in the 1960s, the Euler equations of incompressible fluid flow correspond formally to geodesic equations in a group of volume-preserving diffeomorphisms. Working in an Eulerian framework, we study incompressible flows of shapes as critical paths for action (kinetic energy) along transport paths constrained to have characteristic-function densities. The formal geodesic equations for this problem are Euler equations for incompressible, inviscid potential flow of fluid with zero pressure and surface tension on the free boundary. The problem of minimizing this action exhibits an instability associated with microdroplet formation, with the following outcomes: Any two shapes of equal volume can be approximately connected by an Euler spray---a countable superposition of ellipsoidal geodesics. The infimum of the action is the Wasserstein distance squared, and is almost never attained except in dimension 1. Every Wasserstein geodesic between bounded densities of compact support provides a solution of the (compressible) pressureless Euler system that is a weak limit of (incompressible) Euler sprays. Each such Wasserstein geodesic is also the unique minimizer of a relaxed least-action principle for a two-fluid mixture theory corresponding to incompressible fluid mixed with vacuum.

math.AP

Analysis of $p$-Laplacian Regularization in Semi-Supervised Learning

We investigate a family of regression problems in a semi-supervised setting. The task is to assign real-valued labels to a set of $n$ sample points, provided a small training subset of $N$ labeled points. A goal of semi-supervised learning is to take advantage of the (geometric) structure provided by the large number of unlabeled data when assigning labels. We consider random geometric graphs, with connection radius $ε(n)$, to represent the geometry of the data set. Functionals which model the task reward the regularity of the estimator function and impose or reward the agreement with the training data. Here we consider the discrete $p$-Laplacian regularization. We investigate asymptotic behavior when the number of unlabeled points increases, while the number of training points remains fixed. We uncover a delicate interplay between the regularizing nature of the functionals considered and the nonlocality inherent to the graph constructions. We rigorously obtain almost optimal ranges on the scaling of $ε(n)$ for the asymptotic consistency to hold. We prove that the minimizers of the discrete functionals in random setting converge uniformly to the desired continuum limit. Furthermore we discover that for the standard model used there is a restrictive upper bound on how quickly $ε(n)$ must converge to zero as $n \to \infty$. We introduce a new model which is as simple as the original model, but overcomes this restriction.

math.ST

A Transportation $L^p$ Distance for Signal Analysis

Transport based distances, such as the Wasserstein distance and earth mover's distance, have been shown to be an effective tool in signal and image analysis. The success of transport based distances is in part due to their Lagrangian nature which allows it to capture the important variations in many signal classes. However these distances require the signal to be nonnegative and normalized. Furthermore, the signals are considered as measures and compared by redistributing (transporting) them, which does not directly take into account the signal intensity. Here we study a transport-based distance, called the $TL^p$ distance, that combines Lagrangian and intensity modelling and is directly applicable to general, non-positive and multi-channelled signals. The framework allows the application of existing numerical methods. We give an overview of the basic properties of this distance and applications to classification, with multi-channelled, non-positive one and two-dimensional signals, and color transfer.

cs.CV