SearcharxivSearch

arXiv subjects

Massimo Fornasier

Publications and source records attributed to Massimo Fornasier.

At least 19 recordsLinked to original sources

The nonlocal attraction-repulsion transport equation with power kernels

We study a nonlocal continuity equation on $\mathbb{R}^d$ in which a probability density is driven by the competition between attraction toward a prescribed background measure $\omega$ and self-repulsion among particles, governed respectively by the power-law kernels $\psi_a(x) = |x|^{1+a}$ and $\psi_r(x) = |x|^{1+r}$ with exponents $a, r \in [0,1)$. We establish global Lagrangian well-posedness via a squared-radius regularization, obtaining uniform $L^\infty$ and moment bounds, $W^{n,\infty}$ regularity, and uniqueness in the Lagrangian class. When the initial data is compactly supported and attraction dominates ($a > r$, or $a = r$ with $\omega(\mathbb{R}^d) > 1$), we prove that the support remains uniformly bounded at all time; a counterexample shows this fails for $a = r > 1$. For the attractive-dominant nonquadratic range $0 \leq r \leq a < 1$, we characterize zero-flux stationary states via a free-boundary problem involving a fractional Laplacian operator, reducing the stationarity condition to a fractional exterior Dirichlet problem. This characterization allow us to exhibit explicit examples of stationary measures in dimensions $d \in \{1,2,3\}$. Numerical particle simulations confirm agreement with the theoretical stationary profiles. Finally, we prove that every global solution with bounded energy and uniform moment bounds converges to a zero-flux stationary state.

math.AP

Approximation in Metric Sobolev Spaces: A General Framework

In our recent work [FHS25], we introduced a numerical framework for approximating Sobolev functions on Wasserstein spaces from finite samples, leveraging structural properties established in [FSS23]. The present paper demonstrates that this methodology extends far beyond that specific setting. We identify a general class of metric measure spaces -- including weighted Riemannian manifolds and spaces of measures equipped with the Hellinger--Kantorovich distance -- for which the key hypotheses of Hilbertianity and the existence of a computable algebra of Lipschitz functions hold. Within this abstract framework, we recover and generalize the core approximation results of [FHS25] for recovering functions from random point evaluations. Our main contribution is to show that the combination of theoretical foundations from [FSS23] and algorithmic strategies from [FHS25] is robust enough to apply to a wide variety of infinite-dimensional spaces of current interest.

math.FA

Sharp Rates of MMD Empirical Estimation with Power Kernels

We establish quantitative rates of convergence for the empirical estimation of probability measures by means of the Maximum Mean Discrepancy (MMD) with power kernel $K_q(x,y) = -|x-y|^q$, $q \in (0,2)$. The resulting discrepancy is the classical \emph{energy distance} $$\mathcal E_q^2(\mu, \omega) = -\frac{1}{2}\iint_{\mathbb{R}^d \times \mathbb{R}^d} |x-y|^q \, d(\mu - \omega)(x)\, d(\mu - \omega)(y),$$ and we ask how fast the best $N$-point empirical approximation $\inf_{\mu_N \in \mathcal{P}^N}\mathcal{E}_q(\mu_N,\omega)$ decays as $N \to \infty$. Given a probability measure $\omega$ on $\mathbb{R}^d$ with compact support satisfying an Ahlfors regularity condition of exponent $\beta \in (0,d]$, we prove that the sharp two-sided bound $$\mathcal E_q(\mu_N, \omega) \asymp N^{-\frac{1}{2}\left(1 + \frac{q}{\beta}\right)}$$ holds both for the worst-case empirical measure $\mu_N$ (lower bound, holding for every configuration of $N$ points) and for an optimally chosen empirical measure $\mu_N$ (upper bound). This complements the qualitative consistency result of Fornasier and H\"utter \cite{fornasier2014consistency}, who proved narrow convergence of the minimizers of $\mathcal E_q^2(\cdot, \omega)$ over empirical measures without quantitative rates.

math.PR

From Consensus-Based Optimization to Evolution Strategies: Proof of Global Convergence

Consensus-based optimization (CBO) is a powerful and versatile zero-order multi-particle method designed to provably solve high-dimensional global optimization problems, including those that are genuinely nonconvex or nonsmooth. The method relies on a balance between stochastic exploration and contraction toward a consensus point, which is defined via the Laplace principle as a proxy for the global minimizer. In this paper, we introduce new CBO variants that address practical and theoretical limitations of the original formulation of this novel optimization methodology. First, we propose a model called $\delta$-CBO}, which incorporates nonvanishing diffusion to prevent premature collapse to suboptimal states. We also develop a numerically stable implementation, the Consensus Freezing scheme, that remains robust even for arbitrarily large time steps by freezing the consensus point over time intervals. We connect these models through appropriate asymptotic limits. Furthermore, we derive from the Consensus Freezing scheme by suitable time rescaling and asymptotics a further algorithm, the Consensus Hopping scheme, which can be interpreted as a form of $(1,\lambda)$-Evolution Strategy. For all these schemes, we characterize for the first time the invariant measures and establish global convergence results, including exponential convergence rates.

math.OC

Consensus-based optimization (CBO): Towards Global Optimality in Robotics

Zero-order optimization has recently received significant attention for designing optimal trajectories and policies for robotic systems. However, most existing methods (e.g., MPPI, CEM, and CMA-ES) are local in nature, as they rely on gradient estimation. In this paper, we introduce consensus-based optimization (CBO) to robotics, which is guaranteed to converge to a global optimum under mild assumptions. We provide theoretical analysis and illustrative examples that give intuition into the fundamental differences between CBO and existing methods. To demonstrate the scalability of CBO for robotics problems, we consider three challenging trajectory optimization scenarios: (1) a long-horizon problem for a simple system, (2) a dynamic balance problem for a highly underactuated system, and (3) a high-dimensional problem with only a terminal cost. Our results show that CBO is able to achieve lower costs with respect to existing methods on all three challenging settings. This opens a new framework to study global trajectory optimization in robotics.

cs.RO

Large-Time Analysis of the Langevin Dynamics for Energies Fulfilling Polyak-{\L}ojasiewicz Conditions

In this work, we take a step towards understanding overdamped Langevin dynamics for the minimization of a general class of objective functions $\mathcal{L}$. We establish well-posedness and regularity of the law $\rho_t$ of the process through novel a priori estimates, and, very importantly, we characterize the large-time behavior of $\rho_t$ under truly minimal assumptions on $\mathcal{L}$. In the case of integrable Gibbs density, the law converges to the normalized Gibbs measure. In the non-integrable case, we prove that the law diffuses. The rate of convergence is $\mathcal{O}(1/t)$. Under a Polyak-Lojasiewicz (PL) condition on $\mathcal{L}$, we also derive sharp exponential contractivity results toward the set of global minimizers. Combining these results we provide the first systematic convergence analysis of Langevin dynamics under PL conditions in non-integrable Gibbs settings: a first phase of exponential in time contraction toward the set of minimizers and then a large-time exploration over it with rate $\mathcal{O}(1/t)$.

math.AP

Balanced quasistatic evolutions of critical points in metric spaces

Quasistatic evolutions of critical points of time-dependent energies exhibit piecewise smooth behavior, making them useful for modeling continuum mechanics phenomena like elastic-plasticity and fracture. Traditionally, such evolutions have been derived as vanishing viscosity and inertia limits, leading to balanced viscosity solutions. However, for nonconvex energies, these constructions have been realized in Euclidean spaces and assume non-degenerate critical points. In this paper, we take a different approach by decoupling the time scales of the energy evolution and of the transition to equilibria. Namely, starting from an equilibrium configuration, we let the energy evolve, while keeping frozen the system state; then, we update the state by freezing the energy, while letting the system transit via gradient flow or an approximation of it (e.g., minimizing movement or backward differentiation schemes). This approach has several advantages. It aligns with the physical principle that systems transit through energy-minimizing steady states. It is also fully constructive and computationally implementable, with physical and computational costs governed by appropriate action functionals. Additionally, our analysis is simpler and more general than previous formulations in the literature, as it does not require non-degenerate critical points. Finally, this approach extends to evolutions in locally compact metric path spaces, and our axiomatic presentation allows for various realizations.

math.OC

Regularity and positivity of solutions of the Consensus-Based Optimization equation: unconditional global convergence

Introduced in 2017 \cite{B1-pinnau2017consensus}, Consensus-Based Optimization (CBO) has rapidly emerged as a significant breakthrough in global optimization. This straightforward yet powerful multi-particle, zero-order optimization method draws inspiration from Simulated Annealing and Particle Swarm Optimization. Using a quantitative mean-field approximation, CBO dynamics can be described by a nonlinear Fokker-Planck equation with degenerate diffusion, which does not follow a gradient flow structure. In this paper, we demonstrate that solutions to the CBO equation remain positive and maintain full support. Building on this foundation, we establish the {\it unconditional} global convergence of CBO methods to global minimizers. Our results are derived through an analysis of solution regularity and the proof of existence for smooth, classical solutions to a broader class of drift-diffusion equations, despite the challenges posed by degenerate diffusion.

math.AP

Trade-off Invariance Principle for minimizers of regularized functionals

In this paper, we consider functionals of the form $H_\alpha(u)=F(u)+\alpha G(u)$ with $\alpha\in[0,+\infty)$, where $u$ varies in a set $U\neq\emptyset$ (without further structure). We first revisit a result stating that, excluding at most countably many values of $\alpha$, we have $\inf_{H_\alpha^\star}G= \sup_{H_\alpha^\star}G$, where $H_\alpha^\star := \arg\min_UH_\alpha$, which is assumed to be non-empty. Then, we prove a stronger result that concerns the invariance of the limiting value of the functional $G$ along minimizing sequences for $H_\alpha$, which extends the above Principle to the case $H_\alpha^\star= \emptyset$. Moreover, we show to what extent these findings generalize to multi-regularized functionals and -- in the presence of an underlying differentiable structure -- to critical points. Finally, the main result implies an unexpected consequence for functionals regularized with uniformly convex norms: excluding again at most countably many values of $\alpha$, it turns out that for a minimizing sequence, convergence to a minimizer in the weak or strong sense is equivalent.

math.OC

Constrained Consensus-Based Optimization and Numerical Heuristics for the Few Particle Regime

Consensus-based optimization (CBO) is a versatile multi-particle optimization method for performing nonconvex and nonsmooth global optimizations in high dimensions. Proofs of global convergence in probability have been achieved for a broad class of objective functions in unconstrained optimizations. In this work we adapt the algorithm for solving constrained optimizations on compact and unbounded domains with boundary by leveraging emerging reflective boundary conditions. In particular, we close a relevant gap in the literature by providing a global convergence proof for the many-particle regime comprehensive of convergence rates. On the one hand, for the sake of minimizing running cost, it is desirable to keep the number of particles small. On the other hand, reducing the number of particles implies a diminished capability of exploration of the algorithm. Hence numerical heuristics are needed to ensure convergence of CBO in the few-particle regime. In this work, we also significantly improve the convergence and complexity of CBO by utilizing an adaptive region control mechanism and by choosing geometry-specific random noise. In particular, by combining a hierarchical noise structure with a multigrid finite element method, we are able to compute global minimizers for a constrained $p$-Allen-Cahn problem with obstacles, a very challenging variational problem.

math.OC

Fast training of accurate physics-informed neural networks without gradient descent

Solving time-dependent Partial Differential Equations (PDEs) is one of the most critical problems in computational science. While Physics-Informed Neural Networks (PINNs) offer a promising framework for approximating PDE solutions, their accuracy and training speed are limited by two core barriers: gradient-descent-based iterative optimization over complex loss landscapes and non-causal treatment of time as an extra spatial dimension. We present Frozen-PINN, a novel PINN based on the principle of space-time separation that leverages random features instead of training with gradient descent, and incorporates temporal causality by construction. On eight PDE benchmarks, including challenges such as extreme advection speeds, shocks, and high dimensionality, Frozen-PINNs achieve superior training efficiency and accuracy over state-of-the-art PINNs, often by several orders of magnitude. Our work addresses longstanding training and accuracy bottlenecks of PINNs, delivering quickly trainable, highly accurate, and inherently causal PDE solvers, a combination that prior methods could not realize. Our approach challenges the reliance of PINNs on stochastic gradient-descent-based methods and specialized hardware, leading to a paradigm shift in PINN training and providing a challenging benchmark for the community.

math.NA

A PDE Framework of Consensus-Based Optimization for Objectives with Multiple Global Minimizers

Consensus-based optimization (CBO) is an agent-based derivative-free method for non-smooth global optimization that has been introduced in 2017, leveraging a surprising interplay between stochastic exploration and Laplace principle. In addition to its versatility and effectiveness in handling high-dimensional, non-convex, and non-smooth optimization problems, this approach lends itself well to theoretical analysis. Indeed, its dynamics is governed by a degenerate nonlinear Fokker--Planck equation, whose large time behavior explains the convergence of the method. Recent results provide guarantees of convergence under the restrictive assumption of a unique global minimizer for the objective function. In this work, we propose a novel and simple variation of CBO to tackle non-convex optimization problems with multiple global minimizers. Despite the simplicity of this new model, its analysis is particularly challenging because of its nonlinearity and nonlocal nature. We prove the existence of solutions of the corresponding nonlinear Fokker--Planck equation and we show exponential concentration in time to the set of minimizers made of multiple smooth, convex, and compact components. Our proofs require combining several ingredients, such as delicate geometrical arguments, new variants of a quantitative Laplace principle, ad hoc regularizations and approximations, and regularity theory for parabolic equations. Ultimately, this result suggests that the corresponding CBO algorithm, formulated as an Euler-Maruyama discretization of the underlying empirical stochastic process, tends to converge to multiple global minimizers.

math.AP

Approximation Theory, Computing, and Deep Learning on the Wasserstein Space

The challenge of approximating functions in infinite-dimensional spaces from finite samples is widely regarded as formidable. We delve into the challenging problem of the numerical approximation of Sobolev-smooth functions defined on probability spaces. Our particular focus centers on the Wasserstein distance function, which serves as a relevant example. In contrast to the existing body of literature focused on approximating efficiently pointwise evaluations, we chart a new course to define functional approximants by adopting three machine learning-based approaches: 1. Solving a finite number of optimal transport problems and computing the corresponding Wasserstein potentials. 2. Employing empirical risk minimization with Tikhonov regularization in Wasserstein Sobolev spaces. 3. Addressing the problem through the saddle point formulation that characterizes the weak form of the Tikhonov functional's Euler-Lagrange equation. We furnish explicit and quantitative bounds on generalization errors for each of these solutions. We leverage the theory of metric Sobolev spaces and we combine it with techniques of optimal transport, variational calculus, and large deviation bounds. In our numerical implementation, we harness appropriately designed neural networks to serve as basis functions. These networks undergo training using diverse methodologies. This approach allows us to obtain approximating functions that can be rapidly evaluated after training. Our constructive solutions significantly enhance at equal accuracy the evaluation speed, surpassing that of state-of-the-art methods by several orders of magnitude. This allows evaluations over large datasets several times faster, including training, than traditional optimal transport algorithms. Our analytically designed deep learning architecture slightly outperforms the test error of state-of-the-art CNN architectures on datasets of images.

math.OC

Consensus-Based Optimization with Truncated Noise

Consensus-based optimization (CBO) is a versatile multi-particle metaheuristic optimization method suitable for performing nonconvex and nonsmooth global optimizations in high dimensions. It has proven effective in various applications while at the same time being amenable to a theoretical convergence analysis. In this paper, we explore a variant of CBO, which incorporates truncated noise in order to enhance the well-behavedness of the statistics of the law of the dynamics. By introducing this additional truncation in the noise term of the CBO dynamics, we achieve that, in contrast to the original version, higher moments of the law of the particle system can be effectively bounded. As a result, our proposed variant exhibits enhanced convergence performance, allowing in particular for wider flexibility in choosing the noise parameter of the method as we confirm experimentally. By analyzing the time-evolution of the Wasserstein-$2$ distance between the empirical measure of the interacting particle system and the global minimizer of the objective function, we rigorously prove convergence in expectation of the proposed CBO variant requiring only minimal assumptions on the objective function and on the initialization. Numerical evidences demonstrate the benefit of truncating the noise in CBO.

math.OC

Density of subalgebras of Lipschitz functions in metric Sobolev spaces and applications to Wasserstein Sobolev spaces

We prove a general criterion for the density in energy of suitable subalgebras of Lipschitz functions in the metric-Sobolev space $H^{1,p}(X,\mathsf{d},\mathfrak{m})$ associated with a positive and finite Borel measure $\mathfrak{m}$ in a separable and complete metric space $(X,\mathsf{d})$. We then provide a relevant application to the case of the algebra of cylinder functions in the Wasserstein Sobolev space $H^{1,2}(\mathcal{P}_2(\mathbb{M}),W_{2},\mathfrak{m})$ arising from a positive and finite Borel measure $\mathfrak{m}$ on the Kantorovich-Rubinstein-Wasserstein space $(\mathcal{P}_2(\mathbb{M}),W_{2})$ of probability measures in a finite dimensional Euclidean space, a complete Riemannian manifold, or a separable Hilbert space $\mathbb{M}$. We will show that such a Sobolev space is always Hilbertian, independently of the choice of the reference measure $\mathfrak{m}$ so that the resulting Cheeger energy is a Dirichlet form. We will eventually provide an explicit characterization for the corresponding notion of $\mathfrak{m}$-Wasserstein gradient, showing useful calculus rules and its consistency with the tangent bundle and the $Γ$-calculus inherited from the Dirichlet form.

math.FA

From NeurODEs to AutoencODEs: a mean-field control framework for width-varying Neural Networks

The connection between Residual Neural Networks (ResNets) and continuous-time control systems (known as NeurODEs) has led to a mathematical analysis of neural networks which has provided interesting results of both theoretical and practical significance. However, by construction, NeurODEs have been limited to describing constant-width layers, making them unsuitable for modeling deep learning architectures with layers of variable width. In this paper, we propose a continuous-time Autoencoder, which we call AutoencODE, based on a modification of the controlled field that drives the dynamics. This adaptation enables the extension of the mean-field control framework originally devised for conventional NeurODEs. In this setting, we tackle the case of low Tikhonov regularization, resulting in potentially non-convex cost landscapes. While the global results obtained for high Tikhonov regularization may not hold globally, we show that many of them can be recovered in regions where the loss function is locally convex. Inspired by our theoretical findings, we develop a training method tailored to this specific type of Autoencoders with residual connections, and we validate our approach through numerical experiments conducted on various examples.

math.OC

Gradient is All You Need? How Consensus-Based Optimization can be Interpreted as a Stochastic Relaxation of Gradient Descent

In this paper, we provide a novel analytical perspective on the theoretical understanding of gradient-based learning algorithms by interpreting consensus-based optimization (CBO), a recently proposed multi-particle derivative-free optimization method, as a stochastic relaxation of gradient descent. Remarkably, we observe that through communication of the particles, CBO exhibits a stochastic gradient descent (SGD)-like behavior despite solely relying on evaluations of the objective function. The fundamental value of such link between CBO and SGD lies in the fact that CBO is provably globally convergent to global minimizers for ample classes of nonsmooth and nonconvex objective functions. Hence, on the one side, we offer a novel explanation for the success of stochastic relaxations of gradient descent by furnishing useful and precise insights that explain how problem-tailored stochastic perturbations of gradient descent (like the ones induced by CBO) overcome energy barriers and reach deep levels of nonconvex functions. On the other side, and contrary to the conventional wisdom for which derivative-free methods ought to be inefficient or not to possess generalization abilities, our results unveil an intrinsic gradient descent nature of heuristics. Instructive numerical illustrations support the provided theoretical insights.

cs.LG

Finite Sample Identification of Wide Shallow Neural Networks with Biases

Artificial neural networks are functions depending on a finite number of parameters typically encoded as weights and biases. The identification of the parameters of the network from finite samples of input-output pairs is often referred to as the \emph{teacher-student model}, and this model has represented a popular framework for understanding training and generalization. Even if the problem is NP-complete in the worst case, a rapidly growing literature -- after adding suitable distributional assumptions -- has established finite sample identification of two-layer networks with a number of neurons $m=\mathcal O(D)$, $D$ being the input dimension. For the range $D<m<D^2$ the problem becomes harder, and truly little is known for networks parametrized by biases as well. This paper fills the gap by providing constructive methods and theoretical guarantees of finite sample identification for such wider shallow networks with biases. Our approach is based on a two-step pipeline: first, we recover the direction of the weights, by exploiting second order information; next, we identify the signs by suitable algebraic evaluations, and we recover the biases by empirical risk minimization via gradient descent. Numerical results demonstrate the effectiveness of our approach.

cs.LG