Searcharxiv⌕ Search

arXiv subjects

Maarten V. de Hoop

Publications and source records attributed to Maarten V. de Hoop.

At least 19 recordsLinked to original sources

Target-adapted Green-Bessel SVGD: uniform-in-time propagation of chaos and last-iterate consistency

We prove uniform-in-time propagation of chaos and last-iterate consistency for a target-adapted Stein variational gradient descent (SVGD) flow on compact connected manifolds. The target has a smooth positive density, and the particles start independently from a fixed smooth nonnegative density ratio. The construction uses the Green--Bessel operator $Q_{r,π}=A_π^{-1}(\mathrm{Id}+A_π)^{-r}$ of the reversible target Langevin generator. Sufficient Bessel smoothing gives a scalar kernel with finite diagonal, and a positive matrix lift realizes its potential force as a Stein velocity. Population and empirical flows then dissipate the same finite target discrepancy. Population entropy and the target spectral gap give decay of this discrepancy; a finite-time particle comparison reaches a time after which common-energy monotonicity controls every later time. The resulting expected uniform discrepancy is $O((\log N)^{-1/2})$, with a corresponding logarithmic $W_1$ bound and consistency along every sequence $t_N\to\infty$. We also prove an exact finite-mode approximation theorem with an explicit spatial-resolution error and a feature-factorized particle implementation. For confining Euclidean targets, we establish static kernel and moment results and give a conditional dynamical extension under explicit population-regularity and transport hypotheses.

math.PR↗

Generic Recovery of Permittivity and Permeability in Anisotropic Maxwell Systems

We study the inverse problem of recovering the constitutive tensors of a homogeneous anisotropic electromagnetic medium without magnetoelectric coupling (non-chiral) from its Fresnel surface, the characteristic variety of Maxwell's equations governing electromagnetic wave propagation. For known isotropic permeability, normalized to $μ= I$, we prove that the Fresnel surface uniquely determines the permittivity tensor $\varepsilon$, and that the associated Fresnel polynomial is reducible precisely when $\varepsilon$ has a repeated eigenvalue. For general, positive-definite symmetric tensors $(\varepsilon,μ)$, we prove that the Fresnel polynomial is generically irreducible over $\mathbb{C}$ and we identify the natural gauge symmetry under which it is invariant. Using geometric invariant theory, a powerful tool of modern algebraic geometry, we construct an affine quotient of the parameter space by this gauge action and prove that the induced Fresnel-polynomial map is birational onto its image. We deduce that, outside a proper real algebraic exceptional set, the real Fresnel surface determines $(\varepsilon,μ)$ up to gauge. This establishes generic uniqueness for the inverse problem and, to our knowledge, provides a new application of affine geometric invariant theory to gauge freedom in a PDE inverse problem.

math.AP↗

Riesz-Kernel Stein Variational Gradient Descent: Renormalized Entropy and Long-Time Particle Limits

Stein variational gradient descent (SVGD) transports interacting particles toward a target distribution through deterministic kernelized dynamics. Singular Riesz kernels are attractive because they can provide quantitative population-level convergence, but at the finite-particle level the corresponding Stein energy has infinite self-interaction. We study periodic Riesz SVGD with self-interaction removed and prove a many-particle, long-time sampling theorem. Throughout the range in which the singular Stein energy is locally integrable, under a uniform bound on the initial relative entropy per particle, the time-averaged empirical-measure law converges weakly to the point mass \(δ_π\) at the target as the particle number and any diverging averaging horizon tend to infinity. We also show that the empirical-measure laws induced by invariant particle laws of finite relative entropy converge weakly to \(δ_π\), without a uniform entropy bound. Below the logarithmic singularity threshold, we obtain an explicit algebraic finite-particle error bound. These results extend the joint-entropy approach for smooth-kernel SVGD to singular interactions.

math.AP↗

Classification of singularities of planar slowness surfaces

Slowness surfaces are algebraic varieties arising from propagation of elastic waves. In dimensions $2$, we completely classify the types of singularities slowness surfaces can have. The two types of possible singularities are a transversal self-intersection and a tangential singularity produced by a concentric circle and ellipse that are tangent to each other. To interpret these results analytically, in the case that the slowness surface has transversal self-intersections, we show that the principal symbol of the elastic wave operator is locally smoothly diagonalizable.

math.AG↗

Weyl formulae for some singular metrics with application to acoustic modes in gas giants

This paper is motivated by recent works on inverse problems for acoustic wave propagation in the interior of gas giant planets. In such planets, the speed of sound is isotropic and tends to zero at the surface. Geometrically, this corresponds to a Riemannian manifold with boundary whose metric blows up near the boundary. Here, the spectral analysis of the corresponding Laplace-Beltrami operator is presented and the Weyl law is derived. The involved exponents depend on the Hausdorff dimension which, in the supercritical case, is larger than the topological dimension.

math.AP↗

The head-wave transform

We introduce and study a new integral ray transform called the head-wave transform. The head-wave transform integrates a function along a piecewise linear (in general geodesic) path consisting of three parts. The geometry of such paths corresponds to ray paths of head-waves propagating in a medium with sharp changes in sound speed. The middle part of the ray paths corresponds to gliding along the so-called gliding surface. As our main results, we prove inversion formulas and null space characterizations under multiple different sets of assumptions on the geometry of the gliding surface and the integrand function.

math.CA↗

Stable recovery of a simple irreversible Finsler geometry from travel time data

We show that a simple irreversible Finsler geometry can be recovered uniquely and Lipschitz-stably from its travel time data. We introduce and use a version of Gromov--Hausdorff distance adapted to irreversible metric spaces. In contrast to reversible (e.g. Riemannian) geometry, even the question of stability becomes ill-defined without simplicity.

math.MG↗

Flowers: A Warp Drive for Neural PDE Solvers

We introduce Flowers, a neural architecture for learning PDE solution operators built entirely from multihead warps. Aside from pointwise channel mixing and a multiscale scaffold, Flowers use no Fourier multipliers, no dot-product attention, and no convolutional mixing. Each head predicts a displacement field and warps the mixed input features. Motivated by physics and computational efficiency, displacements are predicted pointwise, without any spatial aggregation, and nonlocality enters only through sparse sampling at source coordinates, one per head. Stacking warps in multiscale residual blocks yields Flowers, which implement adaptive, global interactions at linear cost. We theoretically motivate this design through three complementary lenses: flow maps for conservation laws, waves in inhomogeneous media, and a kinetic-theoretic continuum limit. Flowers achieve excellent performance on a broad suite of 2D and 3D time-dependent PDE benchmarks, particularly flows and waves. A compact 17M-parameter model consistently outperforms Fourier, convolution, and attention-based baselines of similar size, while a 150M-parameter variant improves over recent transformer-based foundation models with much more parameters, data, and training compute.

cs.LG↗

Measure flow path recovery in Bayes Hilbert spaces

We study the ill-posed problem of recovering a probability measure flow from finitely many moving localized sensors using a Bayes Hilbert framework. Relative to a fixed reference probability measure, a probability law is represented by its centered log-ratio coordinates, so that an evolving law becomes a path in a Hilbert space of functions. For sufficiently regular Bayes Hilbert paths, we construct a canonical minimum-energy transport realization of the path by solving a weighted Neumann problem at each time, yielding an intrinsic transport form on tangent directions. We then formulate an inverse problem directly on Bayes Hilbert path space. Linearization of an observation operator yields an observability form, and recoverability is governed by its interaction with the transport geometry through a joint transport--observability form. In the ambient infinite-dimensional setting, we develop a regularized variational theory and identify limitations of localized sensing: mobile sensors can make the joint form injective, but they do not in general yield a coercive stability estimate on the full state space. This obstruction leads naturally to finite-dimensional Bayes Hilbert reductions. There the transport form becomes a kinetic tensor and the linearized observations become reduced sensing matrices, so recoverability can be expressed through explicit Gramian conditions. We show that localized bump sensors detect every fixed reduced direction, that finitely many suitably placed static sensors yield uniform reduced observability, and there exist path-dependent sensor trajectories such that even a single moving sensor can recover the reduced path. Finally, we show that these reduced recovery results lift to approximate ambient recovery for paths that are well approximated by the chosen finite-dimensional subspaces, yielding stable reconstruction up to projection error.

stat.ML↗

Dimension-Free Multimodal Sampling via Preconditioned Annealed Langevin Dynamics

Designing sampling algorithms for multimodal targets that remain stable under refinement of the finite-dimensional approximation of an underlying function-space problem is a central challenge. Annealed Langevin dynamics (ALD) is a natural alternative to classical Langevin in this context, since it is often observed to improve exploration across modes. Yet a gap remains between its empirical success and existing theory: under which conditions can ALD be guaranteed to remain stable across dimensions? In this paper, we bridge this gap by providing a uniform-in-dimension analysis of continuous-time ALD for Gaussian-mixture targets. Along an explicit annealing path obtained by gradually removing Gaussian smoothing from the target, we identify spectral conditions linking the smoothing covariance to the component covariances under which ALD achieves a prescribed accuracy in Kullback-Leibler divergence within a dimension-uniform time horizon. We then establish stability in a perturbative regime with imperfect initialization and approximate scores. Under a misspecified-mixture score model, we show that preconditioning ALD with an operator whose spectrum decays sufficiently fast prevents error terms from accumulating across coordinates and thereby preserves dimension-uniform control.

math.NA↗

Function graph transformers universally approximate operators between function spaces

We study the approximation of nonlinear operators between function spaces by transformers. Our approach is to lift functions to measures supported on their graphs and leverage a recently introduced measure-theoretic view of transformers. A function $h$ is represented by its graph measure $γ_h$, with finite tokens $\{(x_j,h(x_j))\}_{j=1}^N$ being its empirical approximations. We show that this framework elegantly models discretization refinement via convergence of measures and provides a natural setting for operator learning. Within this framework, we introduce function graph transformers, a graph-preserving subclass of measure-theoretic transformers that maps graph measures to graph measures, which is to say that outputs remain single-valued functions. Crucially, this additional structure does not reduce generality: we prove that the resulting graph-preserving maps can be approximated by finite compositions of standard softmax self-attention layers and pointwise MLPs, yielding universal approximation results for broad classes of nonlinear operators. Unlike existing theoretical approaches to operator learning with transformers, the measure-theoretic framework also accommodates regularized negative-order Sobolev inputs for which discretization invariance is particularly challenging, as well as query points on different output domains. Overall, function graph transformers provide a continuum viewpoint and mathematical toolkit for transformer-based operator learning, clarifying the roles of positional encodings, graph structure, regularization, and ensuring consistency across discretizations.

cs.LG↗

Neural equilibria for long-term prediction of nonlinear conservation laws

Nonlinear conservation laws govern a broad class of important physical systems in science and industry and are central to scientific machine learning (SciML). Large general-purpose models offer speed, but replacing the numerical and physical structure of solvers often compromises stability, accuracy, and physical faithfulness. Here, we aim to balance the general inductive bias of conservation with the flexibility and speed of neural networks through a conservation-aware SciML backbone, which we call Neural Discrete Equilibrium (NeurDE). NeurDE places machine learning inside a kinetic solver by learning the local equilibrium closure of a Boltzmann formulation. The kinetic solver still performs transport, relaxation, moment recovery, and conservation; the neural network provides only the nonlinear equilibrium target. We test NeurDE on $6$ conserved systems, including three very challenging subsonic, transonic, and supersonic shock systems. NeurDE outperforms state-of-the-art SciML methods, including neural operators and pretrained SciML foundation models that are $10^4$ and $10^6$ times larger, respectively. Most notably, NeurDE improves upon the numerical method from which it is derived. NeurDE therefore provides a compact target for scientific machine learning in conservative simulation: learn the equilibrium law toward which the system relaxes, not the evolution law itself.

cs.LG↗

Training Infinitely Deep and Wide Transformers

Transformers have become the dominant architecture in modern machine learning, yet the theoretical understanding of their training dynamics remains limited. This paper develops a rigorous mathematical framework for analyzing gradient-based training of transformers in the mean-field regime, where both the depth (number of layers) and width (number of attention heads) tend to infinity. While ResNet training can be understood as controlling a neural ODE, transformer training corresponds to controlling a neural PDE, due to the coupling of multiple token distributions through the attention mechanism. Our mean-field model features two types of measure representations: token distributions evolving through layers and attention parameters at each layer. We establish well-posedness of the forward pass through infinitely deep transformers, characterizing token evolution via flow maps that satisfy ODEs in function spaces. Using adjoint sensitivity analysis, we derive an explicit formula for the conditional Wasserstein gradient of the training risk, involving adjoint variables governed by backward ODEs. We prove the existence and uniqueness of gradient flow curves in the conditional Wasserstein metric space, establishing a rigorous foundation for gradient-based transformer training. A key technical contribution is providing necessary and sufficient conditions for injectivity of the Neural Tangent Kernel (NTK) for attention mechanisms: we show that NTK injectivity is equivalent to linear independence of log-sum-exp functions modulo affine functions, a condition satisfied by diverse token distributions, including discrete distributions, uniform distributions, and Gaussian mixtures. Under this NTK injectivity assumption, we prove that gradient flow converges to global minima when the initial loss is sufficiently small, eliminating spurious local minima from the optimization landscape.

math.OC↗

Dimension-Uniform Discretization Analysis of Preconditioned Annealed Langevin Dynamics for Multimodal Gaussian Mixtures

Obtaining stable diffusion-based samplers in high- and infinite-dimensional settings is challenging because errors can accumulate across high-frequency coordinates and make the dynamics unstable under refinement of the finite-dimensional approximation of the underlying function-space problem. Discretization is a typical source of such errors, and preconditioning with a suitable spectral decay is one way to control their accumulation. In this paper, we study this problem for preconditioned annealed Langevin dynamics (ALD) applied to Gaussian mixtures. We first show that Euler-Maruyama (EM) discretization, by treating the stiff linear part of the annealed score with a forward Euler step, imposes a stability constraint coupling the preconditioner with the annealed covariance scale. Together with the conditions ensuring dimension-uniform control of the annealed dynamics, this constraint forces the initial smoothed law to remain uniformly close to the target across dimensions. We then consider an exponential-integrator scheme that integrates the stiff linear part of the annealed score exactly. Under explicit spectral summability conditions coupling the smoothing covariance, the component covariance spectra, and the preconditioner, we prove a dimension-uniform Kullback-Leibler (KL) bound for this scheme. This bound can be made arbitrarily small, uniformly in dimension, by allowing enough time for annealing and then refining the time mesh accordingly. Importantly, these conditions allow regimes in which the KL divergence between the target and the initial smoothed law diverges with dimension, showing that the restrictions imposed by EM are scheme-dependent rather than intrinsic to ALD.

stat.ML↗

Boundary observability for gas giant metrics

We study the observability of waves on gas giant manifolds which are a class of Riemannian manifolds whose metrics are singular at the boundary. Such manifolds arise naturally in modeling of acoustic wave propagation in gas giant planets.We establish an observability inequality using full boundary measurements given by a Neumann-type trace that is natural in the gas giant setting. The proof proceeds in two steps. First, observability for a general gas giant metric is reduced to the so-called separable case via a perturbation argument. In the separable case, we employ a uniform-in-tangential-frequency analysis combined with an Ingham inequality to prove observability.

math.OC↗

Moving localized observations and Ces{à}ro asymptotic observability for conservative PDEs

We develop a deterministic large-time mechanism yielding Ces{à}ro asymptotic observability inequalities from moving localized observations for conservative evolutions. On each observation interval, exact convexification on a compact measured homogeneous space replaces full observation on the whole observation manifold by a finite convex combination of translates of one prototype subset. A switching realization theorem then turns that static design into a genuinely moving observer, while a Hilbertian tail-reduction proposition shows that interval estimates proved only on growing spectral windows still recover the full conserved energy after Ces{à}ro averaging. The resulting design-to-observability chain applies to interior observations for wave, Klein-Gordon, and Schr{ö}dinger equations on compact measured homogeneous manifolds, to moving boundary caps on the Euclidean ball, and to a singular almost-separated gas-giant boundary model. The framework is especially relevant when each instantaneous observation set is too small for one to expect a finite-time GCC or time-dependent GCC statement.

math.AP↗

On Observation Time for Recovering Latent Hawkes Networks

Dynamics of interacting systems in engineering, society, and nature often evolve over latent networks that govern which entities can interact. We study the problem of inferring these networks from event-based observations, which arise naturally in finance, seismology, and neuroscience. While there is substantial algorithmic work addressing this important problem, theoretical results are scarce. In this paper we ask the following fundamental question: what is the minimum time that one must observe the dynamics in order to exactly recover the underlying network, as a function of the number $d$ of interacting entities? For a class of stationary Hawkes processes with sparse, weak interactions, we prove that an observation time of order $\log d$ is sufficient and necessary. For the upper bound we construct a two-stage estimator that uses clipped and binned event data for screening, followed by a least-squares refinement, and apply concentration bounds derived from the Poisson cluster representation. For the lower bound we combine Fano's inequality with Jacod's Girsanov formula for point processes on a suitable subclass of networks.

math.ST↗

Hybrid operator learning of wave scattering maps in high-contrast media

Surrogate modeling of wave propagation and scattering (i.e. the wave speed and source to wave field map) in heterogeneous media has significant potential in applications such as seismic imaging and inversion. High-contrast settings, such as subsurface models with salt bodies, exhibit strong scattering and phase sensitivity that challenge existing neural operators. We propose a hybrid architecture that decomposes the scattering operator into two separate contributions: a smooth background propagation and a high-contrast scattering correction. The smooth component is learned with a Fourier Neural Operator (FNO), which produces globally coupled feature tokens encoding background wave propagation; these tokens are then passed to a vision transformer, where attention is used to model the high-contrast scattering correction dominated by strong, spatial interactions. Evaluated on high-frequency Helmholtz problems with strong contrasts, the hybrid model achieves substantially improved phase and amplitude accuracy compared to standalone FNOs or transformers, with favorable accuracy-parameter scaling.

eess.SP↗