SearcharxivSearch

arXiv subjects

Michael Herty

Publications and source records attributed to Michael Herty.

At least 19 recordsLinked to original sources

A Mean-Field Theory of Transformers: Well-Posedness of the Coupled Data--Parameter Dynamics and Global Convergence of Training

We develop a rigorous mean-field theory for transformer networks that captures two large-scale limits inherent in the architecture: the number of tokens $N\to\infty$ in the input sequence and the number of attention heads $H\to\infty$ in each layer. It further considers an infinite number of layers leading to a time-continuous formulation. The resulting framework couples two interacting mean-field objects: a token distribution $\mu_t\in\mathcal P(\mathbb R^d)$, which evolves through network depth $t\in[0,T]$ according to a McKean--Vlasov transport equation, and the attention-parameter distribution $\rho_s\in\mathcal P(\Theta)$, that evolves through training time $s\ge0$ according to a Wasserstein gradient flow of the empirical risk, with optional entropic or Tikhonov regularization. We establish a comprehensive analytical foundation for the system coupling the transformer and the training dynamics; in particular, we establish global well-posedness of the resulting nonlinear Fokker--Planck system describing the coupled mean-field and training dynamics. Beyond well-posedness, we connect the mean-field formulation to optimization. For shallow, single-layer attention models, we prove exponential convergence to the entropy-regularized global optimum under a log-Sobolev condition. For genuinely deep, compositional transformers, we establish local linear convergence under a Neural Tangent Kernel non-degeneracy condition.

math.AP

Convergence of time-discrete finite particle consensus based optimization in Hilbert spaces

We study a time-discrete, finite-particle Consensus-Based Optimization (CBO) algorithm in a separable Hilbert space. Our analysis provides convergence guarantees directly for the computable particle system, complementing recent continuous-time and mean-field results in infinite dimensions. Using a common-noise formulation with trace-class covariance, we first establish quantitative pairwise contraction, exponential decay of the expected swarm variance, and almost-sure convergence of all particles to a common consensus state. We then combine estimates on the exponentiated objective functional with a quantitative Laplace principle to show that, for sufficiently large inverse temperature and suitably prepared initial data, the energy of the limiting consensus state can be made arbitrarily close to the global minimum over the active subspace. A key feature of the Hilbert-space formulation is that the stochastic contribution to the convergence estimates is controlled by the trace of the covariance operator and is therefore uniform with respect to the Galerkin dimension. Numerical experiments on an elliptic energy minimization problem with mixed boundary conditions and a PDE-constrained inverse source problem validate the theoretical convergence results and demonstrate stable performance under increasing spatial resolution, in contrast with CBO based on isotropic finite-dimensional noise.

math.OC

A PDE-Based Framework for Generative Modeling Beyond Classical Score-Based Diffusion

We introduce an alternative generative framework based on a nonlinear modification of the classical Ornstein--Uhlenbeck dynamics. The proposed dynamics admits both a microscopic description through an interacting particle system and, in the mean-field limit, a macroscopic formulation given by a nonlinear Fokker--Planck equation with a superlinear drift term. We show that, for suitable choices of the model parameters and sufficiently large initial mass, the forward dynamics exhibits condensation phenomena by proving the loss of $L^2$ regularity of the solution in finite time. Building upon this formulation, we derive a stabilized reverse-time partial differential equation that reconstructs the initial distribution from the asymptotic state of the forward dynamics, thereby extending the generative paradigm beyond the classical score-based framework. Furthermore, we introduce numerical discretizations of both the forward and reverse processes that accurately capture the asymptotic behavior of the continuous model while successfully reconstructing the initial distribution. Numerical experiments in one and two spatial dimensions validate the proposed methodology and illustrate its application to density filtering through successive iterations of the generative process.

math.NA

Ensemble Kalman Inversion as an Inertial Interacting Particle System

Ensemble Kalman Inversion (EKI) is a derivative-free, ensemble-based method for inverse and optimization problems. Its continuous-time formulation can be interpreted as an interacting particle system driven by a Kalman-type preconditioned descent direction. A well-known limitation of this dynamics is the possible premature collapse of the covariance of the ensemble, which makes the method sensitive to the initial ensemble. We introduce a second-order particle system in which the particles evolve according to an inertial dynamics. The model combines a Kalman-type relaxation force with damping, attraction towards the ensemble mean, and a short-range repulsive interaction designed to counteract ensemble collapse. The resulting dynamics can be interpreted as a heavy-ball reformulation of continuous-time EKI enriched by competing attractive and repulsive mechanisms. For linear inverse problems, we analyze the induced mean and fluctuation dynamics and identify a parameter regime in which fully collapsed configurations are linearly unstable. We further characterize asymptotic equilibria through a constrained optimality condition on the subspace retained by the limiting ensemble covariance and derive an exponential decay estimate. Numerical experiments illustrate the effect of inertia and repulsion on the ensemble dynamics and compare the proposed second-order method with first-order EKI-type

math.NA

Sparse and low-rank kinetic distribution estimation

In this paper, we consider methods that allow for memory-efficient storage of high-dimensional distributions and retain certain key features thereof, specifically in a kinetic theory context. We propose an extension to the entropic quadrature method that allows for enforcing sparsity, and propose a new low-rank decomposition approach that ensures preservation of moment information. The methods are applied to several model kinetic distributions, as well as to distributions obtained from high-resolution kinetic simulations of the Vlasov--Maxwell system.

physics.comp-ph

Computation of entanglement for quantum states by a Consensus-Based Optimization method

The computation of quantum entanglement can be formulated as a high-dimensional nonconvex optimization problem with orthogonality constraints. In this work, we propose structure-preserving consensus-based optimization (CBO) methods for entanglement computation, with one approach based on a Hermitian formulation and the other evolving directly on the unitary manifold. To handle the variable dimension of the feasible set, we introduce a cross-dimensional interaction mechanism allowing exchange of information between particles of different sizes. Numerical experiments demonstrate that the proposed methods achieve accurate approximations.

quant-ph

New Scheme Adaption Strategy for Hyperbolic Conservation Laws

We introduce a new scheme adaption strategy for one- and two-dimensional hyperbolic systems of conservation laws. The proposed approach builds upon the adaptive framework introduced in [S. Chu, A. Kurganov, and I. Menshov, Appl. Numer. Math., 209 (2025), pp.155--170], where we first employed the smoothness indicator from [R. Lohner, Comput. Methods. Appl. Mech. Eng., 61 (1987), pp.323--338] to automatically detect ``rough'' and smooth parts of the computed solution, and then used different limiters in the detected regions. This adaptive strategy was based on a threshold needed to sharply separate ``rough'' and smooth regions. In this paper, we propose a different adaption strategy. We use SBM-type limiters and vary one of the limiting parameters continuously to allow a smooth transition between the ``rough'' and smooth areas. This way, compressive and overcompressive limiters are activated in the shock and contact wave vicinities only, while we gradually switch to dissipative limiters in the smooth regions. A series of one- and two-dimensional numerical tests for the Euler equations of gas dynamics demonstrates that the new scheme adaption strategy leads to a higher resolution and reduced numerical dissipation.

math.NA

Consensus-based optimization with $α$-stable jump processes

In this paper, we introduce a novel variant of the CBO method that incorporates jumps according to an $α$-stable stochastic process in a kinetic framework. This extension gives rise to nonlocal stochastic effects, which improve the exploration capabilities of the method. We formulate the method at the particle level, detailing the corresponding stochastic dynamics and its asymptotic behavior. In particular, through a Fourier-based representation, we derive the associated fractional Fokker-Planck equation, which naturally accounts for the nonlocal diffusion behaviors induced by $α$-stable processes. As a central result, we establish a rigorous convergence result for the proposed approach. Finally, we evaluate the performance of the method through a set of numerical experiments. The results demonstrate the effectiveness of the $α$-stable jump process and emphasize its potential advantages over standard diffusion-based methods, particularly in complex optimization settings.

math.OC

Kinetic simulated annealing optimization with entropy-based cooling rate

We present a modified simulated annealing method with a dynamical choice of the cooling temperature. The latter is determined via a closed-loop control and is proven to yield exponential decay of the entropy of the particle system. The analysis is carried out through kinetic equations for interacting particle systems describing the simulated annealing method in an extended phase space. Decay estimates are derived under the quasi-invariant scaling of the resulting system of Boltzmann-type equations to assess the consistency with their mean-field limit. Numerical results are provided to illustrate and support the theoretical findings.

math.OC

A Comparative Study of Low-Dissipation Numerical Schemes for Hyperbolic Conservation Laws

This work provides a comparative assessment of several low-dissipation numerical schemes for hyperbolic conservation laws, highlighting their performance relative to the classical Harten-Lax-van Leer (HLL) schemes. The schemes under consideration include the classical Harten-Lax-van Leer-Contact (HLLC), the recently proposed TV flux splitting, the low-dissipation Central-Upwind (LDCU), and the local characteristic decomposition-based Central-Upwind (LCDCU) schemes. These methods are extended to higher orders of accuracy, up to the fifth order, within both finite-volume and finite-difference frameworks. A series of numerical experiments for the one- and two-dimensional Euler equations of gas dynamics are performed to evaluate the accuracy, robustness, and computational efficiency of the studied schemes. The comparison highlights the trade-offs between resolution of contact and shear waves, robustness in the presence of shocks, and computational cost. The investigated low-dissipation schemes show comparable levels of numerical dissipation, with only subtle differences appearing in selected benchmark problems. The results provide practical guidance for selecting efficient low-dissipation solvers for the simulation of complex compressible flows.

math.NA

A Bayesian Approach to Feedback Control for Hyperbolic Balance Laws

We propose a Bayesian framework for feedback boundary control of hyperbolic balance laws. The method propagates a probability distribution over feedback parameters using Lyapunov decay estimates as a likelihood. For linear models, it recovers available analytical stability results and extends to nonlinear regimes where theory is limited. Using first-order local Lax-Friedrichs (LLF) discretizations, we validate the approach on the decoupled wave system and the linearized Saint-Venant equations, reproducing known stability intervals and mixed boundary couplings. We then treat nonlinear and stochastic problems, including the nonlinear Saint-Venant system, one- and two-dimensional Burgers equations, Burgers equation with random initial data, and nonconservative perturbations with source terms, and show that the inferred stability domains are robust with respect to the indicator and the prior. Finally, we demonstrate transfer to a second-order semi-discrete LLF scheme and to a two-parameter feedback model for laser powder bed fusion with power regulation.

math.NA

Numerical Study of Dissipative Weak Solutions for the Euler Equations of Gas Dynamics

We study dissipative weak (DW) solutions of the Euler equations of gas dynamics using the first-, second-, third-, fifth-, seventh-, and ninth-order local characteristic decomposition-based central-upwind (LCDCU), low-dissipation central-upwind (LDCU), and viscous finite volume (VFV) methods, whose higher-order extensions are obtained via the framework of the alternative weighted essentially non-oscillatory (A-WENO) schemes. These methods are applied to several benchmark problems, including several two-dimensional Riemann problems and a Kelvin-Helmholtz instability test. The numerical results demonstrate that for methods converging only weakly in space and time, the limiting solutions are generalized DW solutions, approximated in the sense of ${\cal K}$-convergence and dependent on the numerical scheme. For all of the studied methods, we compute the associated Young measures and compare the DW solutions using entropy production and energy defect criteria.

math.NA

High-dimensional stochastic finite volumes using the tensor train format

A method for the uncertainty quantification of nonlinear hyperbolic conservation laws with many uncertain parameters is presented. The method combines stochastic finite volume methods and tensor trains in a novel way: the dimensions of physical space and time are kept as full tensors, while all stochastic dimensions are compressed together into a tensor train. The resulting hybrid format has one tensor train for each spatial cell and each time step. The MUSCL scheme is adapted to the proposed hybrid format, and its feasibility is demonstrated through several test cases. For the scalar Burgers' equation, we conduct a convergence study and compare the results with those obtained using the full tensor train format with three stochastic parameters. The equation is then solved for an increasing number of stochastic dimensions.For systems of conservation laws, we focus on the Euler equations. A parameter study and a comparison with the full tensor train format are carried out for the Sod shock tube problem. As a more complex application, we investigate the Shu-Osher problem, which involves intricate wave interactions. The presented method opens new avenues for integrating uncertainty quantification with established numerical schemes for hyperbolic conservation laws.

math.NA

Mean-Field Model for Two-Layer Neural Networks Trained with Consensus-Based Optimization

We study Consensus-Based Optimization (CBO) for two-layer neural network training. We compare the performance of CBO against Adam on two test cases and demonstrate how a hybrid approach, combining CBO with Adam, provides faster convergence than CBO. Additionally, in the context of multi-task learning, we recast CBO into a formulation that offers less memory overhead. The CBO method allows for a mean-field model formulation, which we couple with the mean-field model of the neural network. To this end, we first reformulate CBO within the optimal transport framework. As the number of particles tends to infinity, we lift the corresponding dynamics to the Wasserstein-over-Wasserstein space and show that the variance decreases monotonically. We confirm numerically that both mean-field models converge.

cs.LG

Consensus-based algorithms for stochastic optimization problems

We address an optimization problem where the cost function is the expectation of a random mapping. To tackle the problem two approaches based on the approximation of the objective function by consensus-based particle optimization methods on the search space are developed. The resulting methods are mathematically analyzed using a mean-field approximation and their connection is established. Several numerical experiments show the validity of the proposed algorithms and investigate their rates of convergence.

math.OC

Solving Random Hyperbolic Conservation Laws Using Linear Programming

A novel structure-preserving numerical method to solve random hyperbolic systems of conservation laws is presented. The method uses a concept of generalized, measure-valued solutions to random conservation laws. This yields a linear partial differential equation with respect to the Young measure and allows to compute the approximation based on linear programming problems. We analyze structure-preserving properties of the derived numerical method and discuss its advantages and disadvantages. We numerically demonstrate the approach on the one-dimensional Burgers and isentropic Euler equations and compare with stochastic collocation. In addition, we introduce a discontinuous-flux test in which different entropies used in the linear-program objective select different weak entropy solutions, and we report the corresponding changes in the moments and supports of the Young measure.

math.NA

A modified particle filter that reduces weight collapse

Particle filters are a widely used Monte Carlo based data assimilation technique that estimates the probability distribution of a system's state conditioned on observations through a collection of weights and particles. A known problem for particle filters is weight collapse, or degeneracy, where a single weight attains a value of one while all others are close to zero, thereby collapsing the estimated distribution. We address this issue by introducing a novel modification to the particle filter that is simple to implement and inspired by energy-based diversity measures. Our approach adjusts particle weights to minimize a two-body energy potential, promoting balanced weight distributions and mitigating collapse. We demonstrate the performance of this modified particle filter in a series of numerical experiments with linear and nonlinear dynamical models, where we compare with the classical particle filter and ensemble Kalman filters in the nonlinear case. We find that our new approach improves weight distributions compared to the classical particle filter and thereby improve state estimates.

stat.AP

Turnpike property for hierarchical optimal control problems: from particle systems to hydrodynamic equations

This work is concerned with a hierarchical framework of optimal control problems connecting interacting particle systems, the mean field limit equations, and associated hydrodynamic models. By assuming the existence of solutions, we establish the exponential turnpike property for each level of the hierarchy, showing that optimal trajectories remain close to the associated steady states over long time horizons. The results demonstrate that the exponential turnpike behavior persists consistently across scales, providing a unified connection between microscopic, kinetic, and macroscopic optimal control frameworks.

math.OC