SearcharxivSearch

arXiv subjects

Arnaud Guillin

Publications and source records attributed to Arnaud Guillin.

At least 19 recordsLinked to original sources

Quantifying Uncertainty In Wide Two-Layer Neural Networks: On The Law Of The Limiting Fluctuation Process

Uncertainty quantification in neural networks prediction is a main issue for usual applications. Our approach seeks at reducing computation costs by directly evaluating uncertainty using PDE's information on the asymptotic variance, rather than the deep ensemble method which may be seen as a Monte Carlo estimation of the prediction, requiring the training of multiple networks. We thus study the law of the limiting process describing the random fluctuations around the mean-field limit of wide two-layer neural networks trained by stochastic gradient descent in a weak-noise regime. Building on a recent trajectorial central limit theorem, in which this limit is characterized as the weak solution of a linear stochastic evolution equation, we identify its law explicitly. More precisely, we show that it is a centered Gaussian process in the dual of a weighted Sobolev space, and we derive a closed covariance representation for the finite-dimensional distributions obtained by testing it against smooth functions. This covariance is expressed through the solution of a backward transport equation with a nonlocal source term, whose coefficients are driven by the mean-field trajectory. As a consequence, by testing against the activation function at a fixed input, we obtain an expression for the limiting variance of the corresponding network-output fluctuations. We illustrate this result numerically on a one-dimensional regression example.

cs.NE

Weak Functional Inequalities for Perturbed Measures

This paper is a follow up to an article by two of the authors dedicated to the study of Poincar\'e and logarithmic Sobolev inequalities for measures of the form $d\mu = e^{-U} d\nu$ where $e^{-U}$ is seen as a perturbation of $d\nu$. Application to the same functional inequalities for convolution products are then discussed. In the present paper we investigate similar problems for weaker functional inequalities, namely weak Poincar\'e, weighted Poincar\'e, weak log-Sobolev and weighted log-Sobolev inequalities.

math.PR

Uniform-in-time concentration in two-layer neural networks via transportation inequalities

We quantify, uniformly over time and with high probability, the discrepancy between the predictions of a two-layer neural network trained by stochastic gradient descent (SGD) and their mean-field limit, for quadratic loss and ridge regularization. As a key ingredient, we establish T p transportation inequalities (p $\in$ {1, 2}) for the law of the SGD parameters, with explicit constants independent of the iteration index. We then prove uniform-in-time concentration of the empirical parameter measure around its mean-field limit in the Wasserstein distance W 1 , and we translate these bounds into prediction-error estimates against a fixed test function $\Phi$. We also derive analogous concentration bounds in the sliced-Wasserstein distance SW 1 , leading to dimension-free rates.

cs.NE

Propagation of chaos in Fisher information

We present a new method for proving sharp local propagation of chaos in Fisher Information for particles with smooth interaction and drift. We rely on a new Lemma computing the Fisher Information of two diffusion processes with smooth drifts and fine estimates on the hessian of the law of the solution of the McKean-Vlasov equation. It allows us to obtain a new propagation of chaos in Fisher information, generalizing Lacker's seminal work by using the BBGKY hierarchy to obtain a system of differential inequalities satisfied by both the relative entropy and the Fisher Information of k particles. We also show with a simple Gaussian example that our decay rate is optimal.

math.PR

Diffusion annealed Langevin dynamics: a theoretical study

In this work we study the diffusion annealed Langevin dynamics, a score-based diffusion process recently introduced in the theory of generative models and which is an alternative to the classical overdamped Langevin diffusion. Our goal is to provide a rigorous construction and to study the theoretical efficiency of these models for general base distribution as well as target distribution. As a matter of fact these diffusion processes are a particular case of Nelson processes i.e. diffusion processes with a given flow of time marginals. Providing existence and uniqueness of the solution to the annealed Langevin diffusion leads to proving a Poincar\'e inequality for the conditional distribution of $X$ knowing $X+Z=y$ uniformly in $y$, as recently observed by one of us and her coauthors. Part of this work is thus devoted to the study of such Poincar\'e inequalities. Additionally we show that strengthening the Poincar\'e inequality into a logarithmic Sobolev inequality improves the efficiency of the model.

math.PR

Activity-driven clustering and many-body steady state of jamming run-and-tumble particles

We exactly resolve the three-particle steady state of run-and-tumble particles with jamming interactions, providing the first microscopic description beyond two bodies. The invariant measure, derived via a piecewise-deterministic Markov process description and symmetry principles, reveals persistent, separated, and diffusive regimes ruled by the activity parameter. A geometric cascade of scales in the activity parameter organizes the structural weights, showing the separated phase dominates at finite activity, while non-uniformity plays only a minor role. Extending these results to larger systems, we show that the $N$-body steady state inherits the same organization: the number of clusters becomes sharply defined by the activity value, with crossover boundaries whose slopes diverge with $N$. We also show how the activity plays a role similar to a fugacity conjugate to cluster number, yielding a grand-canonical-like structure emerging directly from the microscopic dynamics. This framework lays the groundwork for a systematic microscopic theory of active many-body steady states.

cond-mat.stat-mech

A new machine learning framework for occupational accidents forecasting with safety inspections integration

We propose a model-agnostic framework for short-term occupational accident forecasting that leverages safety inspections and models accident occurrences as binary time series. The approach generates daily predictions, which are then aggregated into weekly safety assessments for better decision making. To ensure the reliability and operational applicability of the forecasts, we apply a sliding-window cross-validation procedure specifically designed for time series data, combined with an evaluation based on aggregated period-level metrics. Several machine learning algorithms, including logistic regression, tree-based models, and neural networks, are trained and systematically compared within this framework. Across all tested algorithms, the proposed framework reliably identifies upcoming high-risk periods and delivers robust period-level performance, demonstrating that converting safety inspections into binary time series yields actionable, short-term risk signals. The proposed methodology converts routine safety inspection data into clear weekly and daily risk scores, detecting the periods when accidents are most likely to occur. Decision-makers can integrate these scores into their planning tools to classify inspection priorities, schedule targeted interventions, and funnel resources to the sites or shifts classified as highest risk, stepping in before incidents occur and getting the greatest return on safety investments.

cs.LG

Quasi-stationarity of the Dyson Brownian motion with collisions

In this work, we investigate the ergodic behavior of a system of particules, subject to collisions, before it exits a fixed subdomain of its state space. This system is composed of several one-dimensional ordered Brownian particules in interaction with electrostatic repulsions, which is usually referred as the (generalized) Dyson Brownian motion. The starting points of our analysis are the work [E. C{\'e}pa and D. L{\'e}pingle, 1997 Probab. Theory Relat. Fields] which provides existence and uniqueness of such a system subject to collisions via the theory of multivalued SDEs and a Krein-Rutman type theorem derived in [A. Guillin, B. Nectoux, L. Wu, 2020 J. Eur. Math. Soc.].

math.PR

Convergence of non-reversible Markov processes via lifting and flow Poincar{\'e} inequality

We propose a general approach for quantitative convergence analysis of non-reversible Markov processes, based on the concept of second-order lifts and a variational approach to hypocoercivity. To this end, we introduce the flow Poincar{\'e} inequality, a space-time Poincar{\'e} inequality along trajectories of the semigroup, and a general divergence lemma based only on the Dirichlet form of an underlying reversible diffusion. We demonstrate the versatility of our approach by applying it to a pair of run-and-tumble particles with jamming, a model from non-equilibrium statistical mechanics, and several piecewise deterministic Markov processes used in sampling applications, in particular including general stochastic jump kernels.

math.AP

Large deviations of the empirical measures of a strong-Feller Markov process inside a subset and quasi-ergodic distribution

In this work, we establish, for a strong Feller process, the large deviation principle for the occupation measure conditioned not to exit a given subregion. The rate function vanishes only at a unique measure, which is the so-called quasi-ergodic distribution of the process in this subregion. In addition, we show that the rate function is the Dirichlet form in the particular case when the process is reversible. We apply our results to several stochastic processes such as the solutions of elliptic stochastic differential equations driven by a rotationally invariant $\alpha$-stable process, the kinetic Langevin process, and the overdamped Langevin process driven by a Brownian motion.

math.PR

Long-time analysis of a pair of on-lattice and continuous run-and-tumble particles with jamming interactions

Run-and-Tumble Particles (RTPs) are a key model of active matter. They are characterized by alternating phases of linear travel and random direction reshuffling. By this dynamic behavior, they break time reversibility and energy conservation at the microscopic level. It leads to complex out-of-equilibrium phenomena such as collective motion, pattern formation, and motility-induced phase separation (MIPS). In this work, we study two fundamental dynamical models of a pair of RTPs with jamming interactions and provide a rigorous link between their discrete- and continuous-space descriptions. We demonstrate that as the lattice spacing vanishes, the discrete models converge to a continuous RTP model on the torus, described by a Piecewise Deterministic Markov Process (PDMP). This establishes that the invariant measures of the discrete models converge to that of the continuous model, which reveals finite mass at jamming configurations and exponential decay away from them. This indicates effective attraction, which is consistent with MIPS. Furthermore, we quantitatively explore the convergence towards the invariant measure. Such convergence study is critical for understanding and characterizing how MIPS emerges over time. Because RTP systems are non-reversible, usual methods may fail or are limited to qualitative results. Instead, we adopt a coupling approach to obtain more accurate, non-asymptotic bounds on mixing times. The findings thus provide deeper theoretical insights into the mixing times of these RTP systems, revealing the presence of both persistent and diffusive regimes.

math.PR

Long time behavior of killed Feynman-Kac semigroups with singular Schr{\"o}dinger potentials

In this work, we investigate the compactness and the long time behavior of killed Feynman-Kac semigroups of various processes arising from statistical physics with very general singular Schr{\"o}dinger potentials. The processes we consider cover a large class of processes used in statistical physics, with strong links with quantum mechanics and (local or not) Schr{\"o}dinger operators (including e.g. fractional Laplacians). For instance we consider solutions to elliptic differential equations, L{\'e}vy processes, the kinetic Langevin process with locally Lipschitz gradient fields, and systems of interacting L{\'e}vy particles. Our analysis relies on a Perron-Frobenius type theorem derived in a previous work [A. Guillin, B. Nectoux, L. Wu, 2020 J. Eur. Math. Soc.] for Feller kernels and on the tools introduced in [L. Wu, 2004, Probab. Theory Relat. Fields] to compute bounds on the essential spectral radius of a bounded nonnegative kernel.

math.PR

Sharp propagation of chaos for McKean-Vlasov equation with non constant diffusion coefficient

We present a method to obtain sharp local propagation of chaos results for a system of N particles with a diffusion coefficient that it not constant and may depend of the empirical measure. This extends the recent works of Lacker [14] and Wang [24] to the case of non constant diffusions. The proof relies on the BBGKY hierarchy to obtain a system of differential inequalities on the relative entropy of k particles, involving the fisher information.

math.PR

Error estimates between SGD with momentum and underdamped Langevin diffusion

Stochastic gradient descent with momentum is a popular variant of stochastic gradient descent, which has recently been reported to have a close relationship with the underdamped Langevin diffusion. In this paper, we establish a quantitative error estimate between them in the 1-Wasserstein and total variation distances.

stat.ML

Optimization of Complex Process, Based on Design Of Experiments, a Generic Methodology

MicroLED displays are the result of a complex manufacturing chain. Each stage of this process, if optimized, contributes to achieving the highest levels of final efficiencies. Common works carried out by Pollen Metrology, Aledia, and Universit{\'e} Clermont-Auvergne led to a generic process optimization workflow. This software solution offers a holistic approach where stages are chained together for gaining a complete optimal solution. This paper highlights key corners of the methodology, validated by the experiments and process experts: data cleaning and multi-objective optimization.

cs.NE

Central Limit Theorem for Bayesian Neural Network trained with Variational Inference

In this paper, we rigorously derive Central Limit Theorems (CLT) for Bayesian two-layerneural networks in the infinite-width limit and trained by variational inference on a regression task. The different networks are trained via different maximization schemes of the regularized evidence lower bound: (i) the idealized case with exact estimation of a multiple Gaussian integral from the reparametrization trick, (ii) a minibatch scheme using Monte Carlo sampling, commonly known as Bayes-by-Backprop, and (iii) a computationally cheaper algorithm named Minimal VI. The latter was recently introduced by leveraging the information obtained at the level of the mean-field limit. Laws of large numbers are already rigorously proven for the three schemes that admits the same asymptotic limit. By deriving CLT, this work shows that the idealized and Bayes-by-Backprop schemes have similar fluctuation behavior, that is different from the Minimal VI one. Numerical experiments then illustrate that the Minimal VI scheme is still more efficient, in spite of bigger variances, thanks to its important gain in computational complexity.

stat.ML

Generalized Langevin And Nos{\'e}-hoover Processes Absorbed At The Boundary Of A Metastable Domain

In this paper, we prove in a very weak regularity setting existence and uniqueness of quasi-stationary distributions as well as exponential conver- gence towards the quasi-stationary distribution for the generalized Langevin and the Nos{\'e}-Hoover processes, two processes which are widely used in molecular dynamics. The case of singular potentials is considered. With the techniques used in this work, we are also able to greatly improve existing results on quasi-stationary distributions for the kinetic Langevin process to a weak regularity setting.

math.PR

Discrete sticky couplings of functional autoregressive processes

In this paper, we provide bounds in Wasserstein and total variation distances between the distributions of the successive iterates of two functional autoregressive processes with isotropic Gaussian noise of the form $Y_{k+1} = \mathrm{T}_γ(Y_k) + \sqrt{γσ^2} Z_{k+1}$ and $\tilde{Y}_{k+1} = \tilde{\mathrm{T}}_γ(\tilde{Y}_k) + \sqrt{γσ^2} \tilde{Z}_{k+1}$. More precisely, we give non-asymptotic bounds on $ρ(\mathcal{L}(Y_{k}),\mathcal{L}(\tilde{Y}_k))$, where $ρ$ is an appropriate weighted Wasserstein distance or a $V$-distance, uniformly in the parameter $γ$, and on $ρ(π_γ,\tildeπ_γ)$, where $π_γ$ and $\tildeπ_γ$ are the respective stationary measures of the two processes. The class of considered processes encompasses the Euler-Maruyama discretization of Langevin diffusions and its variants. The bounds we derive are of order $γ$ as $γ\to 0$. To obtain our results, we rely on the construction of a discrete sticky Markov chain $(W_k^{(γ)})_{k \in \mathbb{N}}$ which bounds the distance between an appropriate coupling of the two processes. We then establish stability and quantitative convergence results for this process uniformly on $γ$. In addition, we show that it converges in distribution to the continuous sticky process studied in previous work. Finally, we apply our result to Bayesian inference of ODE parameters and numerically illustrate them on two particular problems.

math.PR