SearcharxivSearch

arXiv subjects

Konstantinos Spiliopoulos

Publications and source records attributed to Konstantinos Spiliopoulos.

At least 19 recordsLinked to original sources

Quantitative Fluctuation Analysis for Continuous-Time Stochastic Gradient Descent via Malliavin Calculus

In this paper, we establish a Quantitative Central Limit Theorem (QCLT) for the Stochastic Gradient Descent in Continuous Time (SGDCT) algorithm, whose parameter updates are governed by a stochastic differential equation. We derive an explicit rate at which the SGDCT iterates converge, in the Wasserstein metric, to a critical point of the objective function. This rate is driven primarily by the magnitude of the learning rate: for a fixed convexity constant of the objective function, smaller learning rates lead to slower convergence. Our approach relies on tools from Malliavin calculus. In particular, we apply a second-order Poincaré inequality and obtain explicit bounds by estimating the first- and second-order Malliavin derivatives separately. Controlling the second-order derivative requires several delicate calculations and a careful sequence of decompositions in order to achieve sharp estimates. We complement the theoretical results with several numerical experiments that illustrate the predicted convergence behavior.

math.PR

Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs

The Deep Galerkin Method (DGM) and Physics Informed Neural Networks (PINNs) have become widely-used methods for solving partial differential equations (PDEs) in the rapidly growing field of scientific machine learning. In these methods, a neural network is trained to approximate the PDE solution by using (stochastic) gradient descent to minimize the PDE residual of the neural network. Due to the non-convexity of the PDE residual objective function, the trained neural network may, in principle, only converge to a local minimizer of the objective function (which would not be a solution of the PDE). Therefore, there is a longstanding question regarding the mathematical foundations of these algorithms, and it is highly valuable to establish that the trained neural network will converge to the PDE solution. In this paper, we consider a class of semilinear PDEs with nonlinearities in the solution and its first derivative. For this class of PDEs, we prove that neural networks trained with gradient descent to minimize the PDE residual objective function will converge to the PDE solution as the network width and training time $\rightarrow \infty$.

cs.LG

Optimizing Irreversible Perturbations of the Unadjusted Langevin Algorithm

Irreversible perturbations accelerate the convergence of Langevin dynamics, breaking detailed balance while preserving the invariant measure. The design of optimal irreversible perturbations has been studied in the continuous-time Gaussian setting, but extensions to non-Gaussian target distributions, and the impact of time discretization on the design of optimal perturbations, have not been well understood. Numerical discretizations of Langevin dynamics introduce bias, which is typically exacerbated by irreversible perturbations; handling this interaction demands a joint treatment of acceleration and accuracy. This paper develops a systematic framework for optimizing position-independent irreversible perturbations of the unadjusted Langevin algorithm (ULA). We formulate a constrained optimization problem that simultaneously accounts for mixing efficiency and discretization bias, where the former is characterized by a spectral gap analogue and the latter is quantified via a weighted expected squared jump distance. Within this framework, we derive an explicit characterization of the optimal position-independent irreversible perturbation. Extensive numerical experiments demonstrate that our design yields faster convergence with controlled bias, and improves mean squared estimation errors compared to other choices of irreversible perturbation.

math.NA

Mean Field Limits for Stochastic, Underdamped Reactive Langevin Dynamics Models

We rigorously derive the effective large-population, mean-field dynamics of particle-based reactive Langevin dynamics (PBRLD) models. These models extend particle-based stochastic reaction-diffusion (PBSRD) descriptions by incorporating velocities, inertial effects, and underdamped motion. In Isaacson, Liu, Spiliopoulos, and Yao, SIAP 2026, PBRLD models were formulated and shown to recover Doi volume reactivity PBSRD model in the overdamped limit. In this work we prove convergence of the associated measure-valued stochastic processes, representing species concentration fields on position-velocity phase space, to a deterministic mean-field limit. The limiting equations form a novel system of nonlocal kinetic reaction-diffusion partial integro-differential equations, coupling hypoelliptic transport with reaction terms that retain the spatial and velocity structure of the underlying particle interactions.

math.PR

Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit

A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. As the number of hidden units tends to infinity, the NN training dynamics converge in probability to the solution of a deterministic limit equation involving a ``Newton neural tangent kernel'' (NNTK). Explicit rates characterizing this convergence are provided and, in the infinite-width limit, we prove that the NN converges exponentially fast to the target data (i.e., a global minimizer with zero loss). We show that this convergence is uniform across the frequency spectrum, addressing the spectral bias inherent in gradient descent. The eigenvalues of the NTK for gradient descent accumulate at zero, leading to slow convergence for target data with high-frequency components. In contrast, the NNTK has uniformly lower bounded eigenvalues if the regularization parameter is selected appropriately, allowing Newton's method to converge more quickly for data with high-frequency components. Mathematical challenges that need to be addressed in our analysis include the implicit parameter update of the Newton method with a potentially indefinite Hessian matrix and the fact that the dimension of this linear system of equations tends to infinity as the NN width grows. This complicates deriving the training dynamics in the overparameterized limit as well as proving the convergence of the finite-width dynamics thereto. The analysis identifies a scaling formula for selecting the regularization parameter, which we show can vanish at a suitable rate as the number of hidden units becomes larger. We prove that, for sufficiently large numbers of hidden units, the regularized Hessian remains positive definite during training and the Newton updates for individual NN parameters converge to zero, showing that the model behaves as a linearization around the initialization.

cs.LG

Uniform-in-time quantitative fluctuations of large scale interacting particle systems

We study fluctuations of mean-field interacting particle systems around their McKean--Vlasov limit. Our main result provides a uniform-in-time quantitative central limit theorem for the fluctuation process, with convergence rate of order $N^{-1/2}$ to the corresponding Gaussian limit in the Wasserstein metric. The proof relies on two main ingredients. First, we establish a uniform-in-time weak expansion for specific functionals of the empirical measure around their limiting behavior. This yields, in particular, uniform-in-time control of the convergence of the prelimit variance to its limiting counterpart. We also derive a backward PDE representation of the limiting variance, which is of independent interest. Second, we use Malliavin calculus tools and, in particular, a second-order Poincaré inequality that bounds the Wasserstein distance between the fluctuation process and its Gaussian limit in terms of the first- and second-order Malliavin derivatives of the particle flow. The quantitative convergence rates then follow from a delicate analysis of these derivatives, yielding the sharp estimates required for uniform-in-time control.

math.PR

Mean-Field Analysis of Latent Variable Process Models on Dynamically Evolving Graphs with Feedback Effects

We study the mean-field limit of a generic class of dynamic co-evolving latent space networks motivated by the social and opinion dynamics literature. Such models include $n$ agents, whose opinions are given by latent stochastic processes, and a dynamic network process describing agent interactions. Models in this class incorporate (a) bi-directional feedback between the latent processes and the network process, (b) persistence effects, meaning that the network structure at the current time depends on the value of the latent processes at the current time but also on the network structure at the previous time instance and (c) localized interactions, meaning that individual agents do not have global information. We characterize the distributional limit of a random sample taken from the latent space network as the number of nodes in the network diverges. We describe the rich conditional probabilistic structure of the resulting limiting model which we use to establish the limiting behavior of the following quantities: (i) the empirical measure of the latent process, (ii) a conditional empirical measure relating the latent process to the network process and (iii) the network process graphon. In proving our main results, we derive a general conditional propagation of chaos result, which is of independent interest. Our novel approach to studying the limiting behavior of random samples proves to be a very useful methodology for fully grasping the asymptotic behavior of co-evolving particle systems. Numerical results are included to illustrate the theoretical findings.

math.PR

Scaling Effects and Uncertainty Quantification in Neural Actor Critic Algorithms

We investigate the neural Actor Critic algorithm using shallow neural networks for both the Actor and Critic models. The focus of this work is twofold: first, to compare the convergence properties of the network outputs under various scaling schemes as the network width and the number of training steps tend to infinity; and second, to provide precise control of the approximation error associated with each scaling regime. Previous work has shown convergence to ordinary differential equations with random initial conditions under inverse square root scaling in the network width. In this work, we shift the focus from convergence speed alone to a more comprehensive statistical characterization of the algorithm's output, with the goal of quantifying uncertainty in neural Actor Critic methods. Specifically, we study a general inverse polynomial scaling in the network width, with an exponent treated as a tunable hyperparameter taking values strictly between one half and one. We derive an asymptotic expansion of the network outputs, interpreted as statistical estimators, in order to clarify their structure. To leading order, we show that the variance decays as a power of the network width, with an exponent equal to one half minus the scaling parameter, implying improved statistical robustness as the scaling parameter approaches one. Numerical experiments support this behavior and further suggest faster convergence for this choice of scaling. Finally, our analysis yields concrete guidelines for selecting algorithmic hyperparameters, including learning rates and exploration rates, as functions of the network width and the scaling parameter, ensuring provably favorable statistical behavior.

cs.LG

Kernel Limit for a Class of Recurrent Neural Networks Trained on Ergodic Data Sequences

Mathematical methods are developed to characterize the asymptotics of recurrent neural networks (RNN) as the number of hidden units, data samples in the sequence, hidden state updates, and training steps simultaneously grow to infinity. In the case of an RNN with a simplified weight matrix, we prove the convergence of the RNN to the solution of an infinite-dimensional ODE coupled with the fixed point of a random algebraic equation. The analysis requires addressing several challenges which are unique to RNNs. In typical mean-field applications (e.g., feedforward neural networks), discrete updates are of magnitude $\mathcal{O}(1/N)$ and the number of updates is $\mathcal{O}(N)$. Therefore, the system can be represented as an Euler approximation of an appropriate ODE/PDE, which it will converge to as $N \rightarrow \infty$. However, the RNN hidden layer updates are $\mathcal{O}(1)$. Therefore, RNNs cannot be represented as a discretization of an ODE/PDE and standard mean-field techniques cannot be applied. Instead, we develop a fixed point analysis for the evolution of the RNN memory states, with convergence estimates in terms of the number of update steps and the number of hidden units. The RNN hidden layer is studied as a function in a Sobolev space, whose evolution is governed by the data sequence (a Markov chain), the parameter updates, and its dependence on the RNN hidden layer at the previous time step. Due to the strong correlation between updates, a Poisson equation must be used to bound the fluctuations of the RNN around its limit equation. These mathematical methods give rise to the neural tangent kernel (NTK) limits for RNNs trained on data sequences as the number of data samples and size of the neural network grow to infinity.

cs.LG

Global Convergence of Adjoint-Optimized Neural PDEs

Many engineering and scientific fields have recently become interested in modeling terms in partial differential equations (PDEs) with neural networks, which requires solving the inverse problem of learning neural network terms from observed data in order to approximate missing or unresolved physics in the PDE model. The resulting neural-network PDE model, being a function of the neural network parameters, can be calibrated to the available ground truth data by optimizing over the PDE using gradient descent, where the gradient is evaluated in a computationally efficient manner by solving an adjoint PDE. These neural PDE models have emerged as an important research area in scientific machine learning. In this paper, we study the convergence of the adjoint gradient descent optimization method for training neural PDE models in the limit where both the number of hidden units and the training time tend to infinity. Specifically, for a general class of nonlinear parabolic PDEs with a neural network embedded in the source term, we prove convergence of the trained neural-network PDE solution to the target data (i.e., a global minimizer). The global convergence proof poses a unique mathematical challenge that is not encountered in finite-dimensional neural network convergence analyses due to (i) the neural network training dynamics involving a non-local neural network kernel operator in the infinite-width hidden layer limit where the kernel lacks a spectral gap for its eigenvalues and (ii) the nonlinearity of the limit PDE system, which leads to a non-convex optimization problem in the neural network function even in the infinite-width hidden layer limit (unlike in typical neural network training cases where the optimization problem becomes convex in the large neuron limit). The theoretical results are illustrated and empirically validated by numerical studies.

cs.LG

On the large-time behaviour of affine Volterra processes

We show the existence of a stationary measure for a class of multidimensional stochastic Volterra systems of affine type. These processes are in general not Markovian, a shortcoming which hinders their large-time analysis. We circumvent this issue by lifting the system to a measure-valued stochastic evolution equation introduced by Cuchiero and Teichmann~\cite{CT18}, whence we retrieve the Markov property. Leveraging on the associated generalised Feller property, we extend the Krylov-Bogoliubov theorem to this infinite-dimensional setting and thus establish an approach to the existence of invariant measures. We present concrete examples, including the rough Heston model from Mathematical Finance.

math.PR

A Macroscopically Consistent Reactive Langevin Dynamics Model

Particle-based stochastic reaction-diffusion (PBSRD) models are a popular approach for capturing stochasticity in reaction and transport processes across biological systems. In some contexts, the overdamped approximation inherent in such models may be inappropriate, necessitating the use of more microscopic Langevin Dynamics models for spatial transport. In this work we develop a novel particle-based Reactive Langevin Dynamics (RLD) model, with a focus on deriving reactive interaction kernels that are consistent with the physical constraint of detailed balance of reactive fluxes at equilibrium. We demonstrate that, to leading order, the overdamped limit of the resulting RLD model corresponds to the volume reactivity PBSRD model, of which the well-known Doi model is a particular instance. Our work provides a step towards systematically deriving PBSRD models from more microscopic reaction models, and suggests possible constraints on the latter to ensure consistency between the two physical scales.

physics.bio-ph

Uniform-in-time bounds for a stochastic hybrid system with fast periodic sampling and small white-noise

We study the asymptotic behavior, uniform-in-time, of a non-linear dynamical system under the combined effects of fast periodic sampling with period $δ$ and small white noise of size $\varepsilon,\thinspace 0<\varepsilon,δ\ll 1$. The dynamics depend on both the current and recent measurements of the state, and as such it is not Markovian. Our main results can be interpreted as Law of Large Numbers (LLN) and Central Limit Theorem (CLT) type results. LLN type result shows that the resulting stochastic process is close to an ordinary differential equation (ODE) uniformly in time as $\varepsilon,δ\searrow 0.$ Further, in regards to CLT, we provide quantitative and uniform-in-time control of the fluctuations process. The interaction of the small parameters provides an additional drift term in the limiting fluctuations, which captures both the sampling and noise effects. As a consequence, we obtain a first-order perturbation expansion of the stochastic process along with time-independent estimates on the remainder. The zeroth- and first-order terms in the expansion are given by an ODE and SDE, respectively. Simulation studies that illustrate and supplement the theoretical results are also provided.

math.PR

Uniform attraction and exit problems for stochastic damped wave equations

We consider a class of wave equations with constant damping and polynomial nonlinearities that are perturbed by small, multiplicative, space-time white noise. The equations are defined on a one-dimensional bounded interval with Dirichlet boundary conditions, continuous initial position and distributional initial velocity. In the first part of this work, we study the corresponding deterministic dynamics and prove that certain neighborhoods of asymptotically stable equilibria are uniformly attracting in the topology of uniform convergence. Then, we consider exit problems for local solutions of the stochastic damped wave equations from bounded domains $D$ of uniform attraction. Using tools from large deviations along with novel controllability results, we obtain logarithmic asymptotics for exit times and exit places, in the vanishing noise limit, that are expressed in terms of the corresponding quasipotential. In doing so, we develop arguments that take into account the lack of both smoothing and exact controllability that are inherent to the problem at hand. Moreover, our exit time results provide asymptotic lower bounds for the mean explosion time of local solutions. We introduce a novel notion of "regular" boundary points allowing to avoid the question of boundary smoothness in infinite dimensions and leading to the proof of a large deviations lower bound for the exit place. We illustrate this notion by providing explicit examples for different classes of domains $D$. Conditions under which lower and upper bounds for exit time and exit place logarithmic asymptotic hold, are also presented. In addition, we obtain deterministic stability results for linear damped wave equations that are of independent interest.

math.PR

Convergence Analysis of Real-time Recurrent Learning (RTRL) for a class of Recurrent Neural Networks

Recurrent neural networks (RNNs) are commonly trained with the truncated backpropagation-through-time (TBPTT) algorithm. For the purposes of computational tractability, the TBPTT algorithm truncates the chain rule and calculates the gradient on a finite block of the overall data sequence. Such approximation could lead to significant inaccuracies, as the block length for the truncated backpropagation is typically limited to be much smaller than the overall sequence length. In contrast, Real-time recurrent learning (RTRL) is an online optimization algorithm which asymptotically follows the true gradient of the loss on the data sequence as the number of sequence time steps $t \rightarrow \infty$. RTRL forward propagates the derivatives of the RNN hidden/memory units with respect to the parameters and, using the forward derivatives, performs online updates of the parameters at each time step in the data sequence. RTRL's online forward propagation allows for exact optimization over extremely long data sequences, although it can be computationally costly for models with large numbers of parameters. We prove convergence of the RTRL algorithm for a class of RNNs. The convergence analysis establishes a fixed point for the joint distribution of the data sequence, RNN hidden layer, and the RNN hidden layer forward derivatives as the number of data samples from the sequence and the number of training steps tend to infinity. We prove convergence of the RTRL algorithm to a stationary point of the loss. Numerical studies illustrate our theoretical results. One potential application area for RTRL is the analysis of financial data, which typically involve long time series and models with small to medium numbers of parameters. This makes RTRL computationally tractable and a potentially appealing optimization method for training models. Thus, we include an example of RTRL applied to limit order book data.

cs.LG

Stochastic gradient descent-based inference for dynamic network models with attractors

In Coevolving Latent Space Networks with Attractors (CLSNA) models, nodes in a latent space represent social actors, and edges indicate their dynamic interactions. Attractors are added at the latent level to capture the notion of attractive and repulsive forces between nodes, borrowing from dynamical systems theory. However, CLSNA reliance on MCMC estimation makes scaling difficult, and the requirement for nodes to be present throughout the study period limit practical applications. We address these issues by (i) introducing a Stochastic gradient descent (SGD) parameter estimation method, (ii) developing a novel approach for uncertainty quantification using SGD, and (iii) extending the model to allow nodes to join and leave over time. Simulation results show that our extensions result in little loss of accuracy compared to MCMC, but can scale to much larger networks. We apply our approach to the longitudinal social networks of members of US Congress on the social media platform X. Accounting for node dynamics overcomes selection bias in the network and uncovers uniquely and increasingly repulsive forces within the Republican Party.

stat.ME

Quantitative fluctuation analysis of multiscale diffusion systems via Malliavin calculus

We study fluctuations of small noise multiscale diffusions around their homogenized deterministic limit. We derive quantitative rates of convergence of the fluctuation processes to their Gaussian limits in the appropriate Wasserstein metric requiring detailed estimates of the first and second order Malliavin derivative of the slow component. We study a fully coupled system and the derivation of the quantitative rates of convergence depends on a very careful decomposition of the first and second Malliavin derivatives of the slow and fast component to terms that have different rates of convergence depending on the strength of the noise and timescale separation parameter.

math.PR

Transport map unadjusted Langevin algorithms: learning and discretizing perturbed samplers

Langevin dynamics are widely used in sampling high-dimensional, non-Gaussian distributions whose densities are known up to a normalizing constant. In particular, there is strong interest in unadjusted Langevin algorithms (ULA), which directly discretize Langevin dynamics to estimate expectations over the target distribution. We study the use of transport maps that approximately normalize a target distribution as a way to precondition and accelerate the convergence of Langevin dynamics. We show that in continuous time, when a transport map is applied to Langevin dynamics, the result is a Riemannian manifold Langevin dynamics (RMLD) with metric defined by the transport map. We also show that applying a transport map to an irreversibly-perturbed ULA results in a geometry-informed irreversible perturbation (GiIrr) of the original dynamics. These connections suggest more systematic ways of learning metrics and perturbations, and also yield alternative discretizations of the RMLD described by the map, which we study. Under appropriate conditions, these discretized processes can be endowed with non-asymptotic bounds describing convergence to the target distribution in 2-Wasserstein distance. Illustrative numerical results complement our theoretical claims.

stat.ME