SearcharxivSearch

arXiv subjects

Hailiang Liu

Publications and source records attributed to Hailiang Liu.

At least 19 recordsLinked to original sources

A Mean-Field Theory of Transformers: Well-Posedness of the Coupled Data--Parameter Dynamics and Global Convergence of Training

We develop a rigorous mean-field theory for transformer networks that captures two large-scale limits inherent in the architecture: the number of tokens $N\to\infty$ in the input sequence and the number of attention heads $H\to\infty$ in each layer. It further considers an infinite number of layers leading to a time-continuous formulation. The resulting framework couples two interacting mean-field objects: a token distribution $\mu_t\in\mathcal P(\mathbb R^d)$, which evolves through network depth $t\in[0,T]$ according to a McKean--Vlasov transport equation, and the attention-parameter distribution $\rho_s\in\mathcal P(\Theta)$, that evolves through training time $s\ge0$ according to a Wasserstein gradient flow of the empirical risk, with optional entropic or Tikhonov regularization. We establish a comprehensive analytical foundation for the system coupling the transformer and the training dynamics; in particular, we establish global well-posedness of the resulting nonlinear Fokker--Planck system describing the coupled mean-field and training dynamics. Beyond well-posedness, we connect the mean-field formulation to optimization. For shallow, single-layer attention models, we prove exponential convergence to the entropy-regularized global optimum under a log-Sobolev condition. For genuinely deep, compositional transformers, we establish local linear convergence under a Neural Tangent Kernel non-degeneracy condition.

math.AP

A Unified Variational Framework for Optimal Transport with Lagrangian Costs

We investigate optimal transport distances induced by general Lagrangian action functionals. Extending the classical Monge - Kantorovich and Benamou - Brenier theories, we derive a unified variational framework that connects several equivalent formulations of the induced transport distance, including Lagrangian, Eulerian, convex optimization, Hamilton - Jacobi dual, and Hamiltonian flow formulations. Under standard convexity assumptions on the Lagrangian, we establish the equivalence of these formulations through variational arguments and convex duality. The resulting optimality system reveals a natural Hamiltonian structure on the Wasserstein space, providing a direct link between optimal transport, Hamiltonian dynamics, and optimal control. The proposed framework extends the classical quadratic-cost theory to general Lagrangian costs and offers a unified perspective for the analysis of transport metrics generated by action functionals.

math.OC

Inf-Sup Neural Networks for High Dimensional PDEs

Solving partial differential equations (PDEs) in high dimensions remains challenging due to the curse of dimensionality. We propose a neural-network-based framework that reformulates PDEs as inf--sup optimization problems through the introduction of a Lagrange multiplier. The primal solution and the associated Lagrange multiplier are parameterized by two networks and are computed via an iterative saddle-point optimization procedure. We prove the theoretical equivalence between the proposed optimization formulation and the original PDE problem, and we derive rigorous error estimates that quantify the total approximation error in terms of the network approximation error, statistical (sampling) error, and optimization error. Numerical experiments demonstrate the accuracy, stability, and efficiency of the proposed method for solving high-dimensional PDEs.

math.NA

Neural solver for Wasserstein Geodesics and optimal transport dynamics

In recent years, the machine learning community has increasingly embraced the optimal transport (OT) framework for modeling distributional relationships. In this work, we introduce a sample-based neural solver for computing the Wasserstein geodesic between a source and target distribution, along with the associated velocity field. Building on the dynamical formulation of the optimal transport (OT) problem, we recast the constrained optimization as a minimax problem, using deep neural networks to approximate the relevant functions. This approach not only provides the Wasserstein geodesic but also recovers the OT map, enabling direct sampling from the target distribution. By estimating the OT map, we obtain velocity estimates along particle trajectories, which in turn allow us to learn the full velocity field. The framework is flexible and readily extends to general cost functions, including the commonly used quadratic cost. We demonstrate the effectiveness of our method through experiments on both synthetic and real datasets.

cs.LG

A Primal-Dual Level Set Method for Computing Geodesic Distances

The numerical computation of shortest paths or geodesics on surfaces, along with the associated geodesic distance, has a wide range of applications. Compared to Euclidean distance computation, these tasks are more complex due to the influence of surface geometry on the behavior of shortest paths. This paper introduces a primal-dual level set method for computing geodesic distances. A key insight is that the underlying surface can be implicitly represented as a zero level set, allowing us to formulate a constraint minimization problem. We employ the primal-dual methodology, along with regularization and acceleration techniques, to develop our algorithm. This approach is robust, efficient, and easy to implement. We establish a convergence result for the high-resolution PDE system, and numerical evidence suggests that the method converges to a geodesic in the limit of refinement.

math.NA

A Generalized Energy-Based Adaptive Gradient Method for Optimization

Adaptive Gradient Descent with Energy (AEGD) is a variant of Gradient Descent (GD) designed to address step size sensitivity through an energy-based formulation. AEGD is notable for its unconditional energy stability, ensuring convergence in energy regardless of the initial step size. In this work, we propose the Generalized Energy-Based Adaptive Gradient (gAEGD) method, which extends AEGD by generalizing the energy function beyond the square root form to a broader class of functions. We prove that gAEGD retains the unconditional energy stability property, remains robust to step size selection, and exhibits a two-phase adaptive dynamic: the effective step size first adjusts adaptively, then stabilizes within a range that guarantees decay of the objective function values. We establish an optimal convergence rate of $O(1/k)$ for finding an $\epsilon$-stationary point, along with improved convergence rates for the objective gap under a local Kurdyka-{\L}ojasiewicz (KL) condition. Empirical results support the theoretical analysis and indicate that the generalized energy-based approach preforms effectively and reliably for a broad range of optimization problems.

math.OC

Global Convergence in Neural ODEs: Impact of Activation Functions

Neural Ordinary Differential Equations (ODEs) have been successful in various applications due to their continuous nature and parameter-sharing efficiency. However, these unique characteristics also introduce challenges in training, particularly with respect to gradient computation accuracy and convergence analysis. In this paper, we address these challenges by investigating the impact of activation functions. We demonstrate that the properties of activation functions, specifically smoothness and nonlinearity, are critical to the training dynamics. Smooth activation functions guarantee globally unique solutions for both forward and backward ODEs, while sufficient nonlinearity is essential for maintaining the spectral properties of the Neural Tangent Kernel (NTK) during training. Together, these properties enable us to establish the global convergence of Neural ODEs under gradient descent in overparameterized regimes. Our theoretical findings are validated by numerical experiments, which not only support our analysis but also provide practical guidelines for scaling Neural ODEs, potentially leading to faster training and improved performance in real-world applications.

cs.LG

A positivity-preserving hybrid DDG method for Poisson--Nernst--Planck systems

In earlier work [H. Liu and Z. Wang, J. Comput. Phys., 328(2017)], an arbitrary high-order conservative and energy-dissipative direct discontinuous Galerkin (DDG) scheme was developed. Although this scheme enforced solution positivity using cell averages as reference values, it lacked a theoretical guarantee for the positivity of those cell averages. In this study, we develop a novel arbitrary high-order DDG method with rigorously proven positivity-preserving properties. Specifically, the positivity of the cell averages is ensured through a modified numerical flux in combination with forward Euler time discretization. To achieve point-wise positivity of ion concentrations, we introduce a hybrid algorithm that integrates a positivity-preserving limiter. The proposed method is further extended to higher-dimensional problems with rectangular meshes. Numerical results confirm the scheme's high-order accuracy, guaranteed positivity preservation, and consistent discrete energy dissipation.

math.NA

Variational structure of Fokker-Planck equations with variable mobility

We study Fokker--Planck equations with symmetric, positive definite mobility matrices capturing diffusion in heterogeneous environments. A weighted Wasserstein metric is introduced for which these equations are gradient flows. This metric is shown to emerge from an optimal control problem in the space of probability densities for a class of variable mobility matrices, with the cost function capturing the work dissipated via friction. Using the Nash-Kuiper isometric embedding theorem for Riemannian manifolds, we demonstrate the existence of optimal transport maps. Additionally, we construct a time-discrete variational scheme, establish key properties for the associated minimizing problem, and prove convergence to weak solutions of the associated Fokker-Planck equation.

math.OC

The Onsager principle and structure preserving numerical schemes

We present a natural framework for constructing energy-stable time discretization schemes. By leveraging the Onsager principle, we demonstrate its efficacy in formulating partial differential equation models for diverse gradient flow systems. Furthermore, this principle provides a robust basis for developing numerical schemes that uphold crucial physical properties. Within this framework, several widely used schemes emerge naturally, showing its versatility and applicability.

math.NA

Positivity-preserving and energy-dissipating discontinuous Galerkin methods for nonlinear nonlocal Fokker-Planck equations

This paper is concerned with structure-preserving numerical approximations for a class of nonlinear nonlocal Fokker-Planck equations, which admit a gradient flow structure and find application in diverse contexts. The solutions, representing density distributions, must be non-negative and satisfy a specific energy dissipation law. We design an arbitrary high-order discontinuous Galerkin (DG) method tailored for these model problems. Both semi-discrete and fully discrete schemes are shown to admit the energy dissipation law for non-negative numerical solutions. To ensure the preservation of positivity in cell averages at all time steps, we introduce a local flux correction applied to the DDG diffusive flux. Subsequently, a hybrid algorithm is presented, utilizing a positivity-preserving limiter, to generate positive and energy-dissipating solutions. Numerical examples are provided to showcase the high resolution of the numerical solutions and the verified properties of the DG schemes.

math.NA

Unconditionally energy stable IEQ-FEMs for the Cahn-Hilliard equation and Allen-Cahn equation

In this paper, we present several unconditionally energy-stable invariant energy quadratization (IEQ) finite element methods (FEMs) with linear, first- and second-order accuracy for solving both the Cahn-Hilliard equation and the Allen-Cahn equation. For time discretization, we compare three distinct IEQ-FEM schemes that position the intermediate function introduced by the IEQ approach in different function spaces: finite element space, continuous function space, or a combination of these spaces. Rigorous proofs establishing the existence and uniqueness of the numerical solution, along with analyses of energy dissipation for both equations and mass conservation for the Cahn-Hilliard equation, are provided. The proposed schemes' accuracy, efficiency, and solution properties are demonstrated through numerical experiments.

math.NA

Inf-Sup neural networks for high-dimensional elliptic PDE problems

Solving high dimensional partial differential equations (PDEs) has historically posed a considerable challenge when utilizing conventional numerical methods, such as those involving domain meshes. Recent advancements in the field have seen the emergence of neural PDE solvers, leveraging deep networks to effectively tackle high dimensional PDE problems. This study introduces Inf-SupNet, a model-based unsupervised learning approach designed to acquire solutions for a specific category of elliptic PDEs. The fundamental concept behind Inf-SupNet involves incorporating the inf-sup formulation of the underlying PDE into the loss function. The analysis reveals that the global solution error can be bounded by the sum of three distinct errors: the numerical integration error, the duality gap of the loss function (training error), and the neural network approximation error for functions within Sobolev spaces. To validate the efficacy of the proposed method, numerical experiments conducted in high dimensions demonstrate its stability and accuracy across various boundary conditions, as well as for both semi-linear and nonlinear PDEs.

math.NA

Data-driven optimal control with neural network modeling of gradient flows

Extracting physical laws from observation data is a central challenge in many diverse areas of science and engineering. We propose Optimal Control Neural Networks (OCN) to learn the laws of vector fields in dynamical systems, with no assumption on their analytical form, given data consisting of sampled trajectories. The OCN framework consists of a neural network representation and an optimal control formulation. We provide error bounds for both the solution and the vector field. The bounds are shown to depend on both the training error and the time step between the observation data. We also demonstrate the effectiveness of OCN, as well as its generalization ability, by testing on several canonical systems, including the chaotic Lorenz system.

math.DS

Wide Neural Networks as Gaussian Processes: Lessons from Deep Equilibrium Models

Neural networks with wide layers have attracted significant attention due to their equivalence to Gaussian processes, enabling perfect fitting of training data while maintaining generalization performance, known as benign overfitting. However, existing results mainly focus on shallow or finite-depth networks, necessitating a comprehensive analysis of wide neural networks with infinite-depth layers, such as neural ordinary differential equations (ODEs) and deep equilibrium models (DEQs). In this paper, we specifically investigate the deep equilibrium model (DEQ), an infinite-depth neural network with shared weight matrices across layers. Our analysis reveals that as the width of DEQ layers approaches infinity, it converges to a Gaussian process, establishing what is known as the Neural Network and Gaussian Process (NNGP) correspondence. Remarkably, this convergence holds even when the limits of depth and width are interchanged, which is not observed in typical infinite-depth Multilayer Perceptron (MLP) networks. Furthermore, we demonstrate that the associated Gaussian vector remains non-degenerate for any pairwise distinct input data, ensuring a strictly positive smallest eigenvalue of the corresponding kernel matrix using the NNGP kernel. These findings serve as fundamental elements for studying the training and generalization of DEQs, laying the groundwork for future research in this area.

cs.LG

Adaptive Preconditioned Gradient Descent with Energy

We propose an adaptive step size with an energy approach for a suitable class of preconditioned gradient descent methods. We focus on settings where the preconditioning is applied to address the constraints in optimization problems, such as the Hessian-Riemannian and natural gradient descent methods. More specifically, we incorporate these preconditioned gradient descent algorithms in the recently introduced Adaptive Energy Gradient Descent (AEGD) framework. In particular, we discuss theoretical results on the unconditional energy-stability and convergence rates across three classes of objective functions. Furthermore, our numerical results demonstrate excellent performance of the proposed method on several test bed optimization problems.

math.OC

On critical thresholds for hyperbolic balance law systems

We review the theoretical development in the study of critical thresholds for hyperbolic balance laws. The emphasis is on two classes of systems: Euler-Poisson-alignment (EPA) systems and hyperbolic relaxation systems. We start with an introduction to the `Critical Threshold Phenomena' and study some nonlocal PDE systems, which are important from modeling point of view.

math.AP

A complete characterization of sharp thresholds to spherically symmetric multidimensional pressureless Euler-Poisson systems

The Euler-Poisson (EP) system models the dynamics of a variety of physical processes, including charge transport, collisional plasmas, and certain cosmological wave phenomena. In this work, we establish sharp critical threshold conditions that distinguish global-in-time regularity from finite-time breakdown for solutions of the radially symmetric, multidimensional pressureless EP system. Overall, there are two cases: with and without background ($c>0, c=0$ respectively). For $c>0$, we obtain precise thresholds assuming a periodicity condition. A key feature of our approach is that it extends seamlessly to the zero background case, where we obtain sharp thresholds without imposing any additional assumptions. In particular, the framework accommodates initial velocities that may be negative, allowing the flow to be directed toward the origin. The main analytical challenge of deriving threshold conditions for EP systems stems from the intricate coupling of various local/nonlocal forces. To overcome this, we identify a novel nonlinear quantity that plays a decisive role in the analysis and enables a unified treatment of all relevant scenarios. Our results provide a comprehensive characterization of critical thresholds for the pressureless EP system in multiple dimensions.

math.AP