SearcharxivSearch

arXiv subjects

Yewei Xu

Publications and source records attributed to Yewei Xu.

6 recordsLinked to original sources

Random Coordinate Descent on the Wasserstein Space of Probability Measures

Optimization over the space of probability measures endowed with the Wasserstein-2 geometry is central to modern machine learning and mean-field modeling. However, traditional methods relying on full Wasserstein gradients often suffer from high computational overhead in high-dimensional or ill-conditioned settings. We propose a randomized coordinate descent framework specifically designed for the Wasserstein manifold, introducing both Random Wasserstein Coordinate Descent (RWCD) and Random Wasserstein Coordinate Proximal{-Gradient} (RWCP) for composite objectives. By exploiting coordinate-wise structures, our methods adapt to anisotropic objective landscapes where full-gradient approaches typically struggle. We provide a rigorous convergence analysis across various landscape geometries, establishing guarantees under non-convex, Polyak-{\L}ojasiewicz, and geodesically convex conditions. Our theoretical results mirror the classic convergence properties found in Euclidean space, revealing a compelling symmetry between coordinate descent on vectors and on probability measures. The developed techniques are inherently adaptive to the Wasserstein geometry and offer a robust analytical template that can be extended to other optimization solvers within the space of measures. Numerical experiments on ill-conditioned energies demonstrate that our framework offers significant speedups over conventional full-gradient methods.

stat.ML

Forward Euler for Wasserstein Gradient Flows: Breakdown and Regularization

Wasserstein gradient flows have become a central tool for optimization problems over probability measures. A natural numerical approach is forward-Euler time discretization. We show, however, that even in the simple case where the energy functional is the Kullback-Leibler (KL) divergence against a smooth target density, forward-Euler can fail dramatically: the scheme does not converge to the gradient flow, despite the fact that the first variation $\nabla\frac{\delta F}{\delta\rho}$ remains formally well defined at every step. We identify the root cause as a loss of regularity induced by the discretization, and prove that a suitable regularization of the functional restores the necessary smoothness, making forward-Euler a viable solver that converges in discrete time to the global minimizer.

math.NA

Forward-Euler time-discretization for Wasserstein gradient flows can be wrong

In this note, we examine the forward-Euler discretization for simulating Wasserstein gradient flows. We provide two counter-examples showcasing the failure of this discretization even for a simple case where the energy functional is defined as the KL divergence against some nicely structured probability densities. A simple explanation of this failure is also discussed.

stat.ML

Correcting Auto-Differentiation in Neural-ODE Training

Does the use of auto-differentiation yield reasonable updates for deep neural networks (DNNs)? Specifically, when DNNs are designed to adhere to neural ODE architectures, can we trust the gradients provided by auto-differentiation? Through mathematical analysis and numerical evidence, we demonstrate that when neural networks employ high-order methods, such as Linear Multistep Methods (LMM) or Explicit Runge-Kutta Methods (ERK), to approximate the underlying ODE flows, brute-force auto-differentiation often introduces artificial oscillations in the gradients that prevent convergence. In the case of Leapfrog and 2-stage ERK, we propose simple post-processing techniques that effectively eliminates these oscillations, correct the gradient computation and thus returns the accurate updates.

cs.LG

Singular Vectors in Real Affine Subspaces

We prove inheritance of measure zero property of the set of singular vectors for affine subspaces and submanifolds inside those affine subspaces. We define a notion of $n$-singularity for matrices, which is closely related to the uniform exponent of irrationality. For certain affine subspaces, we show that the set of singular vectors has measure zero if and only if the parametrizing matrix is not $n$-singular. In particular, we show for affine hyperplanes the set of singular vectors has measure zero if and only if the parametrizing matrix is not rational.

math.NT

Singular Vectors and $\psi$-Dirichlet Numbers over Function Field

We show that the only $\psi$-Dirichlet numbers in a function field over a finite field are rational functions, unlike $\psi$-Dirichlet numbers in $\mathbb{R}$. We also prove that there are uncountably many totally irrational singular vectors with large uniform exponent in quadratic surfaces over a positive characteristic field.

math.NT