SearcharxivSearch

arXiv subjects

Wenqing Hu

Publications and source records attributed to Wenqing Hu.

At least 19 recordsLinked to original sources

Coadjoint averaging and cross-scale fluxes in a fast-slow stochastic Euler-Arnold system on $SU(N)$

We study the emergence of (cross-scale) fluxes associated with energy and enstrophy in a stochastic version of Zeitlin's $SU(N)$ approximation of 2-d Euler dynamics. Motivated by the Euler-Arnold formulation, we interpret the nonlinear transport as motion along coadjoint orbits, reflecting an underlying symmetry that preserves all Casimir invariants. We introduce a fast-slow stochastic framework in which rapid mixing is modeled by a structured fast stochastic forcing acting along selected directions. We show that when the fast dynamics preserves the full coadjoint orbit symmetry given by the natural symplectic structure, the averaged system exhibits no nontrivial flux. In contrast, when this symmetry is broken by tangential non-Hamiltonian vector fields on the coadjoint orbit, we identify conditions that yield nonzero fluxes carried by the Euler-Arnold nonlinearity. Thus, coadjoint symmetry suppresses averaged nonlinear flux, whereas its breaking can select a preferred direction of cross-scale transfer.

math.PR

Wave propagation for 1-dimensional reaction-diffusion equations with nonzero random drift

We consider the wave propagation for a reaction-diffusion equation on the real line, with a random drift and Fisher-Kolmogorov-Petrovskii-Piscounov (FKPP) type nonlinear reaction. We show that when the average drift is positive, the asymptotic wave fronts propagating to the positive and negative directions are both pushed in the negative direction, leading to the possibility that both wave fronts propagate toward negative infinity. Our proof is based on the Large Deviations Principle for diffusion processes in random environments, as well as an analysis of the Feynman-Kac formula. Such probabilistic arguments also reveal the underlying physical mechanism of the wave fronts formation: the drift acts as an external field that shifts the (quenched) free-energy reference level without altering the intrinsic fluctuation structure of the system.

math.AP

QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models

Large Language Models (LLMs) have shown strong capabilities across many domains, yet their evaluation in financial quantitative tasks remains fragmented and mostly limited to knowledge-centric question answering. We introduce QuantEval, a benchmark that evaluates LLMs across three essential dimensions of quantitative finance: knowledge-based QA, quantitative mathematical reasoning, and quantitative strategy coding. Unlike prior financial benchmarks, QuantEval integrates a CTA-style backtesting framework that executes model-generated strategies and evaluates them using financial performance metrics, enabling a more realistic assessment of quantitative coding ability. We evaluate some state-of-the-art open-source and proprietary LLMs and observe substantial gaps to human experts, particularly in reasoning and strategy coding. Finally, we conduct large-scale supervised fine-tuning and reinforcement learning experiments on domain-aligned data, demonstrating consistent improvements. We hope QuantEval will facilitate research on LLMs' quantitative finance capabilities and accelerate their practical adoption in real-world trading workflows. We additionally release the full deterministic backtesting configuration (asset universe, cost model, and metric definitions) to ensure strict reproducibility.

cs.CL

MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss

Infrared imaging has emerged as a robust solution for urban object detection under low-light and adverse weather conditions, offering significant advantages over traditional visible-light cameras. However, challenges such as class imbalance, thermal noise, and computational constraints can significantly hinder model performance in practical settings. To address these issues, we evaluate multiple YOLO variants on the FLIR ADAS V2 dataset, ultimately selecting YOLOv8 as our baseline due to its balanced accuracy and efficiency. Building on this foundation, we present \texttt{MS-YOLO} (\textbf{M}obileNetv4 and \textbf{S}lideLoss based on YOLO), which replaces YOLOv8's CSPDarknet backbone with the more efficient MobileNetV4, reducing computational overhead by \textbf{1.5%} while sustaining high accuracy. In addition, we introduce \emph{SlideLoss}, a novel loss function that dynamically emphasizes under-represented and occluded samples, boosting precision without sacrificing recall. Experiments on the FLIR ADAS V2 benchmark show that \texttt{MS-YOLO} attains competitive mAP and superior precision while operating at only \textbf{6.7 GFLOPs}. These results demonstrate that \texttt{MS-YOLO} effectively addresses the dual challenge of maintaining high detection quality while minimizing computational costs, making it well-suited for real-time edge deployment in urban environments.

cs.CV

B2MAPO: A Batch-by-Batch Multi-Agent Policy Optimization to Balance Performance and Efficiency

Most multi-agent reinforcement learning approaches adopt two types of policy optimization methods that either update policy simultaneously or sequentially. Simultaneously updating policies of all agents introduces non-stationarity problem. Although sequentially updating policies agent-by-agent in an appropriate order improves policy performance, it is prone to low efficiency due to sequential execution, resulting in longer model training and execution time. Intuitively, partitioning policies of all agents according to their interdependence and updating joint policy batch-by-batch can effectively balance performance and efficiency. However, how to determine the optimal batch partition of policies and batch updating order are challenging problems. Firstly, a sequential batched policy updating scheme, B2MAPO (Batch by Batch Multi-Agent Policy Optimization), is proposed with a theoretical guarantee of the monotonic incrementally tightened bound. Secondly, a universal modulized plug-and-play B2MAPO hierarchical framework, which satisfies CTDE principle, is designed to conveniently integrate any MARL models to fully exploit and merge their merits, including policy optimality and inference efficiency. Next, a DAG-based B2MAPO algorithm is devised, which is a carefully designed implementation of B2MAPO framework. Comprehensive experimental results conducted on StarCraftII Multi-agent Challenge and Google Football Research demonstrate the performance of DAG-based B2MAPO algorithm outperforms baseline methods. Meanwhile, compared with A2PO, our algorithm reduces the model training and execution time by 60.4% and 78.7%, respectively.

cs.MA

On the Posterior Distribution of a Random Process Conditioned on Empirical Frequencies of a Finite Path: the i.i.d and finite Markov chain case

We obtain the posterior distribution of a random process conditioned on observing the empirical frequencies of a finite sample path. We find under a rather broad assumption on the "dependence structure" of the process, {\em c.f.} independence or Markovian, the posterior marginal distribution of the process at a given time index can be identified as certain empirical distribution computed from the observed empirical frequencies of the sample path. We show that in both cases of discrete-valued i.i.d. sequence and finite Markov chain, a certain "conditional symmetry" given by the observation of the empirical frequencies leads to the desired result on the posterior distribution. Results for both finite-time observations and its asymptotic infinite-time limit are connected via the idea of Gibbs conditioning. Finally, since our results demonstrate a central role of the empirical frequency in understanding the information content of data, we use the Large Deviations Principle (LDP) to construct a general notion of "data-driven entropy", from which one can apply a formalism from the recent study of statistical thermodynamics to data.

math.PR

Wave propagation for reaction-diffusion equations on infinite random trees

The asymptotic wave speed for FKPP type reaction-diffusion equations on a class of infinite random metric trees are considered. We show that a travelling wavefront emerges, provided that the reaction rate is large enough. The wavefront travels at a speed that can be quantified via a variational formula involving the random branching degrees $\vec{d}$ and the random branch lengths $\vec{\ell}$ of the tree. This speed is slower than that of the same equation on the real line $\mathbb{R}$, and we estimate this slow down in terms of $\vec{d}$ and $\vec{\ell}$. Our key idea is to project the Brownian motion on the tree onto a one-dimensional axis along the direction of the wave propagation. The projected process is a multi-skewed Brownian motion, introduced by Ramirez [Multi-skewed Brownian motion and diffusion in layered media, Proc. Am. Math. Soc., Vol. 139, No. 10, pp.3739-3752, 2011], with skewness and interface sets that encode the metric structure $(\vec{d}, \vec{\ell})$ of the tree. Combined with analytic arguments based on the Feynman-Kac formula, this idea connects our analysis of the wavefront propagation to the large deviations principle (LDP) of the multi-skewed Brownian motion with random skewness and random interface set. Our LDP analysis involves delicate estimates for an infinite product of $2\times 2$ random matrices parametrized by $\vec{d}$ and $\vec{\ell}$ and for hitting times of a random walk in random environment. By exhausting all possible shapes of the LDP rate function (action functional), the analytic arguments that bridge the LDP and the wave propagation overcome the random drift effect due to multi-skewness.

math.PR

On the Noisy Gradient Descent that Generalizes as SGD

The gradient noise of SGD is considered to play a central role in the observed strong generalization abilities of deep learning. While past studies confirm that the magnitude and the covariance structure of gradient noise are critical for regularization, it remains unclear whether or not the class of noise distributions is important. In this work we provide negative results by showing that noises in classes different from the SGD noise can also effectively regularize gradient descent. Our finding is based on a novel observation on the structure of the SGD noise: it is the multiplication of the gradient matrix and a sampling noise that arises from the mini-batch sampling procedure. Moreover, the sampling noises unify two kinds of gradient regularizing noises that belong to the Gaussian class: the one using (scaled) Fisher as covariance and the one using the gradient covariance of SGD as covariance. Finally, thanks to the flexibility of choosing noise class, an algorithm is proposed to perform noisy gradient descent that generalizes well, the variant of which even benefits large batch SGD training without hurting generalization.

cs.LG

Stochastic Recursive Momentum Method for Non-Convex Compositional Optimization

We propose a novel stochastic optimization algorithm called STOchastic Recursive Momentum for Compositional (STORM-Compositional) optimization that minimizes the composition of expectations of two stochastic functions, the latter being an optimization problem arising in various important machine learning applications. By introducing the momentum term in the compositional gradient updates, STORM-Compositional operates the stochastic recursive variance-reduced compositional gradients in an exponential-moving average way. This leads to an $O(\varepsilon^{-3})$ complexity upper bound for STORM-Compositional, that matches the best known complexity bounds in previously announced compositional optimization algorithms. At the same time, STORM-Compositional is a single loop algorithm that avoids the typical alternative tuning between large and small batch sizes, as well as recording of checkpoint gradients, that persist in variance-reduced stochastic gradient methods. This allows considerably simpler parameter tuning in numerical experiments, which demonstrates the superiority of STORM-Compositional over other stochastic compositional optimization algorithms.

math.OC

On the fast convergence of random perturbations of the gradient flow

We consider in this work small random perturbations (of multiplicative noise type) of the gradient flow. We prove that under mild conditions, when the potential function is a Morse function with additional strong saddle condition, the perturbed gradient flow converges to the neighborhood of local minimizers in $O(\ln (\varepsilon^{-1}))$ time on the average, where $\varepsilon$ is the scale of the random perturbation. Under a change of time scale, this indicates that for the diffusion process that approximates the stochastic gradient method, it takes (up to logarithmic factor) only a linear time of inverse stepsize to evade from all saddle points. This can be regarded as a manifestation of fast convergence of the discrete-time stochastic gradient method, the latter being used heavily in modern statistical machine learning.

math.PR

On the Global Convergence of Continuous-Time Stochastic Heavy-Ball Method for Nonconvex Optimization

We study the convergence behavior of the stochastic heavy-ball method with a small stepsize. Under a change of time scale, we approximate the discrete method by a stochastic differential equation that models small random perturbations of a coupled system of nonlinear oscillators. We rigorously show that the perturbed system converges to a local minimum in a logarithmic time. This indicates that for the diffusion process that approximates the stochastic heavy-ball method, it takes (up to a logarithmic factor) only a linear time of the square root of the inverse stepsize to escape from all saddle points. This results may suggest a fast convergence of its discrete-time counterpart. Our theoretical results are validated by numerical experiments.

math.PR

Large deviations and averaging for systems of slow--fast stochastic reaction--diffusion equations

We study a large deviation principle for a system of stochastic reaction--diffusion equations (SRDEs) with a separation of fast and slow components and small noise in the slow component. The derivation of the large deviation principle is based on the weak convergence method in infinite dimensions, which results in studying averaging for controlled SRDEs. By appropriate choice of the parameters, the fast process and the associated control that arises from the weak convergence method decouple from each other. We show that in this decoupling case one can use the weak convergence method to characterize the limiting process via a "viable pair" that captures the limiting controlled dynamics and the effective invariant measure simultaneously. The characterization of the limit of the controlled slow-fast processes in terms of viable pair enables us to obtain a variational representation of the large deviation action functional. Due to the infinite--dimensional nature of our set--up, the proof of tightness as well as the analysis of the limit process and in particular the proof of the large deviations lower bound is considerably more delicate here than in the finite--dimensional situation. Smoothness properties of optimal controls in infinite dimensions (a necessary step for the large deviations lower bound) need to be established. We emphasize that many issues that are present in the infinite dimensional case, are completely absent in finite dimensions.

math.PR

On the long-time behavior of a perturbed conservative system with degeneracy

We consider in this work a model conservative system subject to dissipation and Gaussian-type stochastic perturbations. The original conservative system possesses a continuous set of steady states, and is thus degenerate. We characterize the long-time limit of our model system as the perturbation parameter tends to zero. The degeneracy in our model system carries features found in some partial differential equations related, for example, to turbulence problems.

math.PR

Quasi-potential as an implicit regularizer for the loss function in the stochastic gradient descent

We interpret the variational inference of the Stochastic Gradient Descent (SGD) as minimizing a new potential function named the \textit{quasi-potential}. We analytically construct the quasi-potential function in the case when the loss function is convex and admits only one global minimum point. We show in this case that the quasi-potential function is related to the noise covariance structure of SGD via a partial differential equation of Hamilton-Jacobi type. This relation helps us to show that anisotropic noise leads to faster escape than isotropic noise. We then consider the dynamics of SGD in the case when the loss function is non-convex and admits several different local minima. In this case, we demonstrate an example that shows how the noise covariance structure plays a role in "implicit regularization", a phenomenon in which SGD favors some particular local minimum points. This is done through the relation between the noise covariance structure and the quasi-potential function. Our analysis is based on Large Deviations Theory (LDT), and they are validated by numerical experiments.

cs.LG

A convergence analysis of the perturbed compositional gradient flow: averaging principle and normal deviations

We consider in this work a system of two stochastic differential equations named the perturbed compositional gradient flow. By introducing a separation of fast and slow scales of the two equations, we show that the limit of the slow motion is given by an averaged ordinary differential equation. We then demonstrate that the deviation of the slow motion from the averaged equation, after proper rescaling, converges to a stochastic process with Gaussian inputs. This indicates that the slow motion can be approximated in the weak sense by a standard perturbed gradient flow or the continuous-time stochastic gradient descent algorithm that solves the optimization problem for a composition of two functions. As an application, the perturbed compositional gradient flow corresponds to the diffusion limit of the Stochastic Composite Gradient Descent (SCGD) algorithm for minimizing a composition of two expected-value functions in the optimization literatures. For the strongly convex case, such an analysis implies that the SCGD algorithm has the same convergence time asymptotic as the classical stochastic gradient descent algorithm. Thus it validates, at the level of continuous approximation, the effectiveness of using the SCGD algorithm in the strongly convex case.

math.PR

On the diffusion approximation of nonconvex stochastic gradient descent

We study the Stochastic Gradient Descent (SGD) method in nonconvex optimization problems from the point of view of approximating diffusion processes. We prove rigorously that the diffusion process can approximate the SGD algorithm weakly using the weak form of master equation for probability evolution. In the small step size regime and the presence of omnidirectional noise, our weak approximating diffusion process suggests the following dynamics for the SGD iteration starting from a local minimizer (resp.~saddle point): it escapes in a number of iterations exponentially (resp.~almost linearly) dependent on the inverse stepsize. The results are obtained using the theory for random perturbations of dynamical systems (theory of large deviations for local minimizers and theory of exiting for unstable stationary points). In addition, we discuss the effects of batch size for the deep neural networks, and we find that small batch size is helpful for SGD algorithms to escape unstable stationary points and sharp minimizers. Our theory indicates that one should increase the batch size at later stage for the SGD to be trapped in flat minimizers for better generalization.

stat.ML

FWDA: a Fast Wishart Discriminant Analysis with its Application to Electronic Health Records Data Classification

Linear Discriminant Analysis (LDA) on Electronic Health Records (EHR) data is widely-used for early detection of diseases. Classical LDA for EHR data classification, however, suffers from two handicaps: the ill-posed estimation of LDA parameters (e.g., covariance matrix), and the "linear inseparability" of EHR data. To handle these two issues, in this paper, we propose a novel classifier FWDA -- Fast Wishart Discriminant Analysis, that makes predictions in an ensemble way. Specifically, FWDA first surrogates the distribution of inverse covariance matrices using a Wishart distribution estimated from the training data, then "weighted-averages" the classification results of multiple LDA classifiers parameterized by the sampled inverse covariance matrices via a Bayesian Voting scheme. The weights for voting are optimally updated to adapt each new input data, so as to enable the nonlinear classification. Theoretical analysis indicates that FWDA possesses a fast convergence rate and a robust performance on high dimensional data. Extensive experiments on large-scale EHR dataset show that our approach outperforms state-of-the-art algorithms by a large margin.

cs.LG

Itô's formula, the stochastic exponential and change of measure on general time scales

We provide an Itô's formula for stochastic dynamical equation on general time scales. Based on this Itô's formula we give a closed form expression for stochastic exponential on general time scales. We then demonstrate a Girsanov's change of measure formula in the case of general time scales. Our result is being applied to a Brownian motion on the quantum time scale (q-time scale).

math.PR