SearcharxivSearch

arXiv subjects

Du Ouyang

Publications and source records attributed to Du Ouyang.

8 recordsLinked to original sources

Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment

This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al. (2020), Jia and Zhou (2022b). With the logit parameterization for the stochastic policy, we show that it converges almost surely to the optimal arm under an arbitrary constant learning rate. Furthermore, we derive the non-asymptotic regret upper bound when the constant learning rate is below a time-invariant threshold; and the regret bound has order $O(\log T)$. We improve the analysis in Lattimore (2026a) for the same SDE by constructing a novel Lyapunov function and demonstrate the transparency of analyzing policy gradient using the tools in SDEs. In addition, the same Lyapunov function is also helpful in analyzing the discrete-time policy gradient algorithm.

cs.LG

Quasi-Monte Carlo for SDE Simulation: Error Analysis and Dimensionality Reduction

We investigate the numerical simulation of general stochastic differential equations (SDEs) using Quasi-Monte Carlo (QMC) methods. First, we provide a rigorous theoretical analysis of the QMC method applied to the Euler-Maruyama (EM) scheme, establishing that it significantly accelerates the decay of the sampling error and achieves an asymptotically superior convergence rate over the classical Monte Carlo method. Second, the traditional EM scheme exhibits a slow polynomial decay of the discretization error, which necessitates a large number of time steps and leads to a significantly high integration dimension. To address this issue, we propose a Multilevel Stochastic Time Grid (MSTG) method based on Exact Simulation techniques, and we rigorously establish its convergence rate under randomized QMC sampling, proving that it preserves the high-order convergence of the sampling error. In terms of the overall error, the truncation error of the proposed MSTG method exhibits a remarkably fast super-exponential decay. Consequently, to achieve a given accuracy level, our approach requires significantly fewer discretization steps than the EM scheme, thereby drastically reducing the actual integration dimension of the QMC method. This substantial dimensionality reduction strategy greatly enhances the practical efficiency of the QMC algorithm. Numerical experiments fully corroborate the superiority of the proposed approach.

math.NA

A Zeroth-Order Deep Learning Method for Fully Nonlinear Parabolic Partial Differential Equations with Unknown Coefficients

High-dimensional partial differential equations (PDEs) with unknown coefficients arise widely in scientific machine learning, including continuous-time reinforcement learning, yet solving them efficiently in a data-driven way remains challenging. Existing deep learning solvers often rely on repeated automatic differentiation to evaluate differential operators, which can cause instability and amplify derivative errors in high dimensions, while probabilistic methods based on stochastic representations require explicit knowledge of the data-generating dynamics and therefore do not apply to black-box environments. We introduce two types of simulators as data-generating mechanisms, and take a ``representing-then-learning" approach that learns the solutions and their derivatives under settings where the underlying PDE operators are accessible only through simulations and pointwise evaluations. Our representation of derivatives relies on the zeroth-order derivative (ZOD) estimators derived from perturbed Monte Carlo trajectories. This fully model-free approach generates targets for the gradient and Hessian networks using only function evaluations. We provide a statistical learning analysis of the proposed approach, including a bias--variance tradeoff for ZODs. Assuming a standard contraction property of the underlying operator, we establish a non-asymptotic error bound that decomposes the total error into discretization error, approximation error, statistical error, and ZOD bias. Crucially, we derive the sample complexity of the learned representations in (weighted) Sobolev space, characterizing the error up to second-order derivatives. Numerical experiments illustrate the competitive performance of the method in moderate and high dimensions.

cs.LG

Uncertainty quantification using importance-sampled quasi-Monte Carlo with dimension-independent convergence rates

Quasi-Monte Carlo (QMC) integration over unbounded domains $\mathbb{R}^s$ remains challenging due to the high dimensionality of sampling space and the boundary growth of the integrand. In applications such as uncertainty quantification (UQ), the dimension $s$ can reach hundreds or even thousands. To restore the efficiency of quadrature rules in high dimensions, constructive QMC methods like lattice rules have been successfully developed within the framework of weighted function spaces. In contrast to designing problem-specific quadrature points, this paper proposes transforming the underlying integrand to accommodate the off-the-shelf scrambled nets (a construction-free randomized QMC method) via the boundary-damping importance sampling (BDIS) proposed by Pan et al. (2025). We provide a rigorous analysis of the dimension-independent convergence rate of BDIS-based scrambled nets while covering a broader class of unbounded functions than that in Pan et al. (2025). By exploiting the dimension structure of the parametric input random field, the proposed $n$-point quadrature rule achieves a dimension-independent mean squared error rate of $O(n^{-1-α^*+\varepsilon})$ on standard UQ problems in elliptic partial differential equations (PDEs), where $\varepsilon>0$ is arbitrarily small and $α^*\in (0,1)$ reflects the regularity with respect to the parametric variables. Numerical experiments on elliptic PDEs with high-dimensional parameters further demonstrate the effectiveness of the method.

math.NA

Accuracy of Discretely Sampled Stochastic Policies in Continuous-time Reinforcement Learning

Stochastic policies (also known as relaxed controls) are widely used in continuous-time reinforcement learning algorithms. However, executing a stochastic policy and evaluating its performance in a continuous-time environment remain open challenges. This work introduces and rigorously analyzes a policy execution framework that samples actions from a stochastic policy at discrete time points and implements them as piecewise constant controls. We prove that as the sampling mesh size tends to zero, the controlled state process converges weakly to the dynamics with coefficients aggregated according to the stochastic policy. We explicitly quantify the convergence rate based on the regularity of the coefficients and establish an optimal first-order convergence rate for sufficiently regular coefficients. Additionally, we prove a $1/2$-order weak convergence rate that holds uniformly over the sampling noise with high probability, and establish a $1/2$-order pathwise convergence for each realization of the system noise in the absence of volatility control. Building on these results, we analyze the bias and variance of various policy evaluation and policy gradient estimators based on discrete-time observations. Our results provide theoretical justification for the exploratory stochastic control framework in [H. Wang, T. Zariphopoulou, and X.Y. Zhou, J. Mach. Learn. Res., 21 (2020), pp. 1-34].

cs.LG

Quasi-Monte Carlo integration over $\mathbb{R}^s$ with boundary-damping importance sampling

This paper proposes a new importance sampling (IS) that is tailored to quasi-Monte Carlo (QMC) integration over $\mathbb{R}^s$. IS introduces a multiplicative adjustment to the integrand by compensating the sampling from the proposal instead of the target distribution. Improper proposals result in severe adjustment factor for QMC. Our strategy is to first design a adjustment factor to meet desired regularities and then determine a tractable transport map from the standard uniforms to the proposal for using QMC quadrature points as inputs. The transport map has the effect of damping the boundary growth of the resulting integrand so that the effectiveness of QMC can be reclaimed. Under certain conditions on the original integrand, our proposed IS enjoys a fast convergence rate independently of the dimension $s$, making it amenable to high-dimensional problems.

math.NA

Generalization Error Analysis of Deep Backward Dynamic Programming for Solving Nonlinear PDEs

We explore the application of the quasi-Monte Carlo (QMC) method in deep backward dynamic programming (DBDP) (Hure et al. 2020) for numerically solving high-dimensional nonlinear partial differential equations (PDEs). Our study focuses on examining the generalization error as a component of the total error in the DBDP framework, discovering that the rate of convergence for the generalization error is influenced by the choice of sampling methods. Specifically, for a given batch size $m$, the generalization error under QMC methods exhibits a convergence rate of $O(m^{-1+\varepsilon})$, where $\varepsilon$ can be made arbitrarily small. This rate is notably more favorable than that of the traditional Monte Carlo (MC) methods, which is $O(m^{-1/2+\varepsilon})$. Our theoretical analysis shows that the generalization error under QMC methods achieves a higher order of convergence than their MC counterparts. Numerical experiments demonstrate that QMC indeed surpasses MC in delivering solutions that are both more precise and stable.

math.NA

Quasi-Monte Carlo for unbounded integrands with importance sampling

We consider the problem of estimating an expectation $ \mathbb{E}\left[ h(W)\right]$ by quasi-Monte Carlo (QMC) methods, where $ h $ is an unbounded smooth function on $ \mathbb{R}^d $ and $ W$ is a standard normal distributed random variable. To study rates of convergence for QMC on unbounded integrands, we use a smoothed projection operator to project the output of $W$ to a bounded region, which differs from the strategy of avoiding the singularities along the boundary of the unit cube $ [0,1]^d $ in 10.1137/S0036144504441573. The error is then bounded by the quadrature error of the transformed integrand and the projection error. If the function $h(\boldsymbol{x})$ and its mixed partial derivatives do not grow too fast as the Euclidean norm $|\boldsymbol{x}|$ goes to infinity, we obtain an error rate of $O(n^{-1+ε})$ for QMC and randomized QMC (RQMC) with a sample size $n$ and an arbitrarily small $ε>0$. However, the rate turns out to be $O(n^{-1+2M+ε})$ if the functions grow exponentially with a rate of $O(\exp\{M|\boldsymbol{x}|^2\})$ for a constant $M\in(0,1/2)$. Superisingly, we find that using importance sampling with t distribution as the proposal can improve the root mean squared error of RQMC from $O(n^{-1+2M+ε})$ to $O( n^{-3/2+ε})$ for any $M\in(0,1/2)$.

math.NA