SearcharxivSearch

arXiv subjects

Jianya Lu

Publications and source records attributed to Jianya Lu.

8 recordsLinked to original sources

Toward Optimal Statistical Inference in Noisy Linear Quadratic Reinforcement Learning over a Finite Horizon

Recent developments in Reinforcement learning have significantly enhanced sequential decision-making in uncertain environments. Despite their strong performance guarantees, most existing work has focused primarily on improving the operational accuracy of learned control policies and the convergence rates of learning algorithms, with comparatively little attention to uncertainty quantification and statistical inference. Yet, these aspects are essential for assessing the reliability and variability of control policies, especially in high-stakes applications. In this paper, we study statistical inference for the policy gradient (PG) method for noisy Linear Quadratic Reinforcement learning (LQ RL) over a finite time horizon, where linear dynamics with both known and unknown drift parameters are controlled subject to a quadratic cost. We establish the theoretical foundations for statistical inference in LQ RL, deriving exact asymptotics for both the PG estimators and the corresponding objective loss. Furthermore, we introduce a principled inference framework that leverages online bootstrapping to construct confidence intervals for both the learned optimal policy and the corresponding objective losses. The method updates the PG estimates along with a set of randomly perturbed PG estimates as new observations arrive. We prove that the proposed bootstrapping procedure is distributionally consistent and that the resulting confidence intervals achieve both asymptotic and non-asymptotic validity. Notably, our results imply that the quantiles of the exact distribution can be approximated at a rate of $n^{-1/4}$, where $n$ is the number of samples used during the procedure. The proposed procedure is easy to implement and applicable to both offline and fully online settings. Numerical experiments illustrate the effectiveness of our approach across a range of noisy linear dynamical systems.

math.ST

Approximation to Deep Q-Network by Stochastic Delay Differential Equations

Despite the significant breakthroughs that the Deep Q-Network (DQN) has brought to reinforcement learning, its theoretical analysis remains limited. In this paper, we construct a stochastic differential delay equation (SDDE) based on the DQN algorithm and estimate the Wasserstein-1 distance between them. We provide an upper bound for the distance and prove that the distance between the two converges to zero as the step size approaches zero. This result allows us to understand DQN's two key techniques, the experience replay and the target network, from the perspective of continuous systems. Specifically, the delay term in the equation, corresponding to the target network, contributes to the stability of the system. Our approach leverages a refined Lindeberg principle and an operator comparison to establish these results.

cs.LG

Online federated learning framework for classification

In this paper, we develop a novel online federated learning framework for classification, designed to handle streaming data from multiple clients while ensuring data privacy and computational efficiency. Our method leverages the generalized distance-weighted discriminant technique, making it robust to both homogeneous and heterogeneous data distributions across clients. In particular, we develop a new optimization algorithm based on the Majorization-Minimization principle, integrated with a renewable estimation procedure, enabling efficient model updates without full retraining. We provide a theoretical guarantee for the convergence of our estimator, proving its consistency and asymptotic normality under standard regularity conditions. In addition, we establish that our method achieves Bayesian risk consistency, ensuring its reliability for classification tasks in federated environments. We further incorporate differential privacy mechanisms to enhance data security, protecting client information while maintaining model performance. Extensive numerical experiments on both simulated and real-world datasets demonstrate that our approach delivers high classification accuracy, significant computational efficiency gains, and substantial savings in data storage requirements compared to existing methods.

stat.ML

Self-normalized Cram\'er-type Moderate Deviation of Stochastic Gradient Langevin Dynamics

In this paper, we study the self-normalized Cram\'er-type moderate deviation of the empirical measure of the stochastic gradient Langevin dynamics (SGLD). Consequently, we also derive the Berry-Esseen bound for SGLD. Our approach is by constructing a stochastic differential equation (SDE) to approximate the SGLD and then applying Stein's method as developed in [9,19], to decompose the empirical measure into a martingale difference series sum and a negligible remainder term.

math.PR

Distribution estimation and change-point estimation for time series via DNN-based GANs

The generative adversarial networks (GANs) have recently been applied to estimating the distribution of independent and identically distributed data, and have attracted a lot of research attention. In this paper, we use the blocking technique to demonstrate the effectiveness of GANs for estimating the distribution of stationary time series. Theoretically, we derive a non-asymptotic error bound for the Deep Neural Network (DNN)-based GANs estimator for the stationary distribution of the time series. Based on our theoretical analysis, we propose an algorithm for estimating the change point in time series distribution. The two main results are verified by two Monte Carlo experiments respectively, one is to estimate the joint stationary distribution of $5$-tuple samples of a 20 dimensional AR(3) model, the other is about estimating the change point at the combination of two different stationary time series. A real world empirical application to the human activity recognition dataset highlights the potential of the proposed methods.

cs.LG

Almost sure invariance principle of $β-$mixing time series in Hilbert space

Inspired by \citet{Berkes14} and \citet{Wu07}, we prove an almost sure invariance principle for stationary $β-$mixing stochastic processes defined on Hilbert space. Our result can be applied to Markov chain satisfying Meyn-Tweedie type Lyapunov condition and thus generalises the contraction condition in \citet[Example 2.2]{Berkes14}. We prove our main theorem by the big and small blocks technique and an embedding result in \citet{gotze2011estimates}. Our result is further applied to the ergodic Markov chain and functional autoregressive processes.

math.PR

Central limit theorem and Self-normalized Cramér-type moderate deviation for Euler-Maruyama Scheme

We consider a stochastic differential equation and its Euler-Maruyama (EM) scheme, under some appropriate conditions, they both admit a unique invariant measure, denoted by $π$ and $π_η$ respectively ($η$ is the step size of the EM scheme). We construct an empirical measure $Π_η$ of the EM scheme as a statistic of $π_η$, and use Stein's method developed in \citet{FSX19} to prove a central limit theorem of $Π_η$. The proof of the self-normalized Cramér-type moderate deviation (SNCMD) is based on a standard decomposition on Markov chain, splitting $η^{-1/2}(Π_η(.)-π(.))$ into a martingale difference series sum $\mcl H_η$ and a negligible remainder $\mcl R_η$. We handle $\mcl H_η$ by the time-change technique for martingale, while prove that $\mcl R_η$ is exponentially negligible by concentration inequalities, which have their independent interest. Moreover, we show that SNCMD holds for $x = o(η^{-1/6})$, which has the same order as that of the classical result in \citet{shao1999cramer,JSW03}.

math.PR