SearcharxivSearch

arXiv subjects

Guangyan Jia

Publications and source records attributed to Guangyan Jia.

18 recordsLinked to original sources

A note on convergence rate for reflected BSDEs with quadratic generators by penalization method

In this paper, we study the convergence rate between reflected backward stochastic differential equations with quadratic generators and their penalized BSDEs. Using techniques of BMO martingales, we prove the convergence rate is at order $\frac{1}{2}$ as a function of the penalty parameter. Finally, the result is applied to study numerical approximation of reflected BSDEs with sub-quadratic generators by the Euler's polygonal line method.

math.PR

New Second-order Convergent Schemes for Solving decoupled FBSDEs

This paper proposes a new second-order symmetric algorithm for solving decoupled forward-backward stochastic differential equations. Inspired by the alternating direction implicit splitting method for partial differential equations, we split the generator into the sum of two functions. In the computation of the value process Y, explicit and implicit schemes are alternately applied to these two generators, while the algorithms from \citep{ZhaoLi2014} are used for the control process Z. We rigorously prove that the two new schemes have second-order convergence rate. The proposed splitting methods show clear advantages for equations whose generator consists of a linear part plus a nonlinear part, as they reduce the number of iterations required for solving implicit schemes, thereby decreasing computational cost while maintaining second-order convergence. Two numerical examples are provided, including the backward stochastic Riccati equation arising in mean-variance hedging. The numerical results verify the theoretical error analysis and demonstrate the advantage of reduced computational cost compared to the algorithm in \citep{ZhaoLi2014}.

math.NA

Computing stabilizing feedback gains for stochastic linear systems via policy iteration method

In recent years, stabilizing unknown dynamical systems has became a critical problem in control systems engineering. Addressing this for linear time-invariant (LTI) systems is an essential fist step towards solving similar problems for more complex systems. In this paper, we develop a model-free reinforcement learning algorithm to compute stabilizing feedback gains for stochastic LTI systems with unknown system matrices. This algorithm proceeds by solving a series of discounted stochastic linear quadratic (SLQ) optimal control problems via policy iteration (PI). And the corresponding discount factor gradually decreases according to an explicit rule, which is derived from the equivalent condition in verifying the stabilizability. We prove that this method can return a stabilizer after finitely many steps. Finally, a numerical example is provided to illustrate the effectiveness of the proposed method.

math.OC

Quadratic BSDEs with Singular Generators and Unbounded Terminal Conditions: Theory and Applications

We investigate a class of quadratic backward stochastic differential equations (BSDEs) with generators singular in $ y $. First, we establish the existence of solutions and a comparison theorem, thereby extending results in the literature. Additionally, we analyze the stability property and the Feynman-Kac formula, and prove the uniqueness of viscosity solutions for the corresponding singular semilinear partial differential equations (PDEs). Finally, we demonstrate applications in the context of robust control linked to stochastic differential utility and certainty equivalent based on $g$-expectation. In these applications, the coefficient of the quadratic term in the generator captures the level of ambiguity aversion and the coefficient of absolute risk aversion, respectively.

math.PR

Robust policy iteration for continuous-time stochastic $H_\infty$ control problem with unknown dynamics

In this article, we study a continuous-time stochastic $H_\infty$ control problem based on reinforcement learning (RL) techniques that can be viewed as solving a stochastic linear-quadratic two-person zero-sum differential game (LQZSG). First, we propose an RL algorithm that can iteratively solve stochastic game algebraic Riccati equation based on collected state and control data when all dynamic system information is unknown. In addition, the algorithm only needs to collect data once during the iteration process. Then, we discuss the robustness and convergence of the inner and outer loops of the policy iteration algorithm, respectively, and show that when the error of each iteration is within a certain range, the algorithm can converge to a small neighborhood of the saddle point of the stochastic LQZSG problem. Finally, we applied the proposed RL algorithm to two simulation examples to verify the effectiveness of the algorithm.

math.OC

Inverse reinforcement learning by expert imitation for the stochastic linear-quadratic optimal control problem

This article studies inverse reinforcement learning (IRL) for the stochastic linear-quadratic optimal control problem, where two agents are considered. A learner agent does not know the expert agent's performance cost function, but it imitates the behavior of the expert agent by constructing an underlying cost function that obtains the same optimal feedback control as the expert's. We first develop a model-based IRL algorithm, which consists of a policy correction and a policy update from the policy iteration in reinforcement learning, as well as a cost function weight reconstruction based on the inverse optimal control. Then, under this scheme, we propose a model-free off-policy IRL algorithm, which does not need to know or identify the system and only needs to collect the behavior data of the expert agent and learner agent once during the iteration process. Moreover, the proofs of the algorithm's convergence, stability, and non-unique solutions are given. Finally, a simulation example is provided to verify the effectiveness of the proposed algorithm.

math.OC

Convergence of Policy Gradient for Stochastic Linear-Quadratic Control Problem in Infinite Horizon

With the outstanding performance of policy gradient (PG) method in the reinforcement learning field, the convergence theory of it has aroused more and more interest recently. Meanwhile, the significant importance and abundant theoretical researches make the stochastic linear quadratic (SLQ) control problem a starting point for studying PG in model-based learning setting. In this paper, we study the PG method for the SLQ problem in infinite horizon and take a step towards providing rigorous guarantees for gradient methods. Although the cost functional of linear-quadratic problem is typically nonconvex, we still overcome the difficulty based on gradient domination condition and L-smoothness property, and prove exponential/linear convergence of gradient flow/descent algorithm.

math.OC

On the Correspondence and the Risk Contribution for Conditional Coherent and Deviation Risk Measures

We give an axiomatic framework for conditional generalized deviation measures. Under financially reasonable assumptions, we give the correspondence between conditional coherent risk measures and generalized deviation measures. Moreover, we establish the notion of continuous-time risk contribution for conditional coherent risk measures and generalized deviation measures. With the help of the correspondence between these two different types of risk measures, we give a microscopic interpretation of their risk contributions. Particularly, we show that the risk contributions of time-consistent risk measures are still time-consistent. We also demonstrate that the second element of the BSDE solution $(Y, Z)$ associated with $g$-expectation has the meaning of risk contribution.

q-fin.RM

Continuous-Time Risk Contribution of the Terminal Variance and its Related Risk Budgeting Problem

To achieve robustness of risk across different assets, risk parity investing rules, a particular state of risk contributions, have grown in popularity over the previous few decades. To generalize the concept of risk contribution from the simple covariance matrix case to the continuous-time case in which the terminal variance of wealth is used as the risk measure, we characterize risk contributions and marginal risk contributions on various assets as predictable processes using the Gateaux differential and Doleans measure. Meanwhile, the risk contributions we extend here have the aggregation property, namely that total risk can be represented as the aggregation of those among different assets and $(t,ω)$. Subsequently, as an inverse target -- allocating risk, the risk budgeting problem of how to obtain policies whose risk contributions coincide with pre-given risk budgets in the continuous-time case is also explored in this paper. These policies are solutions to stochastic convex optimizations parametrized by the pre-given risk budgets. Moreover, single-period risk budgeting policies are explained as the projection of risk budgeting policies in continuous-time cases. On the application side, volatility-managed portfolios in [Moreira and Muir,2017] can be obtained by risk budgeting optimization; similarly to previous findings, continuous-time mean-variance allocation in [Zhou and Li, 2000] appears to be concentrated in terms of risk contribution.

q-fin.MF

Monetary Risk Measures

In this paper, we study general monetary risk measures (without any convexity or weak convexity). A monetary (respectively, positively homogeneous) risk measure can be characterized as the lower envelope of a family of convex (respectively, coherent) risk measures. The proof does not depend on but easily leads to the classical representation theorems for convex and coherent risk measures. When the law-invariance and the SSD (second-order stochastic dominance)-consistency are involved, it is not the convexity (respectively, coherence) but the comonotonic convexity (respectively, comonotonic coherence) of risk measures that can be used for such kind of lower envelope characterizations in a unified form. The representation of a law-invariant risk measure in terms of VaR is provided.

q-fin.MF

A note on characterizations of G-normal distribution

In this paper, we show that the G-normality of X and Y can be characterized according to the form of f such that the distribution of λ+f(λ)Y does not depend on λ, where Y is an independent copy of X and λ is in the domain of f. Without the condition that Y is identically distributed with X, we still have a similar argument.

math.PR

Invariant representation for stochastic differential operator by BSDEs with uniformly continuous coefficients and its applications

In this paper, we prove that a kind of second order stochastic differential operator can be represented by the limit of solutions of BSDEs with uniformly continuous coefficients. This result is a generalization of the representation for the uniformly continuous generator. With the help of this representation, we obtain the corresponding converse comparison theorem for the BSDEs with uniformly continuous coefficients, and get some equivalent relationships between the properties of the generator $g$ and the associated solutions of BSDEs. Moreover, we give a new proof about $g$-convexity.

math.PR

A uniqueness theorem for solution of BSDEs

In this note, we prove that if $g$ is uniformly continuous in $z$, uniformly with respect to $(\oo,t)$ and independent of $y$, the solution to the backward stochastic differential equation (BSDE) with generator $g$ is unique.

math.PR

Jensen's Inequality for g-Convex Function under g-Expectation

A real valued function defined on}$\mathbb{R}$ {\small is called}$g${\small --convex if it satisfies the following \textquotedblleft generalized Jensen's inequality\textquotedblright under a given}$g${\small -expectation, i.e., }$h(\mathbb{E}^{g}[X])\leq \mathbb{E}% ^{g}[h(X)]${\small, for all random variables}$X$ {\small such that both sides of the inequality are meaningful. In this paper we will give a necessary and sufficient conditions for a }$C^{2}${\small -function being}$% g ${\small -convex. We also studied some more general situations. We also studied}$g${\small -concave and}$g${\small -affine functions.

math.PR