SearcharxivSearch

arXiv subjects

Yuchao Dong

Publications and source records attributed to Yuchao Dong.

At least 19 recordsLinked to original sources

Randomized Optimal Switching Problem and Related Mirror Descent Flow

We study continuous-time reinforcement learning for the optimal switching problem, in which a decision-maker controls a diffusion process by switching among finitely many regimes, incurring both running and transition costs. To enable exploration, we relax the classical deterministic switching control to a randomized framework, where the switching decisions are governed by a continuous-time Markov chain with state-dependent generator, and augment the cost functional with a KL-divergence regularization weighted by a temperature parameter $λ$. Under mild assumptions on the coefficients, we establish that the regularized value function is the unique smooth solution of an elliptic Hamilton--Jacobi--Bellman system, and derive an explicit optimal Gibbs policy given by an exponential transformation of the value function differences across modes. We further prove that the regularized value function approximates the classical optimal value function with error of order $O\left(λ\log \frac{1}λ\right)$, which is consistent with analogous bounds established in other entropy-regularized control problems and is believed to be sharp. To solve the regularized problem numerically, we introduce a mirror descent flow in the dual logarithmic policy space, prove its well-posedness and the monotonic decrease of the value function along the flow, and establish quantitative error bound to the classical optimal value function. For a constant temperature scheduler, the convergence rate is of order $O\left(\frac{1}{e^{λs} - 1}+λ\log\frac1λ\right)$, while under the annealing scheduler $λ_s = \frac{1}{\sqrt{1+s}}$, we obtain the rate $O\left(\frac{\log s}{\sqrt{s}}\right)$, which decays to zero as the flow time $s \to \infty$.

math.OC

A Two-fold Randomization Framework for Impulse Control Problems

We propose and analyze a randomization scheme for a general class of impulse control problems. The solution to this randomized problem is characterized as the fixed point of a compound operator which consists of a regularized nonlocal operator and a regularized stopping operator. This approach allows us to derive a semi-linear Hamilton-Jacobi-Bellman (HJB) equation. Through an equivalent randomization scheme with a Poisson compound measure, we establish a verification theorem that implies the uniqueness of the solution. Via an iterative approach, we prove the existence of the solution. The existence-and-uniqueness result ensures the randomized problem is well-defined. We then demonstrate that our randomized impulse control problem converges to its classical counterpart as the randomization parameter $\pmb λ$ vanishes. This convergence, combined with the value function's $C^{2,α}_{loc}$ regularity, confirms our framework provides a robust approximation and a foundation for developing learning algorithms. Under this framework, we propose an offline reinforcement learning (RL) algorithm. Its policy improvement step is naturally derived from the iterative approach from the existence proof, which enjoys a geometric convergence rate. We implement a model-free version of the algorithm and numerically demonstrate its effectiveness using a widely-studied example. The results show that our RL algorithm can learn the randomized solution, which accurately approximates its classical counterpart. A sensitivity analysis with respect to the volatility parameter $σ$ in the state process effectively demonstrates the exploration-exploitation tradeoff.

math.OC

Data-Driven Merton's Strategies via Policy Randomization

We study Merton's expected utility maximization problem in an incomplete market, characterized by a factor process in addition to the stock price process, where all the model primitives are unknown. The agent under consideration is a price taker who has access only to the stock and factor value processes and the instantaneous volatility. We propose an auxiliary problem in which the agent can invoke policy randomization according to a specific class of Gaussian distributions, and prove that the mean of its optimal Gaussian policy solves the original Merton problem. With randomized policies, we are in the realm of continuous-time reinforcement learning (RL) recently developed in Wang et al. (2020) and Jia and Zhou (2022a, 2022b, 2023), enabling us to solve the auxiliary problem in a data-driven way without having to estimate the model primitives. Specifically, we establish a policy improvement theorem based on which we design both online and offline actor-critic RL algorithms for learning Merton's strategies. A key insight from this study is that RL in general and policy randomization in particular are useful beyond the purpose for exploration -- they can be employed as a technical tool to solve a problem that cannot be otherwise solved by mere deterministic policies. At last, we carry out both simulation and empirical studies in a stochastic volatility environment to demonstrate the decisive outperformance of the devised RL algorithms in comparison to the conventional model-based, plug-in method.

q-fin.PM

Merton's Problem with Recursive Perturbed Utility

The classical Merton investment problem predicts deterministic, state-dependent portfolio rules; however, laboratory and field evidence suggests that individuals often prefer randomized decisions leading to stochastic and noisy choices. Fudenberg et al. (2015) develop the additive perturbed utility theory to explain the preference for randomization in the static setting, which, however, becomes ill-posed or intractable in the dynamic setting. We introduce the recursive perturbed utility (RPU), a special stochastic differential utility that incorporates an entropy-based preference for randomization into a recursive aggregator. RPU endogenizes the intertemporal trade-off between utilities from randomization and bequest via a discounting term dependent on past accumulated randomization, thereby avoiding excessive randomization and yielding a well-posed problem. In a general Markovian incomplete market with CRRA preferences, we prove that the RPU-optimal portfolio policy (in terms of the risk exposure ratio) is Gaussian and can be expressed in closed form, independent of wealth. Its variance is inversely proportional to risk aversion and stock volatility, while its mean is based on the solution to a partial differential equation. Moreover, the mean is the sum of a myopic term and an intertemporal hedging term (against market incompleteness) that intertwines with policy randomization. Finally, we carry out an asymptotic expansion in terms of the perturbed utility weight to show that the optimal mean policy deviates from the classical Merton policy at first order, while the associated relative wealth loss is of a higher order, quantifying the financial cost of the preference for randomization.

q-fin.MF

Extended HJB Equation for Mean-Variance Stopping Problem: Vanishing Regularization Method

This paper studies the time-inconsistent MV optimal stopping problem via a game-theoretic approach to find equilibrium strategies. To overcome the mathematical intractability of direct equilibrium analysis, we propose a vanishing regularization method: first, we introduce an entropy-based regularization term to the MV objective, modeling mixed-strategy stopping times using the intensity of a Cox process. For this regularized problem, we derive a coupled extended Hamilton-Jacobi-Bellman (HJB) equation system, prove a verification theorem linking its solutions to equilibrium intensities, and establish the existence of classical solutions for small time horizons via a contraction mapping argument. By letting the regularization term tend to zero, we formally recover a system of parabolic variational inequalities that characterizes equilibrium stopping times for the original MV problem. This system includes an additional key quadratic term--a distinction from classical optimal stopping, where stopping conditions depend only on comparing the value function to the instantaneous reward.

math.OC

Optimal Carbon Emission Control With Allowances Purchasing

In this paper, we consider a company can simultaneously reduce its emissions and buy carbon allowances at any time. We establish an optimal control model involving two stochastic processes with two control variables, which is a singular control problem. This model can then be converted into a Hamilton-Jacobi-Bellman (HJB) equation, which is a two-dimensional variational equality with gradient barrier, so that the free boundary is a surface. We prove the existence and uniqueness of the solution. Finally, some numerical results are shown.

math.OC

Stability of traveling wave solutions in a credit rating migration Free Boundary Problem

In this paper, we study the stability of traveling wave solutions arising from a credit rating migration problem with a free boundary, After some transformations, we turn the Free Boundary Problem into a fully nonlinear parabolic problem on a fixed domain and establish a rigorous stability analysis of the equilibrium in an exponentially weighted function space. It implies the convergence of the discounted value of bonds that stands as an attenuated traveling wave solution.

math.AP

Randomized Optimal Stopping Problem in Continuous time and Reinforcement Learning Algorithm

In this paper, we study the optimal stopping problem in the so-called exploratory framework, in which the agent takes actions randomly conditioning on current state and an entropy-regularized term is added to the reward functional. Such a transformation reduces the optimal stopping problem to a standard optimal control problem. We derive the related HJB equation and prove its solvability. Furthermore, we give a convergence rate of policy iteration and the comparison to classical optimal stopping problem. Based on the theoretical analysis, a reinforcement learning algorithm is designed and numerical results are demonstrated for several models.

math.OC

Double free boundary problem for defaultable corporate bond with credit rating migration risks and their asymptotic behaviors

In this work, a pricing model for a defaultable corporate bond with credit rating migration risk is established. The model turns out to be a free boundary problem with two free boundaries. The latter are the level sets of the solution but of different kinds. One is from the discontinuous second order term, the other from the obstacle. Existence, uniqueness, and regularity of the solution are obtained. We also prove that two free boundaries are $C^\infty$. The asymptotic behavior of the solution is also considered: we show that it converges to a traveling wave solution when time goes to infinity. Moreover, numerical results are presented.

math.AP

The Relationship between Maximum Principle and Dynamic Programming Principle for Stochastic Recursive Control Problem with Random Coefficients

This paper aims to explore the relationship between maximum principle and dynamic programming principle for stochastic recursive control problem with random coefficients. Under certain regular conditions for the coefficients, the relationship between the Hamilton system with random coefficients and stochastic Hamilton-Jacobi-Bellman equation is obtained. It is very different from the deterministic coefficients case since stochastic Hamilton-Jacobi-Bellman equation is a backward stochastic partial differential equation with solution being a pair of random fields rather than a deterministic function. A linear quadratic recursive utility optimization problem is given as an explicitly illustrated example based on this kind of relationship.

math.OC

Optimal controls of stochastic differential equations with jumps and random coefficients: Stochastic Hamilton-Jacobi-Bellman equations with jumps

In this paper, we study the following nonlinear backward stochastic integral partial differential equation with jumps \begin{equation*} \left\{ \begin{split} -d V(t,x) =&\displaystyle\inf_{u\in U}\bigg\{H(t,x,u, DV(t,x),D Φ(t,x), D^2 V(t,x),\int_E \left(\mathcal I V(t,e,x,u)+Ψ(t,x+g(t,e,x,u))\right)l(t,e)ν(de)) \\ &+\displaystyle\int_{E}\big[\mathcal I V(t,e,x,u)-\displaystyle (g(t, e,x,u), D V(t,x))\big]ν(d e)+\int_{E}\big[\mathcal I Ψ(t,e,x,u)\big]ν(d e)\bigg\}dt\\ &-Φ(t,x)dW(t)-\displaystyle\int_{E} Ψ(t, e,x)\tildeμ(d e,dt),\\ V(T,x)=& \ h(x), \end{split} \right. \end{equation*} where $\tilde μ$ is a Poisson random martingale measure, $W$ is a Brownian motion, and $\mathcal I$ is a non-local operator to be specified later. The function $H$ is a given random mapping, which arises from a corresponding non-Markovian optimal control problem. This equation appears as the stochastic Hamilton-Jacobi-Bellman equation, which characterizes the value function of the optimal control problem with a recursive utility cost functional. The solution to the equation is a predictable triplet of random fields $(V,Φ,Ψ)$. We show that the value function, under some regularity assumptions, is the solution to the stochastic HJB equation; and a classical solution to this equation is the value function and gives the optimal control. With some additional assumptions on the coefficients, an existence and uniqueness result in the sense of Sobolev space is shown by recasting the backward stochastic partial integral differential equation with jumps as a backward stochastic evolution equation in Hilbert spaces with Poisson jumps.

math.OC

Weak Limits of Random Coefficient Autoregressive Processes and their Application in Ruin Theory

We prove that a large class of discrete-time insurance surplus processes converge weakly to a generalized Ornstein-Uhlenbeck process, under a suitable re-normalization and when the time-step goes to 0. Motivated by ruin theory, we use this result to obtain approximations for the moments, the ultimate ruin probability and the discounted penalty function of the discrete-time process.

math.PR

Backward Stochastic Riccati Equation with Jumps associated with Stochastic Linear Quadratic Optimal Control with Jumps and Random Coefficients

In this paper, we investigate the solvability of matrix valued Backward stochastic Riccati equations with jumps (BSREJ), which is associated with a stochastic linear quadratic (SLQ) optimal control problem with random coefficients and driven by both Brownian motion and Poisson jumps. By dynamic programming principle, Doob-Meyer decomposition and inverse flow technique, the existence and uniqueness of the solution for the BSREJ is established. The difficulties addressed to this issue not only are brought from the high nonlinearity of the generator of the BSREJ like the case driven only by Brownian motion, but also from that i) the inverse flow of the controlled linear stochastic differential equation driven by Poisson jumps may not exist without additional technical condition, and ii) how to show the inverse matrix term involving jump process in the generator is well-defined. Utilizing the structure of the optimal problem, we overcome these difficulties and establish the existence of the solution. In additional, a verification theorem for BSREJ is given which implies the uniqueness of the solution.

math.OC

Utility maximization for L{é}vy switching models

This article is devoted to the maximisation of HARA utilities of L{é}vy switching process on finite time interval via dual method. We give the description of all f-divergence minimal martingale measures in initially enlarged filtration, the expression of their Radon-Nikodym densities involving Hellinger and Kulback-Leibler processes, the expressions of the optimal strategies in progressively enlarged filtration for the maximisation of HARA utilities as well as the values of the corresponding maximal expected utilities. The example of Brownian switching model is presented to give the financial interpretation of the results.

math.PR

The Obstacle Problem for Quasilinear Stochastic PDEs with Neumann boundary condition

We prove the existence and uniqueness of solution of the obstacle problem for quasilinear stochastic partial differential equations (OSPDEs for short) with Neumann boundary condition. Our method is based on the analytical technics coming from parabolic potential theory. The solution is expressed as a pair $(u,ν)$ where $u$ is a predictable continuous process which takes values in a proper Sobolev space and $ν$ is a random regular measure satisfying minimal Skohorod condition.

math.PR

Second-Order Necessary Conditions for Optimal Control with Recursive Utilities

The necessary conditions for an optimal control of a stochastic control problem with recursive utilities is investigated. The first order condition is the the well-known Pontryagin type maximum principle. When the optimal control satisfying such first-order necessary condition is singular in some sense, certain type of the second-order necessary condition will come in naturally. The aim of this paper is to explore such kind of conditions for our optimal control problem.

math.OC

Constrained LQ problem with a random jump and application to portfolio selection

In this paper, we consider a constrained stochastic linear-quadratic (LQ) optimal control problem where the control is constrained in a closed cone. The state process is governed by a controlled SDE with random coefficients. Moreover, there is a random jump of the state process. In mathematical finance, the random jump often represents the default of a counter party. Thanks to the Itô-Tanaka formula, optimal control and optimal value can be obtained by solutions of a system of backward stochastic differential equations (BSDEs). The solvability of the BSDEs is obtained by solving a recursive system of BSDEs driven by the Brownian motions. We also apply the result to the mean variance portfolio selection problem in which the stock price can be affected by the default of a counterparty.

math.OC