SearcharxivSearch

arXiv subjects

Zuo Quan Xu

Publications and source records attributed to Zuo Quan Xu.

At least 19 recordsLinked to original sources

Linear-quadratic mixed Stackelberg-zero-sum game for mean-field regime switching system

Motivated by a green finance problem, a linear-quadratic Stackelberg differential game for a regime switching system involving one leader and two followers is studied. The two followers engage in a zero-sum differential game, and both the state system and the cost functional incorporate a conditional mean-field term. Applying continuation method and induction method, we first establish the existence and uniqueness of the conditional mean-field forward-backward stochastic differential equations with Markovian switching. Based on it, we prove the unique solvability of Hamiltonian systems associated with the two followers and the leader. Moreover, utilizing stochastic maximum principle, decoupling approach and optimal filtering technique, we obtain the optimal feedback strategies for the two followers and the leader. Employing the theoretical results, we solve the green finance problem with some numerical simulations.

math.OC

Multi-Asset Liquidation in Dark Pools with Adverse Selection

We study the optimal liquidation of a multi-asset portfolio using both a traditional exchange and dark pools in the presence of quadratic adverse-selection costs. The problem leads to a matrix-valued backward stochastic differential equation with jumps and a singular terminal condition. We establish existence and uniqueness of its solution and use it to characterize the value function and the optimal liquidation strategy. The uniqueness result is the main mathematical contribution and strengthens the existing theory even in simpler special cases; the existence result is also new. For a two-asset model, we distinguish the roles of asset correlation, own-asset adverse selection, and cross-asset spillover in adverse-selection costs. Under diagonal temporary impact and in the absence of cross-asset spillover, an initially well-diversified portfolio remains well diversified during optimal liquidation and, for a fixed sign of the correlation, its liquidation cost is strictly decreasing in the magnitude of the correlation. By contrast, under the same diagonal-impact specification, under explicit conditions and sufficiently close to the liquidation horizon, cross-asset spillover makes a well-diversified portfolio more costly to liquidate than its poorly diversified sign-reversed counterpart and causes sufficiently unbalanced well-diversified portfolios to become poorly diversified with positive probability. Separately, without requiring diagonal temporary impact, we show that, in the absence of cross-asset spillover, own-asset adverse selection introduces an explicit shrinkage factor in the optimal dark-pool order relative to the order minimizing the post-execution continuation value. Finally, we derive an explicit condition under which a dark-pool execution transforms a poorly diversified portfolio into a well-diversified one.

q-fin.MF

Model-Free Q-Learning for Infinite-Horizon Stochastic Linear Quadratic Problems with Regime Switching

This paper addresses infinite-horizon continuous-time stochastic linear quadratic optimal control problems with regime switching. We propose a paradigm shift from model-based design by adopting an adaptive dynamic programming approach, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data. The theoretical core of our work consists of a complete proof of the equivalence between the on- and off-policy architectures, alongside a rigorous analysis establishing the stability of the closed-loop system and the convergence of the algorithms to the optimal solution. For computational tractability, we implement these algorithms using vectorization and Kronecker product algebra. The theoretical results are corroborated by numerical case studies that clearly demonstrate the operational effectiveness and practical feasibility of the proposed model-free control strategy.

math.OC

Inverse Optimal Control for Linear Quadratic Problem with Poisson Jumps: Model-Free Inverse Reinforcement Learning Approaches

This paper addresses the inverse optimal control (IOC) problem for stochastic linear systems subject to both Brownian motion and Poisson jumps, using an inverse reinforcement learning (IRL) framework. Given a target feedback gain from an expert, the objective is to identify an equivalent cost functional-specifically, the set of all cost weights-that yields this same gain. To solve this problem when system dynamics are unknown, we propose two model-free, off-policy IRL algorithms that operate entirely from data, circumventing the need to solve the generalized algebraic Riccati equation or compute the cost weights analytically. The first is an inverse Q-learning algorithm that constructs data-driven equations from expert demonstrations to compute the Q-function matrix, with equivalent cost weights updated algebraically and without requiring additional trajectory data. The second is a model-free off-policy inverse policy iteration algorithm that leverages data collected under an initial stabilizing policy, offering a complementary approach suited to different data availability scenarios. Crucially, by decoupling the data-collection behavior policies from the policies being iteratively updated, both algorithms can learn equivalent cost weights from sufficiently excited trajectories without identifying the system dynamics or jump intensity. Numerical simulations validate the effectiveness of the proposed methods.

math.OC

Stochastic LQ Optimal Control with Random Coefficients and a Terminal Mean-Field Cost

This paper investigates a multidimensional non-homogeneous stochastic linear-quadratic optimal control problem featuring random coefficients and a terminal mean-field term in the cost functional, enabling its direct application to mean-variance models in financial engineering. Employing the Lagrangian duality method together with a decomposition approach for linear backward stochastic differential equations, we provide two types of sufficient conditions for solvability and derive the corresponding optimal controls. In particular, in the deterministic-coefficient case, our condition is weaker than the standard condition found in the existing literature on mean-field stochastic LQ problems. Finally, a numerical example drawn from optimal portfolio selection with multiple assets under mean-variance utility demonstrates the applicability of our results.

math.OC

Dividend ratcheting and capital injection under the Cramér-Lundberg model: Strong solution and optimal strategy

We consider an optimal dividend payout problem for an insurance company whose surplus follows the classical Cramér-Lundberg model. The dividend rate is subject to a ratcheting constraint (i.e., it must be nondecreasing over time), and the company may inject capital at a proportional cost to avoid ruin. This problem gives rise to a stochastic control problem with a self-path-dependent control constraint, costly capital injections, and jump-diffusion dynamics. The associated Hamilton-Jacobi-Bellman (HJB) equation is a partial integro-differential variational inequality featuring both a nonlocal integral term and a gradient constraint. We develop a systematic probabilistic and PDE-based approach to solve this HJB equation. By discretizing the space of admissible dividend rates, we construct a sequence of approximating regime-switching systems of ordinary integro-differential equations. Through careful a priori estimates and a limiting argument, we prove the existence and uniqueness of a \emph{strong solution} in a suitable space. This regularity result is fundamental: it allows us to characterize the optimal dividend policy via a switching free boundary and to construct an explicit optimal feedback control strategy. To the best of our knowledge, this is the first complete solution -- comprising both the value function and an implementable optimal strategy -- for a dividend ratcheting problem with capital injection under the Cramér-Lundberg model. Our work advances the mathematical theory of optimal stochastic control beyond the standard viscosity solution framework, providing a rigorous foundation for dividend policy design in economics.

math.OC

$α$-robust utility maximization with intractable claims: A quantile optimization approach

This paper studies an $α$-robust utility maximization problem where an investor faces an intractable claim -- an exogenous contingent claim with known marginal distribution but unspecified dependence structure with financial market returns. The $α$-robust criterion interpolates between worst-case ($α=0$) and best-case ($α=1$) evaluations, generalizing both extremes through a continuous ambiguity attitude parameter. For weighted exponential utilities, we establish via rearrangement inequalities and comonotonicity theory that the $α$-robust risk measure is law-invariant, depending only on marginal distributions. This transforms the dynamic stochastic control problem into a concave static quantile optimization over a convex domain. We derive optimality conditions via calculus of variations and characterize the optimal quantile as the solution to a two-dimensional first-order ordinary differential equation system, which is a system of variational inequalities with mixed boundary conditions, enabling numerical solution. Our framework naturally accommodates additional risk constraints such as Value-at-Risk and Expected Shortfall. Numerical experiments reveal how ambiguity attitude, market conditions, and claim characteristics interact to shape optimal payoffs.

q-fin.PM

Ergodic McKean-Vlasov Games: Verification Theorems and Linear-Quadratic Applications

This paper investigates two-player ergodic nonzero-sum stochastic differential games with McKean-Vlasov dynamics. We establish a verification theorem connecting solutions of coupled Hamilton-Jacobi-Bellman (HJB) Master equations to Nash equilibria, characterized through an auxiliary control problem defined on the measure space. A key contribution is showing that the value functions are uniquely determined (up to an additive constant) by the uniqueness of the invariant measure of the optimal state process. The theory is applied to Linear-Quadratic-Gaussian (LQG) settings, where explicit solutions to the Master equations are derived by exploiting their polynomial structure in measure variables.

math.OC

Optimal dividend payout with path-dependent drawdown constraint

This paper studies an optimal dividend problem with a drawdown constraint in a Brownian motion model, requiring the dividend payout rate to remain above a fixed proportion of its historical maximum. This leads to a path-dependent stochastic control problem, as the admissible control depends on its own past values. The associated Hamilton-Jacobi-Bellman (HJB) equation is a novel two-dimensional variational inequality with a gradient constraint, a type of problem previously only analyzed in the literature using viscosity solution techniques. In contrast, this paper employs delicate PDE methods to establish the existence of a strong solution. This stronger regularity allows us to explicitly characterize an optimal feedback control strategy, expressed in terms of two free boundaries and the running maximum surplus process. Furthermore, we derive key properties of the value function and the free boundaries, including boundedness and continuity. Numerical examples are provided to verify the theoretical results and to offer new financial insights.

q-fin.MF

De Finetti's problem with fixed transaction costs and regime switching

In this paper, we examine a modified version of de Finetti's optimal dividend problem, incorporating fixed transaction costs and altering the surplus process by introducing two-valued drift and two-valued volatility coefficients. This modification aims to capture the transitions or adjustments in the company's financial status. We identify the optimal dividend strategy, which maximizes the expected total net dividend payments (after accounting for transaction costs) until ruin, as a two-barrier impulsive dividend strategy. Notably, the optimal strategy can be explicitly determined for almost all scenarios involving different drifts and volatility coefficients. Our primary focus is on exploring how changes in drift and volatility coefficients influence the optimal dividend strategy.

q-fin.MF

Competitive optimal portfolio selection under mean-variance criterion

We investigate a portfolio selection problem involving multi competitive agents, each exhibiting mean-variance preferences. Unlike classical models, each agent's utility is determined by their relative wealth compared to the average wealth of all agents, introducing a competitive dynamic into the optimization framework. To address this game-theoretic problem, we first reformulate the mean-variance criterion as a constrained, non-homogeneous stochastic linear-quadratic control problem and derive the corresponding optimal feedback strategies. The existence of Nash equilibria is shown to depend on the well-posedness of a complex, coupled system of equations. Employing decoupling techniques, we reduce the well-posedness analysis to the solvability of a novel class of multi-dimensional linear backward stochastic differential equations (BSDEs). We solve a new type of nonlinear BSDEs (including the above linear one as a special case) using fixed-point theory. Depending on the interplay between market and competition parameters, three distinct scenarios arise: (i) the existence of a unique Nash equilibrium, (ii) the absence of any Nash equilibrium, and (iii) the existence of infinitely many Nash equilibria. These scenarios are rigorously characterized and discussed in detail.

math.OC

Learning to Optimally Stop Diffusion Processes, with Financial Applications

We study optimal stopping for diffusion processes with unknown model primitives within the continuous-time reinforcement learning (RL) framework developed by Wang et al. (2020), and present applications to option pricing and portfolio choice. By penalizing the corresponding variational inequality formulation, we transform the stopping problem into a stochastic optimal control problem with two actions. We then randomize controls into Bernoulli distributions and add an entropy regularizer to encourage exploration. We derive a semi-analytical optimal Bernoulli distribution, based on which we devise RL algorithms using the martingale approach established in Jia and Zhou (2022a). We establish a policy improvement theorem and prove the fast convergence of the resulting policy iterations. We demonstrate the effectiveness of the algorithms in pricing finite-horizon American put options, solving Merton's problem with transaction costs, and scaling to high-dimensional optimal stopping problems. In particular, we show that both the offline and online algorithms achieve high accuracy in learning the value functions and characterizing the associated free boundaries.

math.OC

Optimal control of stochastic homogenous systems

This paper investigates a new class of homogeneous stochastic control problems with cone control constraints, extending the classical homogeneous stochastic linear-quadratic (LQ) framework to encompass nonlinear system dynamics and non-quadratic cost functionals. We demonstrate that, analogous to the LQ case, the optimal controls and value functions for these generalized problems are intimately connected to a novel class of highly nonlinear backward stochastic differential equations (BSDEs). We establish the existence and uniqueness of solutions to these BSDEs under three distinct sets of conditions, employing techniques such as truncation functions and logarithmic transformations. Furthermore, we derive explicit feedback representations for the optimal controls and value functions in terms of the solutions to these BSDEs, supported by rigorous verification arguments. Our general solvability conditions allow us to recover many known results for homogeneous LQ problems, including both standard and singular cases, as special instances of our framework.

math.OC

Optimal mean-variance portfolio selection under regime-switching-induced stock price shocks

In this paper, we investigate mean-variance (MV) portfolio selection problems with jumps in a regime-switching financial model. The novelty of our approach lies in allowing not only the market parameters -- such as the interest rate, appreciation rate, volatility, and jump intensity -- to depend on the market regime, but also in permitting stock prices to experience jumps when the market regime switches, in addition to the usual micro-level jumps. This modeling choice is motivated by empirical observations that stock prices often exhibit sharp declines when the market shifts from a ``bullish'' to a ``bearish'' regime, and vice versa. By employing the completion-of-squares technique, we derive the optimal portfolio strategy and the efficient frontier, both of which are characterized by three systems of multi-dimensional ordinary differential equations (ODEs). Among these, two systems are linear, while the first one is an $\ell$-dimensional, fully coupled, and highly nonlinear Riccati equation. In the absence of regime-switching-induced stock price shocks, these systems reduce to simple linear ODEs. Thus, the introduction of regime-switching-induced stock price shocks adds significant complexity and challenges to our model. Additionally, we explore the MV problem under a no-shorting constraint. In this case, the corresponding Riccati equation becomes a $2\ell$-dimensional, fully coupled, nonlinear ODE, for which we establish solvability. The solution is then used to explicitly express the optimal portfolio and the efficient frontier.

q-fin.PM

Infinite horizon discounted LQ optimal control problems for mean-field switching diffusions

This paper investigates an infinite horizon discounted linear-quadratic (LQ) optimal control problem for stochastic differential equations (SDEs) incorporating regime switching and mean-field interactions. The regime switching is modeled by a finite-state Markov chain acting as common noise, while the mean-field interactions are characterized by the conditional expectation of the state process given the history of the Markov chain. To address system stability in the infinite horizon setting, a discounted factor is introduced. Within this framework, the well-posedness of the state equation and adjoint equation -- formulated as infinite horizon mean-field forward and backward SDEs with Markov chains, respectively -- is established, along with the asymptotic behavior of their solutions as time approaches infinity. A candidate optimal feedback control law is formally derived based on two algebraic Riccati equations (AREs), which are introduced for the first time in this context. The solvability of these AREs is proven through an approximation scheme involving a sequence of Lyapunov equations, and the optimality of the proposed feedback control law is rigorously verified using the completion of squares method. Finally, numerical experiments are conducted to validate the theoretical findings, including solutions to the AREs, the optimal control process, and the corresponding optimal (conditional) state trajectory. This work provides a comprehensive framework for solving infinite horizon discounted LQ optimal control problems in the presence of regime switching and mean-field interactions, offering both theoretical insights and practical computational tools.

math.OC

A System of BSDEs with Singular Terminal Values Arising in Optimal Liquidation with Regime Switching

We study a stochastic control problem with regime switching arising in an optimal liquidation problem with dark pools and multiple regimes. The new feature of this model is that it introduces a system of BSDEs with jumps and with singular terminal values, which appears in literature for the first time. The existence result for this system is obtained. As a result, we solve the stochastic control problem with regime switching. More importantly, the uniqueness result of this system is also obtained, in contrast to merely minimal solutions established in most related literature.

q-fin.MF

Constrained stochastic linear quadratic control under regime switching with controlled jump size

In this paper, we examine a stochastic linear-quadratic control problem characterized by regime switching and Poisson jumps. All the coefficients in the problem are random processes adapted to the filtration generated by Brownian motion and the Poisson random measure for each given regime. The model incorporates two distinct types of controls: the first is a conventional control that appears in the continuous diffusion component, while the second is an unconventional control, dependent on the variable $z$, which influences the jump size in the jump diffusion component. Both controls are constrained within general closed cones. By employing the Meyer-Itô formula in conjunction with a generalized squares completion technique, we rigorously and explicitly derive the optimal value and optimal feedback control. These depend on solutions to certain multi-dimensional fully coupled stochastic Riccati equations, which are essentially backward stochastic differential equations with jumps (BSDEJs). We establish the existence of a unique nonnegative solution to the BSDEJs. One of the major tools used in the proof is the newly established comparison theorems for multidimensional BSDEJs.

math.OC

Stochastic optimal self-path-dependent control: A new type of variational inequality and its viscosity solution

In this paper, we explore a new class of stochastic control problems characterized by specific control constraints. Specifically, the admissible controls are subject to the ratcheting constraint, meaning they must be non-decreasing over time and are thus self-path-dependent. This type of problems is common in various practical applications, such as optimal consumption problems in financial engineering and optimal dividend payout problems in actuarial science. Traditional stochastic control theory does not readily apply to these problems due to their unique self-path-dependent control feature. To tackle this challenge, we introduce a new class of Hamilton-Jacobi-Bellman (HJB) equations, which are variational inequalities concerning the derivative of a new spatial argument that represents the historical maximum control value. Under the standard Lipschitz continuity condition, we demonstrate that the value functions for these self-path-dependent control problems are the unique solutions to their corresponding HJB equations in the viscosity sense.

math.OC