SearcharxivSearch

arXiv subjects

Sebastien Lleo

Publications and source records attributed to Sebastien Lleo.

8 recordsLinked to original sources

Reinforcement Learning for Risk-Sensitive Investment Management: a Free Energy--Entropy Duality Approach

This paper develops a reinforcement-learning approach to continuous-time risk-sensitive benchmarked asset allocation in a partly model-based setting. The benchmarked problem does not directly fit the standard Markovian stochastic-control template: the state is uncontrolled, whereas the terminal reward contains a controlled It\^o integral. We use free energy-entropy duality to reformulate the problem as a linear-quadratic-Gaussian stochastic differential game under an equivalent probability measure, yielding explicit finite- and infinite-horizon saddle-point solutions. This structure guides a continuous-time $q$-learning actor-critic method: the quadratic value function motivates the critic, while the affine saddle-point controls motivate deterministic actors for the portfolio allocation and adversarial control. The learned allocation admits an economic interpretation through fractional Kelly decompositions. A proof-of-concept implementation calibrated to U.S. equity data shows that the actors learn the optimal policy with high accuracy and reveals a favorable asymmetry: the portfolio actor receives a cleaner learning signal than the auxiliary adversarial actor.

q-fin.PM

Risk-Sensitive Investment Management via Free Energy-Entropy Duality

We study a benchmarked risk-sensitive portfolio problem in a factor-based setting to bring together three strands of the literature: benchmarked risk-sensitive investment management, the Kuroda-Nagai change-of-measure method, and the free energy-entropy duality of Dai Pra et al. (1996). We show that the duality yields a direct solution of the benchmarked problem by reformulating it as a linear-quadratic-Gaussian stochastic differential game under a suitable equivalent probability measure, with an entropic regularization. The resulting value function is quadratic, the optimal controls are explicit affine feedback maps, and the optimal allocation admits two complementary interpretations: as a fractional Kelly strategy and as a Kelly portfolio adjusted via the entropic regularization. This formulation, therefore, contributes both a direct analytical route to the solution and a clearer interpretation of risk sensitivity, thereby embedding the classical Kuroda-Nagai change-of-measure approach within a more general framework. An added benefit of this formulation is that it is suitable for implementation via an RL algorithm. A simple implementation on U.S. equity data illustrates the tractability of the framework and numerically confirms the equivalence of the two approaches.

q-fin.PM

Exploratory Randomization for Discrete-Time Risk-Sensitive Benchmarked Investment Management with Reinforcement Learning

This paper bridges reinforcement learning (RL) and risk-sensitive stochastic control by introducing a tractable exploration mechanism for policy search in risk-sensitive portfolio management, with known and unknown model parameters, that yields an endogenous relative-entropy regularization. We construct a discrete-time risk-sensitive benchmarked investment model. This model combines a factor-based asset universe with periodic portfolio rebalancing. Exploration is incorporated through user-specified Gaussian perturbations to baseline (exploitative) controls. The risk-sensitive stochastic control problem is solved analytically using the Free Energy-Entropy Duality. The Duality recasts the control problem as a linear-quadratic-Gaussian game and introduces a natural penalty for exploration. This approach yields simple sufficiency conditions for optimality. It also induces intuitive bounds on exploration based on risk sensitivity, asset covariance, and rebalancing frequency. Additionally, the optimal investment strategy can be interpreted through the lens of fractional Kelly strategies. By connecting risk-sensitive control theory and RL, this work provides a principled parametric family for policy-gradient implementations, guiding the design of RL methods.

q-fin.PM

Exploratory Randomization for Discrete-Time Linear Exponential Quadratic Gaussian (LEQG) Problem

We investigate exploratory randomization for an extended linear-exponential-quadratic-Gaussian (LEQG) control problem in discrete time. This extended control problem is related to the structure of risk-sensitive investment management applications. We introduce exploration through a randomization of the control. Next, we apply the duality between free energy and relative entropy to reduce the LEQG problem to an equivalent risk-neutral LQG control problem with an entropy regularization term, see, e.g. Dai Pra et al. (1996), for which we present a solution approach based on Dynamic Programming. Our approach, based on the energy-entropy duality may also be considered as leading to a justification for the use, in the literature, of an entropy regularization when applying a randomized control.

math.OC

Jump-Diffusion Risk-Sensitive Asset Management II: Jump-Diffusion Factor Model

In this article we extend earlier work on the jump-diffusion risk-sensitive asset management problem [SIAM J. Fin. Math. (2011) 22-54] by allowing jumps in both the factor process and the asset prices, as well as stochastic volatility and investment constraints. In this case, the HJB equation is a partial integro-differential equation (PIDE). By combining viscosity solutions with a change of notation, a policy improvement argument and classical results on parabolic PDEs we prove that the HJB PIDE admits a unique smooth solution. A verification theorem concludes the resolution of this problem.

q-fin.PM

Jump-Diffusion Risk-Sensitive Asset Management I: Diffusion Factor Model

This paper considers a portfolio optimization problem in which asset prices are represented by SDEs driven by Brownian motion and a Poisson random measure, with drifts that are functions of an auxiliary diffusion factor process. The criterion, following earlier work by Bielecki, Pliska, Nagai and others, is risk-sensitive optimization (equivalent to maximizing the expected growth rate subject to a constraint on variance.) By using a change of measure technique introduced by Kuroda and Nagai we show that the problem reduces to solving a certain stochastic control problem in the factor process, which has no jumps. The main result of the paper is to show that the risk-sensitive jump diffusion problem can be fully characterized in terms of a parabolic Hamilton-Jacobi-Bellman PDE rather than a PIDE, and that this PDE admits a classical C^{1,2} solution.

q-fin.PM

Jump-Diffusion Risk-Sensitive Asset Management

This paper considers a portfolio optimization problem in which asset prices are represented by SDEs driven by Brownian motion and a Poisson random measure, with drifts that are functions of an auxiliary diffusion 'factor' process. The criterion, following earlier work by Bielecki, Pliska, Nagai and others, is risk-sensitive optimization (equivalent to maximizing the expected growth rate subject to a constraint on variance.) By using a change of measure technique introduced by Kuroda and Nagai we show that the problem reduces to solving a certain stochastic control problem in the factor process, which has no jumps. The main result of the paper is that the Hamilton-Jacobi-Bellman equation for this problem has a classical solution. The proof uses Bellman's "policy improvement" method together with results on linear parabolic PDEs due to Ladyzhenskaya et al.

q-fin.PM

Risk Sensitive Investment Management with Affine Processes: a Viscosity Approach

In this paper, we extend the jump-diffusion model proposed by Davis and Lleo to include jumps in asset prices as well as valuation factors. The criterion, following earlier work by Bielecki, Pliska, Nagai and others, is risk-sensitive optimization (equivalent to maximizing the expected growth rate subject to a constraint on variance.) In this setting, the Hamilton- Jacobi-Bellman equation is a partial integro-differential PDE. The main result of the paper is to show that the value function of the control problem is the unique viscosity solution of the Hamilton-Jacobi-Bellman equation.

q-fin.PM