SearcharxivSearch

arXiv · 2412.10692

Continuous-time optimal investment with portfolio constraints: a reinforcement learning approach

Abstract

In a reinforcement learning (RL) framework, we study the exploratory version of the continuous time expected utility (EU) maximization problem with a portfolio constraint that includes widely-used financial regulations such as short-selling constraints and borrowing prohibition. The optimal feedback policy of the exploratory unconstrained classical EU problem is shown to be Gaussian. In the case where the portfolio weight is constrained to a given interval, the corresponding constrained optimal exploratory policy follows a truncated Gaussian distribution. We verify that the closed form optimal solution obtained for logarithmic utility and quadratic utility for both unconstrained and constrained situations converge to the non-exploratory expected utility counterpart when the exploration weight goes to zero. Finally, we establish a policy improvement theorem and devise an implementable reinforcement learning algorithm by casting the optimal problem in a martingale framework. Our numerical examples show that exploration leads to an optimal wealth process that is more dispersedly distributed with heavier tail compared to that of the case without exploration. This effect becomes less significant as the exploration parameter is smaller. Moreover, the numerical implementation also confirms the intuitive understanding that a broader domain of investment opportunities necessitates a higher exploration cost. Notably, when subjected to both short-selling and money borrowing constraints, the exploration cost becomes negligible compared to the unconstrained case.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Huy Chau, Duy Nguyen, Thai Nguyen. 2024-12-14. Continuous-time optimal investment with portfolio constraints: a reinforcement learning approach. https://arxiv.org/abs/2412.10692

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Variance-Optimal Hedging in the Rough Hawkes--Heston Model

We study variance-optimal stock hedging and the convergence of approximate strategies in the rough Hawkes--Heston model. Starting from the model's affine conditional transform and the affine Volterra jump framework, we obtain semi-explicit hedges for European calls and a representation of the minimum quadratic error through the Galtchouk--Kunita--Watanabe projection. Our main approximation result keeps the original stock, variance driver, and information flow fixed while regularizing the kernel used to evaluate the hedge. To handle singular memory and common marked jumps, we construct the approximate holdings from histories available before trading and preserve the conditional transform's random modulus envelope. Riccati--Volterra stability and weighted truncation then yield convergence in the original stock's trading norm on compact Fourier intervals. For calls, a joint choice of kernel regularization and Fourier cutoff gives convergence of the initial capitals and strategies, uniform-in-time square-mean convergence of continuous-time gains, and convergence of the terminal mean-square error to the variance-optimal value. A numerical experiment with shifted fractional kernels illustrates the construction on common original-market paths.

q-fin.MF

Numeraire Invariance of Entropy-Projected Martingale Measures

Let \(P\) be a fixed physical law and let \(Q\) be an equivalent martingale measure selected from the martingale-measure set associated with a chosen numeraire. A change of numeraire maps \(Q\) to \(T_LQ\), where \(d(T_LQ)=L\,dQ\) and \(L\) is the terminal likelihood ratio. The forward relative-entropy projection minimizing \(D_{\mathrm{KL}}(P\Vert Q)\) commutes with this transform because its objective changes only by the constant \(-E_P\log L\). The minimal entropy martingale measure (MEMM) orientation \(D_{\mathrm{KL}}(Q\Vert P)\) does not have this property, and a trinomial counterexample shows that independently recomputed MEMMs need not be likelihood compatible. We make two economic consequences explicit. First, the two entropy orientations are precisely the \(Q\)-dependent terms in the classical convex-dual objectives for logarithmic and exponential utility, respectively. Second, likelihood compatibility is equivalent to equality of the pricing functionals obtained in the two numeraires. Hence the forward selectors value every integrable claim consistently across numeraires, whereas the two MEMMs in the counterexample assign different prices to a nonreplicable digital claim. We also prove a finite-state class-level characterization: uniform invariance over the elementary one-period likelihood-ratio families forces a smooth convex \(f\)-divergence to be logarithmic, up to scaling and affine equivalence. Finally, in finite-state markets, the forward projection exists under the usual strictly positive feasible-point condition; its density \(dP/dQ^*\) is attainable log-optimal terminal wealth, and the minimum forward entropy equals maximal expected log growth.

q-fin.MF

The Delta of a Variance Swap

We define the variance swap delta as the sensitivity of the price of variance to a change in underlying price. We use Carr-Madan spanning formulas to analyze this sensitivity when the implied volatility smile curve may depend on the underlying price. We show that the variance swap total delta is zero for the class of smile curves that are pure functions of (log) moneyness, which goes against the empirical observation that variance is up when the market is down. We propose a simple modification of the smile to correct this issue.

q-fin.MF