SearcharxivSearch

arXiv · 2109.04001

Deep Reinforcement Learning for Equal Risk Pricing and Hedging under Dynamic Expectile Risk Measures

Abstract

Recently equal risk pricing, a framework for fair derivative pricing, was extended to consider dynamic risk measures. However, all current implementations either employ a static risk measure that violates time consistency, or are based on traditional dynamic programming solution schemes that are impracticable in problems with a large number of underlying assets (due to the curse of dimensionality) or with incomplete asset dynamics information. In this paper, we extend for the first time a famous off-policy deterministic actor-critic deep reinforcement learning (ACRL) algorithm to the problem of solving a risk averse Markov decision process that models risk using a time consistent recursive expectile risk measure. This new ACRL algorithm allows us to identify high quality time consistent hedging policies (and equal risk prices) for options, such as basket options, that cannot be handled using traditional methods, or in context where only historical trajectories of the underlying assets are available. Our numerical experiments, which involve both a simple vanilla option and a more exotic basket option, confirm that the new ACRL algorithm can produce 1) in simple environments, nearly optimal hedging policies, and highly accurate prices, simultaneously for a range of maturities 2) in complex environments, good quality policies and prices using reasonable amount of computing resources; and 3) overall, hedging strategies that actually outperform the strategies produced using static risk measures when the risk is evaluated at later points of time.

Explore related subjects

Keep this discovery

BibTeXRIS

Saeed Marzban, Erick Delage, Jonathan Yumeng Li. 2021-09-09. Deep Reinforcement Learning for Equal Risk Pricing and Hedging under Dynamic Expectile Risk Measures. https://arxiv.org/abs/2109.04001

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Pricing and Hedging of Discretely Monitored Asian Options in the Volterra-Heston Model

We develop semi-closed pricing formulas and lifted-model hedging methods for discretely monitored geometric and arithmetic Asian options in the Volterra-Heston stochastic volatility model. Exploiting the affine Volterra structure, we derive a tractable transform for the joint law of the terminal log-price and the discretely monitored geometric average. This transform yields semi-closed pricing formulas for geometric Asian options, which in turn provide effective control variates for Monte Carlo valuation of arithmetic Asian options. Under the stated real-moment and affine-transform hypotheses, we also derive the Galtchouk-Kunita-Watanabe decomposition for Fourier-representable payoffs and obtain a variance-optimal hedge in terms of the Riccati-Volterra equation and the forward-variance curve. Using N-factor Markovian approximations, we obtain a finite-dimensional numerical implementation for hedging Asian options. Our numerical experiments document factor convergence for a regular non-Markovian kernel and the effect of rebalancing frequency on hedging error. In the Heston benchmark, geometric Asian controls substantially reduce the variance of arithmetic-Asian price estimates and improve the finite-sample stability of regression-based hedging relative to direct regression.

q-fin.PR

Beyond Lognormal Sums: A Four-Moment Probability Framework for Basket and Spread Option Pricing

Basket options are difficult to value under correlated lognormal dynamics because weighted sums and differences of lognormal variables have no tractable distribution. This paper develops a probability-based four-moment framework that separates the exact pricing representation from the distributional approximation. A change of measure first writes a basket price as a linear combination of probabilities. For a standard basket with one positive weight, these probabilities become CDF values of positive correlated lognormal sums. Each sum is approximated by a shifted lognormal variance mixture matched to its first four moments. For an unrestricted mixed-sign basket, a signed shifted lognormal proxy gives an analytical call-price formula. We state admissibility conditions, provide a practical root-selection rule, establish the main strike-based financial properties of the direct proxy, and derive exact pricing-error identities in terms of cumulative distribution function (CDF) discrepancies. The numerical analysis combines standard-basket benchmarks with an empirical application to a normalized $3{:}2{:}1$ crack spread constructed from RBOB gasoline, ULSD or heating oil, and WTI futures. The results show that the probability reformulation and the fourth-moment condition improve the distributional fit and pricing accuracy, particularly when maturity and tail asymmetry increase. The framework remains analytical, transparent, and suitable for repeated valuation across strikes and maturities.

q-fin.PR

When to Sell an Asset? - A Distribution Builder Approach

We consider the question of the optimal timing of the sale of an asset with stochastic dynamics. Our analysis is based on the method of the distribution builder introduced by Sharpe, Goldstein and Blythe [SGB00] for the purpose of optimal portfolio selection. Instead of specifying a utility function or risk aversion coefficient, this tool directly elicits the target distribution of the investor. We show how the problem of an optimal asset sale is in this setting linked to the problem of finding a Skorokhod embedding of a distribution into a diffusion process. In the case where the asset process follows a geometric Brownian motion and a specific family of distributions is targeted, one can observe a risk-return tradeoff.

q-fin.PR