SearcharxivSearch

arXiv · 2008.12275

Market-making with reinforcement-learning (SAC)

Abstract

The paper explores the application of a continuous action space soft actor-critic (SAC) reinforcement learning model to the area of automated market-making. The reinforcement learning agent receives a simulated flow of client trades, thus accruing a position in an asset, and learns to offset this risk by either hedging at simulated "exchange" spreads or by attracting an offsetting client flow by changing offered client spreads (skewing the offered prices). The question of learning minimum spreads that compensate for the risk of taking the position is being investigated. Finally, the agent is posed with a problem of learning to hedge a blended client trade flow resulting from independent price processes (a "portfolio" position). The position penalty method is introduced to improve the convergence. An Open-AI gym-compatible hedge environment is introduced and the Open AI SAC baseline RL engine is being used as a learning baseline.

Explore related subjects

Keep this discovery

BibTeXRIS

Alexey Bakshaev. 2020-08-27. Market-making with reinforcement-learning (SAC). https://arxiv.org/abs/2008.12275

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Pricing and Hedging of Discretely Monitored Asian Options in the Volterra-Heston Model

We develop semi-closed pricing formulas and lifted-model hedging methods for discretely monitored geometric and arithmetic Asian options in the Volterra-Heston stochastic volatility model. Exploiting the affine Volterra structure, we derive a tractable transform for the joint law of the terminal log-price and the discretely monitored geometric average. This transform yields semi-closed pricing formulas for geometric Asian options, which in turn provide effective control variates for Monte Carlo valuation of arithmetic Asian options. Under the stated real-moment and affine-transform hypotheses, we also derive the Galtchouk-Kunita-Watanabe decomposition for Fourier-representable payoffs and obtain a variance-optimal hedge in terms of the Riccati-Volterra equation and the forward-variance curve. Using N-factor Markovian approximations, we obtain a finite-dimensional numerical implementation for hedging Asian options. Our numerical experiments document factor convergence for a regular non-Markovian kernel and the effect of rebalancing frequency on hedging error. In the Heston benchmark, geometric Asian controls substantially reduce the variance of arithmetic-Asian price estimates and improve the finite-sample stability of regression-based hedging relative to direct regression.

q-fin.PR

Beyond Lognormal Sums: A Four-Moment Probability Framework for Basket and Spread Option Pricing

Basket options are difficult to value under correlated lognormal dynamics because weighted sums and differences of lognormal variables have no tractable distribution. This paper develops a probability-based four-moment framework that separates the exact pricing representation from the distributional approximation. A change of measure first writes a basket price as a linear combination of probabilities. For a standard basket with one positive weight, these probabilities become CDF values of positive correlated lognormal sums. Each sum is approximated by a shifted lognormal variance mixture matched to its first four moments. For an unrestricted mixed-sign basket, a signed shifted lognormal proxy gives an analytical call-price formula. We state admissibility conditions, provide a practical root-selection rule, establish the main strike-based financial properties of the direct proxy, and derive exact pricing-error identities in terms of cumulative distribution function (CDF) discrepancies. The numerical analysis combines standard-basket benchmarks with an empirical application to a normalized $3{:}2{:}1$ crack spread constructed from RBOB gasoline, ULSD or heating oil, and WTI futures. The results show that the probability reformulation and the fourth-moment condition improve the distributional fit and pricing accuracy, particularly when maturity and tail asymmetry increase. The framework remains analytical, transparent, and suitable for repeated valuation across strikes and maturities.

q-fin.PR

When to Sell an Asset? - A Distribution Builder Approach

We consider the question of the optimal timing of the sale of an asset with stochastic dynamics. Our analysis is based on the method of the distribution builder introduced by Sharpe, Goldstein and Blythe [SGB00] for the purpose of optimal portfolio selection. Instead of specifying a utility function or risk aversion coefficient, this tool directly elicits the target distribution of the investor. We show how the problem of an optimal asset sale is in this setting linked to the problem of finding a Skorokhod embedding of a distribution into a diffusion process. In the case where the asset process follows a geometric Brownian motion and a specific family of distributions is targeted, one can observe a risk-return tradeoff.

q-fin.PR