Searcharxiv⌕ Search

arXiv subjects

Igor Halperin

Publications and source records attributed to Igor Halperin.

At least 37 records · Page 2Linked to original sources

Phases of MANES: Multi-Asset Non-Equilibrium Skew Model of a Strongly Non-Linear Market with Phase Transitions

This paper presents an analytically tractable and practically-oriented model of non-linear dynamics of a multi-asset market in the limit of a large number of assets. The asset price dynamics are driven by money flows into the market from external investors, and their price impact. This leads to a model of a market as an ensemble of interacting non-linear oscillators with the Langevin dynamics. In a homogeneous portfolio approximation, the mean field treatment of the resulting Langevin dynamics produces the McKean-Vlasov equation as a dynamic equation for market returns. Due to the strong non-linearity of the McKean-Vlasov equation, the resulting dynamics give rise to ergodicity breaking and first- or second-order phase transitions under variations of model parameters. Using a tractable potential of the Non-Equilibrium Skew (NES) model previously suggested by the author for a single-stock case, the new Multi-Asset NES (MANES) model enables an analytically tractable framework for a multi-asset market. The equilibrium expected market log-return is obtained as a self-consistent mean field of the McKean-Vlasov equation, and derived in closed form in terms of parameters that are inferred from market prices of S&P 500 index options. The model is able to accurately fit the market data for either a benign or distressed market environments, while using only a single volatility parameter.

q-fin.CP↗

Combining Reinforcement Learning and Inverse Reinforcement Learning for Asset Allocation Recommendations

We suggest a simple practical method to combine the human and artificial intelligence to both learn best investment practices of fund managers, and provide recommendations to improve them. Our approach is based on a combination of Inverse Reinforcement Learning (IRL) and RL. First, the IRL component learns the intent of fund managers as suggested by their trading history, and recovers their implied reward function. At the second step, this reward function is used by a direct RL algorithm to optimize asset allocation decisions. We show that our method is able to improve over the performance of individual fund managers.

cs.LG↗

Non-Equilibrium Skewness, Market Crises, and Option Pricing: Non-Linear Langevin Model of Markets with Supersymmetry

This paper presents a tractable model of non-linear dynamics of market returns using a Langevin approach. Due to non-linearity of an interaction potential, the model admits regimes of both small and large return fluctuations. Langevin dynamics are mapped onto an equivalent quantum mechanical (QM) system. Borrowing ideas from supersymmetric quantum mechanics (SUSY QM), a parameterized ground state wave function (WF) of this QM system is used as a direct input to the model, which also fixes a non-linear Langevin potential. Using a two-component Gaussian mixture as a ground state WF with an asymmetric double well potential produces a tractable low-parametric model with interpretable parameters, referred to as the NES (Non-Equilibrium Skew) model. Supersymmetry (SUSY) is then used to find time-dependent solutions of the model in an analytically tractable way. Additional approximations give rise to a final practical version of the NES model, where real-measure and risk-neutral return distributions are given by three component Gaussian mixtures. This produces a closed-form approximation for option pricing in the NES model by a mixture of three Black-Scholes prices, providing accurate calibration to option prices for either benign or distressed market environments, while using only a single volatility parameter. These results stand in stark contrast to the most of other option pricing models such as local, stochastic, or rough volatility models that need more complex specifications of noise to fit the market data.

q-fin.CP↗

Distributional Offline Continuous-Time Reinforcement Learning with Neural Physics-Informed PDEs (SciPhy RL for DOCTR-L)

This paper addresses distributional offline continuous-time reinforcement learning (DOCTR-L) with stochastic policies for high-dimensional optimal control. A soft distributional version of the classical Hamilton-Jacobi-Bellman (HJB) equation is given by a semilinear partial differential equation (PDE). This `soft HJB equation' can be learned from offline data without assuming that the latter correspond to a previous optimal or near-optimal policy. A data-driven solution of the soft HJB equation uses methods of Neural PDEs and Physics-Informed Neural Networks developed in the field of Scientific Machine Learning (SciML). The suggested approach, dubbed `SciPhy RL', thus reduces DOCTR-L to solving neural PDEs from data. Our algorithm called Deep DOCTR-L converts offline high-dimensional data into an optimal policy in one step by reducing it to supervised learning, instead of relying on value iteration or policy iteration methods. The method enables a computable approach to the quality control of obtained policies in terms of both their expected returns and uncertainties about their values.

cs.LG↗

G-Learner and GIRL: Goal Based Wealth Management with Reinforcement Learning

We present a reinforcement learning approach to goal based wealth management problems such as optimization of retirement plans or target dated funds. In such problems, an investor seeks to achieve a financial goal by making periodic investments in the portfolio while being employed, and periodically draws from the account when in retirement, in addition to the ability to re-balance the portfolio by selling and buying different assets (e.g. stocks). Instead of relying on a utility of consumption, we present G-Learner: a reinforcement learning algorithm that operates with explicitly defined one-step rewards, does not assume a data generation process, and is suitable for noisy data. Our approach is based on G-learning - a probabilistic extension of the Q-learning method of reinforcement learning. In this paper, we demonstrate how G-learning, when applied to a quadratic reward and Gaussian reference policy, gives an entropy-regulated Linear Quadratic Regulator (LQR). This critical insight provides a novel and computationally tractable tool for wealth management tasks which scales to high dimensional portfolios. In addition to the solution of the direct problem of G-learning, we also present a new algorithm, GIRL, that extends our goal-based G-learning approach to the setting of Inverse Reinforcement Learning (IRL) where rewards collected by the agent are not observed, and should instead be inferred. We demonstrate that GIRL can successfully learn the reward parameters of a G-Learner agent and thus imitate its behavior. Finally, we discuss potential applications of the G-Learner and GIRL algorithms for wealth management and robo-advising.

q-fin.PM↗

QLBS: Q-Learner in the Black-Scholes(-Merton) Worlds

This paper presents a discrete-time option pricing model that is rooted in Reinforcement Learning (RL), and more specifically in the famous Q-Learning method of RL. We construct a risk-adjusted Markov Decision Process for a discrete-time version of the classical Black-Scholes-Merton (BSM) model, where the option price is an optimal Q-function, while the optimal hedge is a second argument of this optimal Q-function, so that both the price and hedge are parts of the same formula. Pricing is done by learning to dynamically optimize risk-adjusted returns for an option replicating portfolio, as in the Markowitz portfolio theory. Using Q-Learning and related methods, once created in a parametric setting, the model is able to go model-free and learn to price and hedge an option directly from data, and without an explicit model of the world. This suggests that RL may provide efficient data-driven and model-free methods for optimal pricing and hedging of options, once we depart from the academic continuous-time limit, and vice versa, option pricing methods developed in Mathematical Finance may be viewed as special cases of model-based Reinforcement Learning. Further, due to simplicity and tractability of our model which only needs basic linear algebra (plus Monte Carlo simulation, if we work with synthetic data), and its close relation to the original BSM model, we suggest that our model could be used for benchmarking of different RL algorithms for financial trading applications

q-fin.CP↗

"Quantum Equilibrium-Disequilibrium": Asset Price Dynamics, Symmetry Breaking, and Defaults as Dissipative Instantons

We propose a simple non-equilibrium model of a financial market as an open system with a possible exchange of money with an outside world and market frictions (trade impacts) incorporated into asset price dynamics via a feedback mechanism. Using a linear market impact model, this produces a non-linear two-parametric extension of the classical Geometric Brownian Motion (GBM) model, that we call the "Quantum Equilibrium-Disequilibrium" (QED) model. The QED model gives rise to non-linear mean-reverting dynamics, broken scale invariance, and corporate defaults. In the simplest one-stock (1D) formulation, our parsimonious model has only one degree of freedom, yet calibrates to both equity returns and credit default swap spreads. Defaults and market crashes are associated with dissipative tunneling events, and correspond to instanton (saddle-point) solutions of the model. When market frictions and inflows/outflows of money are neglected altogether, "classical" GBM scale-invariant dynamics with an exponential asset growth and without defaults are formally recovered from the QED dynamics. However, we argue that this is only a formal mathematical limit, and in reality the GBM limit is non-analytic due to non-linear effects that produce both defaults and divergence of perturbation theory in a small market friction parameter.

q-fin.ST↗

Market Self-Learning of Signals, Impact and Optimal Trading: Invisible Hand Inference with Free Energy

We present a simple model of a non-equilibrium self-organizing market where asset prices are partially driven by investment decisions of a bounded-rational agent. The agent acts in a stochastic market environment driven by various exogenous "alpha" signals, agent's own actions (via market impact), and noise. Unlike traditional agent-based models, our agent aggregates all traders in the market, rather than being a representative agent. Therefore, it can be identified with a bounded-rational component of the market itself, providing a particular implementation of an Invisible Hand market mechanism. In such setting, market dynamics are modeled as a fictitious self-play of such bounded-rational market-agent in its adversarial stochastic environment. As rewards obtained by such self-playing market agent are not observed from market data, we formulate and solve a simple model of such market dynamics based on a neuroscience-inspired Bounded Rational Information Theoretic Inverse Reinforcement Learning (BRIT-IRL). This results in effective asset price dynamics with a non-linear mean reversion - which in our model is generated dynamically, rather than being postulated. We argue that our model can be used in a similar way to the Black-Litterman model. In particular, it represents, in a simple modeling framework, market views of common predictive signals, market impacts and implied optimal dynamic portfolio allocations, and can be used to assess values of private signals. Moreover, it allows one to quantify a "market-implied" optimal investment strategy, along with a measure of market rationality. Our approach is numerically light, and can be implemented using standard off-the-shelf software such as TensorFlow.

q-fin.CP↗

The QLBS Q-Learner Goes NuQLear: Fitted Q Iteration, Inverse RL, and Option Portfolios

The QLBS model is a discrete-time option hedging and pricing model that is based on Dynamic Programming (DP) and Reinforcement Learning (RL). It combines the famous Q-Learning method for RL with the Black-Scholes (-Merton) model's idea of reducing the problem of option pricing and hedging to the problem of optimal rebalancing of a dynamic replicating portfolio for the option, which is made of a stock and cash. Here we expand on several NuQLear (Numerical Q-Learning) topics with the QLBS model. First, we investigate the performance of Fitted Q Iteration for a RL (data-driven) solution to the model, and benchmark it versus a DP (model-based) solution, as well as versus the BSM model. Second, we develop an Inverse Reinforcement Learning (IRL) setting for the model, where we only observe prices and actions (re-hedges) taken by a trader, but not rewards. Third, we outline how the QLBS model can be used for pricing portfolios of options, rather than a single option in isolation, thus providing its own, data-driven and model independent solution to the (in)famous volatility smile problem of the Black-Scholes model.

q-fin.CP↗

Inverse Reinforcement Learning for Marketing

Learning customer preferences from an observed behaviour is an important topic in the marketing literature. Structural models typically model forward-looking customers or firms as utility-maximizing agents whose utility is estimated using methods of Stochastic Optimal Control. We suggest an alternative approach to study dynamic consumer demand, based on Inverse Reinforcement Learning (IRL). We develop a version of the Maximum Entropy IRL that leads to a highly tractable model formulation that amounts to low-dimensional convex optimization in the search for optimal model parameters. Using simulations of consumer demand, we show that observational noise for identical customers can be easily confused with an apparent consumer heterogeneity.

q-fin.CP↗

Keep It Real: Tail Probabilities of Compound Heavy-Tailed Distributions

We propose an analytical approach to the computation of tail probabilities of compound distributions whose individual components have heavy tails. Our approach is based on the contour integration method, and gives rise to a representation of the tail probability of a compound distribution in the form of a rapidly convergent one-dimensional integral involving a discontinuity of the imaginary part of its moment generating function across a branch cut. The latter integral can be evaluated in quadratures, or alternatively represented as an asymptotic expansion. Our approach thus offers a viable (especially at high percentile levels) alternative to more standard methods such as Monte Carlo or the Fast Fourier Transform, traditionally used for such problems. As a practical application, we use our method to compute the operational Value at Risk (VAR) of a financial institution, where individual losses are modeled as spliced distributions whose large loss components are given by power-law or lognormal distributions. Finally, we briefly discuss extensions of the present formalism for calculation of tail probabilities of compound distributions made of compound distributions with heavy tails.

q-fin.CP↗

USLV: Unspanned Stochastic Local Volatility Model

We propose a new framework for modeling stochastic local volatility, with potential applications to modeling derivatives on interest rates, commodities, credit, equity, FX etc., as well as hybrid derivatives. Our model extends the linearity-generating unspanned volatility term structure model by Carr et al. (2011) by adding a local volatility layer to it. We outline efficient numerical schemes for pricing derivatives in this framework for a particular four-factor specification (two "curve" factors plus two "volatility" factors). We show that the dynamics of such a system can be approximated by a Markov chain on a two-dimensional space (Z_t,Y_t), where coordinates Z_t and Y_t are given by direct (Kroneker) products of values of pairs of curve and volatility factors, respectively. The resulting Markov chain dynamics on such partly "folded" state space enables fast pricing by the standard backward induction. Using a nonparametric specification of the Markov chain generator, one can accurately match arbitrary sets of vanilla option quotes with different strikes and maturities. Furthermore, we consider an alternative formulation of the model in terms of an implied time change process. The latter is specified nonparametrically, again enabling accurate calibration to arbitrary sets of vanilla option quotes.

q-fin.PR↗

Pricing options on illiquid assets with liquid proxies using utility indifference and dynamic-static hedging

This work addresses the problem of optimal pricing and hedging of a European option on an illiquid asset Z using two proxies: a liquid asset S and a liquid European option on another liquid asset Y. We assume that the S-hedge is dynamic while the Y-hedge is static. Using the indifference pricing approach we derive a HJB equation for the value function, and solve it analytically (in quadratures) using an asymptotic expansion around the limit of the perfect correlation between assets Y and Z. While in this paper we apply our framework to an incomplete market version of the credit-equity Merton's model, the same approach can be used for other asset classes (equity, commodity, FX, etc.), e.g. for pricing and hedging options with illiquid strikes or illiquid exotic options.

q-fin.PR↗

Implied Multi-Factor Model for Bespoke CDO Tranches and other Portfolio Credit Derivatives

This paper introduces a new semi-parametric approach to the pricing and risk management of bespoke CDO tranches, with a particular attention to bespokes that need to be mapped onto more than one reference portfolio. The only user input in our framework is a multi-factor model (a "prior" model hereafter) for index portfolios, such as CDX.NA.IG or iTraxx Europe, that are chosen as benchmark securities for the pricing of a given bespoke CDO. Parameters of the prior model are fixed, and not tuned to match prices of benchmark index tranches. Instead, our calibration procedure amounts to a proper reweightening of the prior measure using the Minimum Cross Entropy method. As the latter problem reduces to convex optimization in a low dimensional space, our model is computationally efficient. Both the static (one-period) and dynamic versions of the model are presented. The latter can be used for pricing and risk management of more exotic instruments referencing bespoke portfolios, such as forward-starting tranches or tranche options, and for calculation of credit valuation adjustment (CVA) for bespoke tranches.

q-fin.PR↗

BSLP: Markovian Bivariate Spread-Loss Model for Portfolio Credit Derivatives

BSLP is a two-dimensional dynamic model of interacting portfolio-level loss and spread (more exactly, loss intensity) processes. The model is similar to the top-down HJM-like frameworks developed by Schonbucher (2005) and Sidenius-Peterbarg-Andersen (SPA) (2005), however is constructed as a Markovian, short-rate intensity model. This property of the model enables fast lattice methods for pricing various portfolio credit derivatives such as tranche options, forward-starting tranches, leveraged super-senior tranches etc. A non-parametric model specification is used to achieve nearly perfect calibration to liquid tranche quotes across strikes and maturities. A non-dynamic version of the model obtained in the zero volatility limit of stochastic intensity is useful on its own as an arbitrage-free interpolation model to price non-standard index tranches off the standard ones.

q-fin.PR↗

Climbing Down from the Top: Single Name Dynamics in Credit Top Down Models

In the top-down approach to multi-name credit modeling, calculation of singe name sensitivities appears possible, at least in principle, within the so-called random thinning (RT) procedure which dissects the portfolio risk into individual contributions. We make an attempt to construct a practical RT framework that enables efficient calculation of single name sensitivities in a top-down framework, and can be extended to valuation and risk management of bespoke tranches. Furthermore, we propose a dynamic extension of the RT method that enables modeling of both idiosyncratic and default-contingent individual spread dynamics within a Monte Carlo setting in a way that preserves the portfolio "top"-level dynamics. This results in a model that is not only calibrated to tranche and single name spreads, but can also be tuned to approximately match given levels of spread volatilities and correlations of names in the portfolio.

q-fin.PR↗

Bayesian Entropic Inverse Theory Approach to Implied Option Pricing with Noisy Data

A popular approach to nonparametric option pricing is the Minimum Cross Entropy (MCE) method based on minimization of the relative Kullback-Leibler entropy of the price density distribution and a given reference density, with observable option prices serving as constraints. When market prices are noisy, the MCE method tends to overfit the data and often becomes unstable. We propose a non-parametric option pricing method whose input are noisy market prices of arbitrary number of European options with arbitrary maturities. Implied transition densities are calculated using the Bayesian inverse theory with entropic priors, with a reference density which may be estimated by the algorithm itself. In the limit of zero noise, our approach is shown to reduce to the canonical MCE method generalized to a multi-period case. The method can be used for a non-parametric pricing of American/Bermudan options with a possible weak path dependence.

cond-mat.stat-mech↗