SearcharxivSearch

arXiv subjects

Sebastian Jaimungal

Publications and source records attributed to Sebastian Jaimungal.

At least 19 recordsLinked to original sources

Equilibrium Liquidity and Risk Offsetting in Decentralised Markets

We study the economic viability of liquidity provision in decentralised exchanges (DEXs) within a structural framework in which market outcomes are endogenous. We formulate strategic interactions as a sequential game: a risk-averse liquidity provider (LP) sets the supply of liquidity in the DEX and a costly dynamic replication strategy in a centralised exchange (CEX), price-sensitive traders determine trading volumes, and arbitrageurs align prices. We establish existence of equilibrium under general trading functions. We show that DEX liquidity depth is a central instrument for risk management, because the LP adjusts liquidity ex ante to manage exposure. In addition to the classical trade-off between liquidity demand and adverse selection, we identify two further determinants of the viability of liquidity provision: the ratio of risk aversion to replication costs and private information. The ratio governs the aggressiveness of replication: greater relative risk aversion reduces risk but also lowers equilibrium liquidity and its mean profitability. Private information has a non-monotonic effect. For moderate price movements, speculative benefits increase liquidity. For large price movements, anticipated adverse selection and replication costs lead to thinner markets.

q-fin.TR

Deep Learning and Elicitability for McKean-Vlasov FBSDEs With Common Noise

We present a novel numerical method for solving McKean--Vlasov forward--backward stochastic differential equations (MV--FBSDEs) with common noise, combining Picard iterations, elicitability and deep learning. The key innovation involves elicitability to derive a pathwise loss function, enabling efficient training of neural networks to approximate both the backward process and the conditional expectations arising from common noise, without requiring computationally expensive nested Monte Carlo simulations. The mean-field interaction term is parameterized via a recurrent neural network trained to minimize an elicitable score, while the backward process is approximated through a hybrid feedforward and recurrent network representing the decoupling field. We validate the algorithm on a systemic-risk inter-bank borrowing and lending model, where analytical solutions exist, demonstrating accurate recovery of the true solution. We further extend the model to quantile-mediated interactions, showcasing the flexibility of the elicitability framework beyond conditional means or moments. Finally, we apply the method to a non-stationary Aiyagari--Bewley--Huggett economic growth model with endogenous interest rates, illustrating its applicability to complex mean-field games without closed-form solutions.

cs.LG

Deep reinforcement learning for optimal trading with partial information

Reinforcement Learning (RL) applied to financial problems has been the subject of a lively area of research. The use of RL for optimal trading strategies that exploit latent information in the market is, to the best of our knowledge, not widely tackled. In this paper we study an optimal trading problem, where a trading signal follows an Ornstein-Uhlenbeck process with regime-switching dynamics. We employ a blend of RL and Recurrent Neural Networks (RNN) in order to make the most at extracting underlying information from the trading signal with latent parameters. The latent parameters driving mean reversion, speed, and volatility are filtered from observations of the signal, and trading strategies are derived via RL. To address this problem, we propose three Deep Deterministic Policy Gradient (DDPG)-based algorithms that integrate Gated Recurrent Unit (GRU) networks to capture temporal dependencies in the signal. The first, a one -step approach (hid-DDPG), directly encodes hidden states from the GRU into the RL trader. The second and third are two-step methods: one (prob-DDPG) makes use of posterior regime probability estimates, while the other (reg-DDPG) relies on forecasts of the next signal value. Through extensive simulations with increasingly complex Markovian regime dynamics for the trading signal's parameters, as well as an empirical application to equity pair trading, we find that prob-DDPG achieves superior cumulative rewards and exhibits more interpretable strategies. By contrast, reg-DDPG provides limited benefits, while hid-DDPG offers intermediate performance with less interpretable strategies. Our results show that the quality and structure of the information supplied to the agent are crucial: embedding probabilistic insights into latent regimes substantially improves both profitability and robustness of reinforcement learning-based trading strategies.

q-fin.TR

Deviations from Tradition: Stylized Facts in the Era of DeFi

Decentralized Exchanges (DEXs) are now a significant component of the financial world where billions of dollars are traded daily. Differently from traditional markets, which are typically based on Limit Order Books, DEXs typically work as Automated Market Makers, and, since the implementation of Uniswap v3, feature concentrated liquidity. By investigating the twenty-four most active pools in Uniswap v3 during 2023 and 2024, we empirically study how this structural change in the organization of the markets modifies the well-studied stylized facts of prices, liquidity, and order flow observed in traditional markets. We find a series of new statistical regularities in the distributions and cross-autocorrelation functions of these variables that we are able to associate either with the market structure (e.g., the execution of orders in blocks) or with the intense activity of Maximal Extractable Value searchers, such as Just-in-Time liquidity providers and sandwich attackers.

q-fin.TR

Multi-Agent Reinforcement Learning for Greenhouse Gas Offset Credit Markets

Climate change is a major threat to the future of humanity, and its impacts are being intensified by excess man-made greenhouse gas emissions. One method governments can employ to control these emissions is to provide firms with emission limits and penalize any excess emissions above the limit. Excess emissions may also be offset by firms who choose to invest in carbon reducing and capturing projects. These projects generate offset credits which can be submitted to a regulating agency to offset a firm's excess emissions, or they can be traded with other firms. In this work, we characterize the finite-agent Nash equilibrium for offset credit markets. As computing Nash equilibria is an NP-hard problem, we utilize the modern reinforcement learning technique Nash-DQN to efficiently estimate the market's Nash equilibria. We demonstrate not only the validity of employing reinforcement learning methods applied to climate themed financial markets, but also the significant financial savings emitting firms may achieve when abiding by the Nash equilibria through numerical experiments.

q-fin.MF

Model Combination in Risk Sharing under Ambiguity

We consider the problem of an agent who faces losses in continuous time over a finite time horizon and may choose to share some of these losses with a counterparty. The agent is uncertain about the true loss distribution and has multiple models for the losses, characterized by a finite set of probability measures. Their goal is to optimize a mean-variance type criterion with model combination under ambiguity through risk sharing. We construct such a criterion using the chi-squared divergence, exploiting a dual representation to expand the state space, yielding a time consistent problem. Assuming a Cram\'er-Lundberg loss model, we fully characterize the optimal risk sharing contract and the agent's wealth process under the optimal strategy. We prove that the strategy we obtain is admissible and that the value function satisfies the appropriate verification conditions. Furthermore, we show that the model combination problem is equivalent to the monotone mean-variance problem of Maccheroni et al. (2009) under a composite probability measure that depends on the agent's reference probability measures. Finally, we apply the optimal strategy to an insurance setting in a simulation example and provide numerical illustrations of the results.

q-fin.RM

Broker-Trader Partial Information Nash-Equilibria

We study partial information Nash equilibrium between a broker and an informed trader. In this setting, the informed trader, who possesses knowledge of a trading signal, trades multiple assets with the broker in a dealer market. Simultaneously, the broker offloads these assets in a lit exchange where their actions impact the asset prices. The broker, however, only observes aggregate prices and cannot distinguish between underlying trends and volatility. Both the broker and the informed trader aim to maximize their penalized expected wealth. Using convex analysis, we characterize the Nash equilibrium and demonstrate its existence and uniqueness. Furthermore, we establish that this equilibrium corresponds to the solution of a nonstandard system of forward-backward stochastic differential equations (FBSDEs) that involves the two differing filtrations. For short enough time horizons, we prove that a unique solution of this system exists. Finally, under quite general assumptions, we show that the solution to the FBSDE system admits a polynomial approximation in the strength of the transient impact to arbitrary order, and prove that the error is controlled.

q-fin.MF

Robust Reinforcement Learning with Dynamic Distortion Risk Measures

In a reinforcement learning (RL) setting, the agent's optimal strategy heavily depends on her risk preferences and the underlying model dynamics of the training environment. These two aspects influence the agent's ability to make well-informed and time-consistent decisions when facing testing environments. In this work, we devise a framework to solve robust risk-aware RL problems where we simultaneously account for environmental uncertainty and risk with a class of dynamic robust distortion risk measures. Robustness is introduced by considering all models within a Wasserstein ball around a reference model. We estimate such dynamic robust risk measures using neural networks by making use of strictly consistent scoring functions, derive policy gradient formulae using the quantile representation of distortion risk measures, and construct an actor-critic algorithm to solve this class of robust risk-aware RL problems. We demonstrate the performance of our algorithm on a portfolio allocation example.

cs.LG

Nash Equilibrium between Brokers and Traders

We study the perfect information Nash equilibrium between a broker and her clients -- an informed trader and an uniformed trader. In our model, the broker trades in the lit exchange where trades have instantaneous and transient price impact with exponential resilience, while both clients trade with the broker. The informed trader and the broker maximise expected wealth subject to inventory penalties, while the uninformed trader is not strategic and sends the broker random buy and sell orders. We characterise the Nash equilibrium of the trading strategies with the solution to a coupled system of forward-backward stochastic differential equations (FBSDEs). We solve this system explicitly and study the effect of information, profitability, and inventory control in the trading strategies of the broker and the informed trader.

q-fin.TR

Kullback-Leibler Barycentre of Stochastic Processes

We consider the problem where an agent aims to combine the views and insights of different experts' models. Specifically, each expert proposes a diffusion process over a finite time horizon. The agent then combines the experts' models by minimising the weighted Kullback--Leibler divergence to each of the experts' models. We show existence and uniqueness of the barycentre model and prove an explicit representation of the Radon--Nikodym derivative relative to the average drift model. We further allow the agent to include their own constraints, resulting in an optimal model that can be seen as a distortion of the experts' barycentre model to incorporate the agent's constraints. We propose two deep learning algorithms to approximate the optimal drift of the combined model, allowing for efficient simulations. The first algorithm aims at learning the optimal drift by matching the change of measure, whereas the second algorithm leverages the notion of elicitability to directly estimate the value function. The paper concludes with an extended application to combine implied volatility smile models that were estimated on different datasets.

q-fin.MF

Nash Equilibria in Greenhouse Gas Offset Credit Markets

One approach to reducing greenhouse gas (GHG) emissions is to incentivize carbon capturing and carbon reducing projects while simultaneously penalising excess GHG output. In this work, we present a novel market framework and characterise the optimal behaviour of GHG offset credit (OC) market participants in both single-player and two-player settings. The single player setting is posed as an optimal stopping and control problem, while the two-player setting is posed as optimal stopping and mixed-Nash equilibria problem. We demonstrate the importance of acting optimally using numerical solutions and Monte Carlo simulations and explore the differences between the homogeneous and heterogeneous players. In both settings, we find that market participants benefit from optimal OC trading and OC generation.

q-fin.GN

Learning conditional distributions on continuous spaces

We investigate sample-based learning of conditional distributions on multi-dimensional unit boxes, allowing for different dimensions of the feature and target spaces. Our approach involves clustering data near varying query points in the feature space to create empirical measures in the target space. We employ two distinct clustering schemes: one based on a fixed-radius ball and the other on nearest neighbors. We establish upper bounds for the convergence rates of both methods and, from these bounds, deduce optimal configurations for the radius and the number of neighbors. We propose to incorporate the nearest neighbors method into neural network training, as our empirical analysis indicates it has better performance in practice. For efficiency, our training process utilizes approximate nearest neighbors search with random binary space partitioning. Additionally, we employ the Sinkhorn algorithm and a sparsity-enforced transport plan. Our empirical findings demonstrate that, with a suitably designed structure, the neural network has the ability to adapt to a suitable level of Lipschitz continuity locally. For reproducibility, our code is available at \url{https://github.com/zcheng-a/LCD_kNN}.

stat.ML

Risk-Averse Markov Decision Processes through a Distributional Lens

By adopting a distributional viewpoint on law-invariant convex risk measures, we construct dynamics risk measures (DRMs) at the distributional level. We then apply these DRMs to investigate Markov decision processes, incorporating latent costs, random actions, and weakly continuous transition kernels. Furthermore, the proposed DRMs allow risk aversion to change dynamically. Under mild assumptions, we derive a dynamic programming principle and show the existence of an optimal policy in both finite and infinite time horizons. Moreover, we provide a sufficient condition for the optimality of deterministic actions. For illustration, we conclude the paper with examples from optimal liquidation with limit order books and autonomous driving.

math.OC

The Price of Information

When an investor is faced with the option to purchase additional information regarding an asset price, how much should she pay? To address this question, we solve for the indifference price of information in a setting where a trader maximizes her expected utility of terminal wealth over a finite time horizon. If she does not purchase the information, then she solves a partial information stochastic control problem, while, if she does purchase the information, then she pays a cost and receives partial information about the asset's trajectory. We further demonstrate that when the investor can purchase the information at any stopping time prior to the end of the trading horizon, she chooses to do so at a deterministic time.

q-fin.MF

Exploratory Control with Tsallis Entropy for Latent Factor Models

We study optimal control in models with latent factors where the agent controls the distribution over actions, rather than actions themselves, in both discrete and continuous time. To encourage exploration of the state space, we reward exploration with Tsallis Entropy and derive the optimal distribution over states - which we prove is $q$-Gaussian distributed with location characterized through the solution of an FBS$Δ$E and FBSDE in discrete and continuous time, respectively. We discuss the relation between the solutions of the optimal exploration problems and the standard dynamic optimal control solution. Finally, we develop the optimal policy in a model-agnostic setting along the lines of soft $Q$-learning. The approach may be applied in, e.g., developing more robust statistical arbitrage trading strategies.

q-fin.MF

FuNVol: A Multi-Asset Implied Volatility Market Simulator using Functional Principal Components and Neural SDEs

We introduce a new approach for generating sequences of implied volatility (IV) surfaces across multiple assets that is faithful to historical prices. We do so using a combination of functional data analysis and neural stochastic differential equations (SDEs) combined with a probability integral transform penalty to reduce model misspecification. We demonstrate that learning the joint dynamics of IV surfaces and prices produces market scenarios that are consistent with historical features and lie within the sub-manifold of surfaces that are essentially free of static arbitrage. Finally, we demonstrate that delta hedging using the simulated surfaces generates profit and loss (P&L) distributions that are consistent with realised P&Ls.

q-fin.CP

Robust Risk-Aware Option Hedging

The objectives of option hedging/trading extend beyond mere protection against downside risks, with a desire to seek gains also driving agent's strategies. In this study, we showcase the potential of robust risk-aware reinforcement learning (RL) in mitigating the risks associated with path-dependent financial derivatives. We accomplish this by leveraging a policy gradient approach that optimises robust risk-aware performance criteria. We specifically apply this methodology to the hedging of barrier options, and highlight how the optimal hedging strategy undergoes distortions as the agent moves from being risk-averse to risk-seeking. As well as how the agent robustifies their strategy. We further investigate the performance of the hedge when the data generating process (DGP) varies from the training DGP, and demonstrate that the robust strategies outperform the non-robust ones.

q-fin.CP

Stressing Dynamic Loss Models

Stress testing, and in particular, reverse stress testing, is a prominent exercise in risk management practice. Reverse stress testing, in contrast to (forward) stress testing, aims to find an alternative but plausible model such that under that alternative model, specific adverse stresses (i.e. constraints) are satisfied. Here, we propose a reverse stress testing framework for dynamic models. Specifically, we consider a compound Poisson process over a finite time horizon and stresses composed of expected values of functions applied to the process at the terminal time. We then define the stressed model as the probability measure under which the process satisfies the constraints and which minimizes the Kullback-Leibler divergence to the reference compound Poisson model. We solve this optimization problem, prove existence and uniqueness of the stressed probability measure, and provide a characterization of the Radon-Nikodym derivative from the reference model to the stressed model. We find that under the stressed measure, the intensity and the severity distribution of the process depend on time and state. We illustrate the dynamic stress testing by considering stresses on VaR and both VaR and CVaR jointly and provide illustrations of how the stochastic process is altered under these stresses. We generalize the framework to multivariate compound Poisson processes and stresses at times other than the terminal time. We illustrate the applicability of our framework by considering ``what if'' scenarios, where we answer the question: What is the severity of a stress on a portfolio component at an earlier time such that the aggregate portfolio exceeds a risk threshold at the terminal time? Furthermore, for general constraints, we propose an algorithm to simulate sample paths under the stressed measure, thus allowing to compare the effects of stresses on the dynamics of the process.

q-fin.RM