SearcharxivSearch

arXiv subjects

Yadh Hafsi

Publications and source records attributed to Yadh Hafsi.

10 recordsLinked to original sources

Stochastic Maximum Principle for McKean-Vlasov Control with Discrete Path Dependence

We study a class of McKean-Vlasov control problems with discrete path dependence. The coefficients and cost functional may depend on finitely many past values of the controlled state, observed at fixed deterministic times, and on their joint law. We establish well-posedness of the controlled state equation and derive necessary and sufficient optimality conditions through a stochastic Pontryagin maximum principle. The adjoint process is characterized by a backward stochastic differential equation with jumps at fixed observation times. Each jump represents the conditional sensitivity of future costs with respect to the corresponding observed state and its distribution. In the linear-quadratic case, we prove global solvability of the resulting forward-backward system by combining a continuation argument with a mean-field Riccati reduction. We finally discuss an application to time-series generation, where the terminal cost is given by a kernel-based discrepancy between the law of the sampled path and a target path distribution.

math.OC

Deep-MKV-TS: Path-Dependent McKean--Vlasov Control for Financial Time Series Generation

We introduce Deep-MKV-TS, a path-dependent McKean-Vlasov framework for financial scenario generation. The stochastic dynamics are chosen by matching selected path and volatility features of generated scenarios to those observed in the data. Starting from an interpretable reference model, Deep-MKV-TS preserves the reference drift and adjusts its volatility, while a regularization penalty limits unnecessary departures from the calibrated dynamics. We solve the resulting control problem using a neural, sample-based implementation of the stochastic maximum principle. We validate the method against an exactly computable oracle. On Heston and Heston-mixture models, Deep-MKV-TS substantially reduces path-dependent and volatility-related deficiencies of the reference model. In delayed-volatility experiments, the correction remains effective as the forecasting horizon increases, while direct training becomes less reliable. On held-out intraday equity-index futures, the corrected model improves conditional forecasts relative to the reference and reaches a level of performance comparable to flexible generative and historical baselines. The resulting scenarios also support greater exposure than the reference under a fixed drawdown-risk target. These results show that path-dependent McKean-Vlasov control can enrich an interpretable reference model without replacing it.

q-fin.CP

Policy Gradient Learning for Distributionally Robust Markov Decision Processes under Wasserstein Ambiguity

We study finite-horizon Markov decision processes under distributional uncertainty in the transition kernels and develop a policy-gradient framework for Wasserstein distributionally robust control. Ambiguity is modeled by Wasserstein balls of common radius centered at state--action-dependent nominal transition kernels, leading to a max--min problem over randomized policies and admissible transition laws. Because the worst-case transition law depends implicitly on the policy parameters, the standard policy-gradient argument does not apply directly. We address this difficulty by combining the dynamic programming recursion with Wasserstein duality and a primal envelope argument. In general, the right and left directional derivatives of the one-step worst-case value are obtained by taking the minimum or maximum expected downstream value derivative over the set of worst-case transition laws. In finite state--action spaces, this set is characterized through the optimal face of a transport linear program, yielding an exact directional-derivative recursion. Under the required stability conditions and uniqueness of the active dual and transport optimizers, the derivative becomes linear in the policy perturbation and admits an explicit vector valued policy-gradient recursion. Building on this representation, we propose a robust actor--critic implementation and evaluate it on benchmark examples.

math.OC

Learning Generative Dynamics with Soft Law Constraints: A McKean-Vlasov FBSDE Approach

We propose a generative framework for learning stochastic dynamics from endpoint and intermediate distributional observations. The method formulates generation as a McKean-Vlasov control problem in which terminal and time-marginal laws are enforced through soft energy constraints. The associated optimality system is a forward-backward stochastic differential equation (FBSDE) whose backward component receives a continuous drift induced by the marginal law penalties. This provides a principled alternative to hard interpolation or optimal transport maps between observed distributions: the model learns a stochastic path law whose dynamics remain globally coupled through the mean-field objective. We derive the reduced FBSDE system for quadratic control cost and constant diffusion, connecting terminal and marginal law flat derivatives to score-like training signals. The resulting neural solver is evaluated on low-dimensional distributional benchmarks, where it recovers smooth stochastic paths matching prescribed marginal laws. In a higher-dimensional ALAE latent space, endpoint supervision is used as a qualitative stress test for transporting non-smiling faces toward smiling ones in a pretrained representation. We then use articulated human motion as a structured high-dimensional case study on a curated AMASS low-to-high position dataset, using SMPL-H pose sequences and reduced pose representations. The experiments show that soft marginal law constraints can produce coherent stochastic trajectories whose intermediate distributions follow the observed evolution of human motion. The code is available at https://github.com/murex/deep-mkv-gen/tree/main.

math.OC

Trading in CEXs and DEXs with Priority Fees and Stochastic Delays

We develop a mixed control framework that combines absolutely continuous controls with impulse interventions subject to stochastic execution delays. The model extends current impulse control formulations by allowing (i) the controller to choose the mean of the stochastic delay of their impulses, and allowing (ii) for multiple pending orders, so that several impulses can be submitted and executed asynchronously at random times. The framework is motivated by an optimal trading problem between centralized (CEX) and decentralized (DEX) exchanges. In DEXs, traders control the distribution of the execution delay through the priority fee paid, introducing a fundamental trade-off between delays, uncertainty, and costs. We study the optimal trading problem of an agent exploiting trading signals in CEXs and DEXs. From a mathematical perspective, we derive the associated dynamic programming principle of this new class of impulse control problems, and establish the viscosity properties of the corresponding quasi-variational inequalities. From a financial perspective, our model provides insights on how to carry out execution across CEXs and DEXs, highlighting how traders manage latency risk optimally through priority fee selection. We show that employing the optimal priority fee has a significant outperformance over non-strategic fee selection.

q-fin.TR

Reinforcement Learning in Queue-Reactive Models: Application to Optimal Execution

We investigate the use of Reinforcement Learning for the optimal execution of meta-orders, where the objective is to execute incrementally large orders while minimizing implementation shortfall and market impact over an extended period of time. Departing from traditional parametric approaches to price dynamics and impact modeling, we adopt a model-free, data-driven framework. Since policy optimization requires counterfactual feedback that historical data cannot provide, we employ the Queue-Reactive Model to generate realistic and tractable limit order book simulations that encompass transient price impact, and nonlinear and dynamic order flow responses. Methodologically, we train a Double Deep Q-Network agent on a state space comprising time, inventory, price, and depth variables, and evaluate its performance against established benchmarks. Numerical simulation results show that the agent learns a policy that is both strategic and tactical, adapting effectively to order book conditions and outperforming standard approaches across multiple training configurations. These findings provide strong evidence that model-free Reinforcement Learning can yield adaptive and robust solutions to the optimal execution problem.

q-fin.TR

Optimal Execution under Liquidity Uncertainty

We study an optimal execution strategy for purchasing a large block of shares over a fixed time horizon. The execution problem is subject to a general price impact that gradually dissipates due to market resilience. We allow for general limit order book shapes to characterize instantaneous market impact. To model the resilience dynamics, we introduce a stochastic process that governs the rate at which the deviation between the impacted and unaffected prices decays. This volume-effect process reflects fluctuations in market activity that drive the pace of liquidity replenishment. Additionally, we incorporate stochastic liquidity variations through a regime-switching Markov chain to capture abrupt shifts in market conditions. We study this singular control problem, where the trader optimally determines the timing and rate of purchases to minimize execution costs. The associated value function to this optimization problem is shown to satisfy a system of variational Hamilton-Jacobi-Bellman inequalities. Moreover, we establish that it is the unique viscosity solution to this HJB system and study the analytical properties of the free boundary separating the execution and continuation regions. To illustrate our results, we present numerical examples under different limit-order book configurations, highlighting the interplay between price impact, resilience dynamics, and stochastic liquidity regimes in shaping the optimal execution strategy.

q-fin.MF

Optimal Execution with Reinforcement Learning

This study investigates the development of an optimal execution strategy through reinforcement learning, aiming to determine the most effective approach for traders to buy and sell inventory within a finite time horizon. Our proposed model leverages input features derived from the current state of the limit order book and operates at a high frequency to maximize control. To simulate this environment and overcome the limitations associated with relying on historical data, we utilize the multi-agent market simulator ABIDES, which provides a diverse range of depth levels within the limit order book. We present a custom MDP formulation followed by the results of our methodology and benchmark the performance against standard execution strategies. Results show that the reinforcement learning agent outperforms standard strategies and offers a practical foundation for real-world trading applications.

q-fin.TR

Optimal Execution under Incomplete Information

We study optimal liquidation strategies under partial information for a single asset within a finite time horizon. We propose a model tailored for high-frequency trading, capturing price formation driven solely by order flow through mutually stimulating marked Hawkes processes. The model assumes a limit order book framework, accounting for both permanent price impact and transient market impact. Importantly, we incorporate liquidity as a hidden Markov process, influencing the intensities of the point processes governing bid and ask prices. Within this setting, we formulate the optimal liquidation problem as an impulse control problem. We elucidate the dynamics of the hidden Markov chain's filter and determine the related normalized filtering equations. We then express the value function as the limit of a sequence of auxiliary continuous functions, defined recursively. This characterization enables the use of a dynamic programming principle for optimal stopping problems and the determination of an optimal strategy. It also facilitates the development of an implementable algorithm to approximate the original liquidation problem. We enrich our analysis with numerical results and visualizations of candidate optimal strategies.

q-fin.MF

Uncovering Market Disorder and Liquidity Trends Detection

The primary objective of this paper is to conceive and develop a new methodology to detect notable changes in liquidity within an order-driven market. We study a market liquidity model which allows us to dynamically quantify the level of liquidity of a traded asset using its limit order book data. The proposed metric holds potential for enhancing the aggressiveness of optimal execution algorithms, minimizing market impact and transaction costs, and serving as a reliable indicator of market liquidity for market makers. As part of our approach, we employ Marked Hawkes processes to model trades-through which constitute our liquidity proxy. Subsequently, our focus lies in accurately identifying the moment when a significant increase or decrease in its intensity takes place. We consider the minimax quickest detection problem of unobservable changes in the intensity of a doubly-stochastic Poisson process. The goal is to develop a stopping rule that minimizes the robust Lorden criterion, measured in terms of the number of events until detection, for both worst-case delay and false alarm constraint. We prove our procedure's optimality in the case of a Cox process with simultaneous jumps, while considering a finite time horizon. Finally, this novel approach is empirically validated by means of real market data analyses.

q-fin.MF