SearcharxivSearch

arXiv subjects

Emmanuel Bacry

Publications and source records attributed to Emmanuel Bacry.

At least 19 recordsLinked to original sources

The Log S-fBM model: Statistical analysis

The Log S-fBM model, introduced by Wu et al., is a stochastic volatility model whose log volatility is a stationary fractional Brownian motion (S-fBM): a stationary Gaussian process with power-decaying autocovariance driven by the Hurst exponent $H$, and variance scaled by an intermittency coefficient. A key property is that it reconciles rough volatility, where $H$ is typically near $0.1$ (see Gatheral et al.), with multifractal volatility, where $H$ is close to $0$ as in Bacry, Muzy et al.: the model's volatility measure converges to a multifractal random measure as $H\to0$. Numerical findings in Wu et al. show intermittency of order $0.02$ across financial assets, motivating a small intermittency approximation of log volatility moments for calibration via the general method of moments (GMM). In this work, we conduct a statistical analysis of the Log S-fBM model. We derive scaling properties of the S-fBM process and the Log S-fBM integrated volatility measure, present deviation inequalities with tail distributions sensitive to $H$ and intermittency, and develop a hypothesis test for the null Hurst exponent, i.e.\ rough versus multifractal dynamics. Finally, we revisit scale invariance of the log volatility increment process via explicit small-intermittency formulas, reproducing analogous properties in both regimes.

q-fin.ST

PARHAF, a human-authored corpus of clinical reports for fictitious patients in French

The development of clinical natural language processing (NLP) systems is severely hampered by the sensitive nature of medical records, which restricts data sharing under stringent privacy regulations, particularly in France and the broader European Union. To address this gap, we introduce PARHAF, a large open-source corpus of clinical documents in French. PARHAF comprises expert-authored clinical reports describing realistic yet entirely fictitious patient cases, making it anonymous and freely shareable by design. The corpus was developed using a structured protocol that combined clinician expertise with epidemiological guidance from the French National Health Data System (SNDS), ensuring broad clinical coverage. A total of 104 medical residents across 18 specialties authored and peer-reviewed the reports following predefined clinical scenarios and document templates. The corpus contains 7394 clinical reports covering 5009 patient cases across a wide range of medical and surgical specialties. It includes a general-purpose component designed to approximate real-world hospitalization distributions, and four specialized subsets that support information-extraction use cases in oncology, infectious diseases, and diagnostic coding. Documents are released under a CC-BY open license, with a portion temporarily embargoed to enable future benchmarking under controlled conditions. PARHAF provides a valuable resource for training and evaluating French clinical language models in a fully privacy-preserving setting, and establishes a replicable methodology for building shareable synthetic clinical corpora in other languages and health systems.

cs.CL

From rough to multifractal multidimensional volatility: A multidimensional Log S-fBM model

We introduce the multivariate Log S-fBM model (mLog S-fBM), extending the univariate framework proposed by Wu \textit{et al.} to the multidimensional setting. We define the multidimensional Stationary fractional Brownian motion (mS-fBM), characterized by marginals following S-fBM dynamics and a specific cross-covariance structure. It is parametrized by a correlation scale $T$, marginal-specific intermittency parameters and Hurst exponents, as well as their multidimensional counterparts: the co-intermittency matrix and the co-Hurst matrix. The mLog S-fBM is constructed by modeling volatility components as exponentials of the mS-fBM, preserving the dependence structure of the Gaussian core. We demonstrate that the model is well-defined for any co-Hurst matrix with entries in $[0, \frac{1}{2}[$, supporting vanishing co-Hurst parameters to bridge rough volatility and multifractal regimes. We generalize the small intermittency approximation technique to the multivariate setting to develop an efficient Generalized Method of Moments calibration procedure, estimating cross-covariance parameters for pairs of marginals. We validate it on synthetic data and apply it to S\&P 500 market data, modeling stock return fluctuations. Diagonal estimates of the stock Hurst matrix, corresponding to single-stock log-volatility Hurst exponents, are close to 0, indicating multifractal behavior, while co-Hurst off-diagonal entries are close to the Hurst exponent of the S\&P 500 index ($H \approx 0.12$), and co-intermittency off-diagonal entries align with univariate intermittency estimates.

q-fin.ST

KANFormer for Predicting Fill Probabilities via Survival Analysis in Limit Order Books

This paper introduces KANFormer, a novel deep-learning-based model for predicting the time-to-fill of limit orders by leveraging both market- and agent-level information. KANFormer combines a Dilated Causal Convolutional network with a Transformer encoder, enhanced by Kolmogorov-Arnold Networks (KANs), which improve nonlinear approximation. Unlike existing models that rely solely on a series of snapshots of the limit order book, KANFormer integrates the actions of agents related to LOB dynamics and the position of the order in the queue to more effectively capture patterns related to execution likelihood. We evaluate the model using CAC 40 index futures data with labeled orders. The results show that KANFormer outperforms existing works in both calibration (Right-Censored Log-Likelihood, Integrated Brier Score) and discrimination (C-index, time-dependent AUC). We further analyze feature importance over time using SHAP (SHapley Additive exPlanations). Our results highlight the benefits of combining rich market signals with expressive neural architectures to achieve accurate and interpretabl predictions of fill probabilities.

cs.AI

A Nested Factor Model for Equity Markets: Reconciling Multifractal Stock Returns and Rough Index Volatilities

The Nested factor model was introduced by Chicheportiche et al. to represent non-linear correlations between stocks. Stock returns are explained by a standard factor model, but the (log)-volatilities of factors and residuals are themselves decomposed into factor modes, with a common dominant volatility mode affecting both market and sector factors but also residuals. Here, we consider the case of a single factor where the only dominant log-volatility mode is rough, with a Hurst exponent $H \simeq 0.11$ and the log-volatility residuals are ''super-rough'' or ''multifractal'', with $H \simeq 0$. We demonstrate that such a construction naturally accounts for the somewhat surprising stylized fact reported by Wu et al. , where it has been observed that the Hurst exponents of stock indexes are large compared to those of individual stocks. We propose a statistical procedure to estimate the Hurst factor exponent from the stock returns dynamics together with theoretical guarantees of its consistency. We demonstrate the effectiveness of our approach through numerical experiments and apply it to daily stock data from the S&P500 index. The estimated roughness exponents for both the factor and idiosyncratic components validate the assumptions underlying our model.

q-fin.ST

No Tick-Size Too Small: A General Method for Modelling Small Tick Limit Order Books

Tick-sizes not only influence the granularity of the price formation process but also affect market agents' behavior. We investigate the disparity in the microstructural properties of the Limit Order Book (LOB) across a basket of assets with different relative tick-sizes. A key contribution of this study is the identification of several stylized facts, which are used to differentiate between large, medium, and small-tick assets, along with clear metrics for their measurement. We provide cross-asset visualizations to illustrate how these attributes vary with relative tick-size. Further, we propose a Hawkes Process model that {\color{black}not only fits well for large-tick assets, but also accounts for }sparsity, multi-tick level price moves, and the shape of the LOB in small-tick assets. Through simulation studies, we demonstrate the {\color{black} versatility} of the model and identify key variables that determine whether a simulated LOB resembles a large-tick or small-tick asset. Our tests show that stylized facts like sparsity, shape, and relative returns distribution can be smoothly transitioned from a large-tick to a small-tick asset using our model. We test this model's assumptions, showcase its challenges and propose questions for further directions in this area of research.

q-fin.TR

Validation of a new, minimally-invasive, software smartphone device to predict sleep apnea and its severity: transversal study

Obstructive sleep apnea (OSA) is frequent and responsible for cardiovascular complications and excessive daytime sleepiness. It is underdiagnosed due to the difficulty to access the gold standard for diagnosis, polysomnography (PSG). Alternative methods using smartphone sensors could be useful to increase diagnosis. The objective is to assess the performances of Apneal, an application that records the sound using a smartphone's microphone and movements thanks to a smartphone's accelerometer and gyroscope, to estimate patients' AHI. In this article, we perform a monocentric proof-of-concept study with a first manual scoring step, and then an automatic detection of respiratory events from the recorded signals using a sequential deep-learning model which was released internally at Apneal at the end of 2022 (version 0.1 of Apneal automatic scoring of respiratory events), in adult patients during in-hospital polysomnography.46 patients (women 34 per cent, mean BMI 28.7 kg per m2) were included. For AHI superior to 15, sensitivity of manual scoring was 0.91, and positive predictive value (PPV) 0.89. For AHI superior to 30, sensitivity was 0.85, PPV 0.94. We obtained an AUC-ROC of 0.85 and an AUC-PR of 0.94 for the identification of AHI superior to 15, and AUC-ROC of 0.95 and AUC-PR of 0.93 for AHI superior to 30. Promising results are obtained for the automatic annotations of events.This article shows that manual scoring of smartphone-based signals is possible and accurate compared to PSG-based scorings. Automatic scoring method based on a deep learning model provides promising results. A larger multicentric validation study, involving subjects with different SAHS severity is required to confirm these results.

eess.SP

Liquidity takers behavior representation through a contrastive learning approach

Thanks to the access to the labeled orders on the CAC40 data from Euronext, we are able to analyze agents' behaviors in the market based on their placed orders. In this study, we construct a self-supervised learning model using triplet loss to effectively learn the representation of agent market orders. By acquiring this learned representation, various downstream tasks become feasible. In this work, we utilize the K-means clustering algorithm on the learned representation vectors of agent orders to identify distinct behavior types within each cluster.

q-fin.ST

The self-exciting nature of the bid-ask spread dynamics

The bid-ask spread, which is defined by the difference between the best selling price and the best buying price in a Limit Order Book at a given time, is a crucial factor in the analysis of financial securities. In this study, we propose a "State-dependent Spread Hawkes model" (SDSH) that accounts for various spread jump sizes and incorporates the impact of the current spread state on its intensity functions. We apply this model to the high-frequency data from the Cac40 Euronext market and capture several statistical properties, such as the spread distributions, inter-event time distributions, and spread autocorrelation functions. We illustrate the ability of the SDSH model to forecast spread values at short-term horizons.

q-fin.TR

From Rough to Multifractal volatility: the log S-fBM model

We introduce a family of random measures $M_{H,T} (d t)$, namely log S-fBM, such that, for $H>0$, $M_{H,T}(d t) = e^{\omega_{H,T}(t)} d t$ where $\omega_{H,T}(t)$ is a Gaussian process that can be considered as a stationary version of an $H$-fractional Brownian motion. Moreover, when $H \to 0$, one has $M_{H,T}(d t) \rightarrow {\widetilde M}_{T}(d t)$ (in the weak sense) where ${\widetilde M}_{T}(d t)$ is the celebrated log-normal multifractal random measure (MRM). Thus, this model allows us to consider, within the same framework, the two popular classes of multifractal ($H = 0$) and rough volatility ($0<H < 1/2$) models. The main properties of the log S-fBM are discussed and their estimation issues are addressed. We notably show that the direct estimation of $H$ from the scaling properties of $\ln(M_{H,T}([t, t+\tau]))$, at fixed $\tau$, can lead to strongly over-estimating the value of $H$. We propose a better GMM estimation method which is shown to be valid in the high-frequency asymptotic regime. When applied to a large set of empirical volatility data, we observe that stock indices have values around $H=0.1$ while individual stocks are characterized by values of $H$ that can be very close to $0$ and thus well described by a MRM. We also bring evidence that unlike the log-volatility variance $\nu^2$ whose estimation appears to be poorly reliable (though used widely in the rough volatility literature), the estimation of the so-called "intermittency coefficient" $\lambda^2$, which is the product of $\nu^2$ and the Hurst exponent $H$, appears to be far more reliable leading to values that seem to be universal for respectively all individual stocks and all stock indices.

q-fin.ST

Predicting the Solar Potential of Rooftops using Image Segmentation and Structured Data

Estimating the amount of electricity that can be produced by rooftop photovoltaic systems is a time-consuming process that requires on-site measurements, a difficult task to achieve on a large scale. In this paper, we present an approach to estimate the solar potential of rooftops based on their location and architectural characteristics, as well as the amount of solar radiation they receive annually. Our technique uses computer vision to achieve semantic segmentation of roof sections and roof objects on the one hand, and a machine learning model based on structured building features to predict roof pitch on the other hand. We then compute the azimuth and maximum number of solar panels that can be installed on a rooftop with geometric approaches. Finally, we compute precise shading masks and combine them with solar irradiation data that enables us to estimate the yearly solar potential of a rooftop.

cs.CV

About contrastive unsupervised representation learning for classification and its convergence

Contrastive representation learning has been recently proved to be very efficient for self-supervised training. These methods have been successfully used to train encoders which perform comparably to supervised training on downstream classification tasks. A few works have started to build a theoretical framework around contrastive learning in which guarantees for its performance can be proven. We provide extensions of these results to training with multiple negative samples and for multiway classification. Furthermore, we provide convergence guarantees for the minimization of the contrastive training error with gradient descent of an overparametrized deep neural encoder, and provide some numerical experiments that complement our theoretical findings

cs.LG

ZiMM: a deep learning model for long term and blurry relapses with non-clinical claims data

This paper considers the problems of modeling and predicting a long-term and ``blurry'' relapse that occurs after a medical act, such as a surgery. The relapse is observed only indirectly, in a ``blurry'' fashion, through longitudinal prescriptions of drugs over a long period of time after the medical act. We introduce a new model, called ZiMM (Zero-inflated Mixture of Multinomial distributions) in order to capture long-term and blurry relapses. On top of it, we build an end-to-end deep-learning architecture called ZiMM Encoder-Decoder (ZiMM ED) that can learn from the complex, irregular, highly heterogeneous and sparse patterns of health events that are observed through a claims-only database. ZiMM ED is applied on a ``non-clinical'' claims database, that contains only timestamped reimbursement codes for drug purchases, medical procedures and hospital diagnoses, the only available clinical feature being the age of the patient. This setting is more challenging than a setting where bedside clinical signals are available. Our motivation for using such a non-clinical claims database is its exhaustivity population-wise, compared to clinical electronic health records coming from a single or a small set of hospitals. Indeed, we consider a dataset containing the claims of almost \emph{all French citizens} who had surgery for prostatic problems, with a history between 1.5 and 5 years. We consider a long-term (18 months) relapse (urination problems still occur despite surgery), which is blurry since it is observed only through the reimbursement of a specific set of drugs for urination problems. Our experiments show that ZiMM ED improves several baselines, including non-deep learning and deep-learning approaches, and that it allows working on such a dataset with minimal preprocessing work.

cs.LG

SCALPEL3: a scalable open-source library for healthcare claims databases

This article introduces SCALPEL3, a scalable open-source framework for studies involving Large Observational Databases (LODs). Its design eases medical observational studies thanks to abstractions allowing concept extraction, high-level cohort manipulation, and production of data formats compatible with machine learning libraries. SCALPEL3 has successfully been used on the SNDS database (see Tuppin et al. (2017)), a huge healthcare claims database that handles the reimbursement of almost all French citizens. SCALPEL3 focuses on scalability, easy interactive analysis and helpers for data flow analysis to accelerate studies performed on LODs. It consists of three open-source libraries based on Apache Spark. SCALPEL-Flattening allows denormalization of the LOD (only SNDS for now) by joining tables sequentially in a big table. SCALPEL-Extraction provides fast concept extraction from a big table such as the one produced by SCALPEL-Flattening. Finally, SCALPEL-Analysis allows interactive cohort manipulations, monitoring statistics of cohort flows and building datasets to be used with machine learning libraries. The first two provide a Scala API while the last one provides a Python API that can be used in an interactive environment. Our code is available on GitHub. SCALPEL3 allowed to extract successfully complex concepts for studies such as Morel et al (2017) or studies with 14.5 million patients observed over three years (corresponding to more than 15 billion healthcare events and roughly 15 TeraBytes of data) in less than 49 minutes on a small 15 nodes HDFS cluster. SCALPEL3 provides a sharp interactive control of data processing through legible code, which helps to build studies with full reproducibility, leading to improved maintainability and audit of studies performed on LODs.

cs.DC

Queue-reactive Hawkes models for the order flow

In this work we introduce two variants of multivariate Hawkes models with an explicit dependency on various queue sizes aimed at modeling the stochastic time evolution of a limit order book. The models we propose thus integrate the influence of both the current book state and the past order flow. The first variant considers the flow of order arrivals at a specific price level as independent from the other one and describes this flow by adding a Hawkes component to the arrival rates provided by the continuous time Markov "Queue Reactive" model of Huang et al. Empirical calibration using Level-I order book data from Eurex future assets (Bund and DAX) show that the Hawkes term dramatically improves the pure "Queue-Reactive" model not only for the description of the order flow properties (as e.g. the statistics of inter-event times) but also with respect to the shape of the queue distributions. The second variant we introduce describes the joint dynamics of all events occurring at best bid and ask sides of some order book during a trading day. This model can be considered as a queue dependent extension of the multivariate Hawkes order-book model of Bacry et al. We provide an explicit way to calibrate this model either with a Maximum-Likelihood method or with a Least-Square approach. Empirical estimation from Bund and DAX level-I order book data allow us to recover the main features of Hawkes interactions uncovered in Bacry et al. but also to unveil their joint dependence on bid and ask queue sizes. We notably find that while the market order or mid-price changes rates can mainly be functions on the volume imbalance this is not the case for the arrival rate of limit or cancel orders. Our findings also allows us to clearly bring to light various features that distinguish small and large tick assets.

q-fin.TR

Dual optimization for convex constrained objectives without the gradient-Lipschitz assumption

The minimization of convex objectives coming from linear supervised learning problems, such as penalized generalized linear models, can be formulated as finite sums of convex functions. For such problems, a large set of stochastic first-order solvers based on the idea of variance reduction are available and combine both computational efficiency and sound theoretical guarantees (linear convergence rates). Such rates are obtained under both gradient-Lipschitz and strong convexity assumptions. Motivated by learning problems that do not meet the gradient-Lipschitz assumption, such as linear Poisson regression, we work under another smoothness assumption, and obtain a linear convergence rate for a shifted version of Stochastic Dual Coordinate Ascent (SDCA) that improves the current state-of-the-art. Our motivation for considering a solver working on the Fenchel-dual problem comes from the fact that such objectives include many linear constraints, that are easier to deal with in the dual. Our approach and theoretical findings are validated on several datasets, for Poisson regression and another objective coming from the negative log-likelihood of the Hawkes process, which is a family of models which proves extremely useful for the modeling of information propagation in social networks and causality inference.

stat.ML

Disentangling and quantifying market participant volatility contributions

Thanks to the access to labeled orders on the Cac40 index future provided by Euronext, we are able to quantify market participants contributions to the volatility in the diffusive limit. To achieve this result we leverage the branching properties of Hawkes point processes. We find that fast intermediaries (e.g., market maker type agents) have a smaller footprint on the volatility than slower, directional agents. The branching structure of Hawkes processes allows us to examine also the degree of endogeneity of each agent behavior. We find that high-frequency traders are more endogenously driven than other types of agents.

q-fin.TR

Tick: a Python library for statistical learning, with a particular emphasis on time-dependent modelling

Tick is a statistical learning library for Python~3, with a particular emphasis on time-dependent models, such as point processes, and tools for generalized linear models and survival analysis. The core of the library is an optimization module providing model computational classes, solvers and proximal operators for regularization. tick relies on a C++ implementation and state-of-the-art optimization algorithms to provide very fast computations in a single node multi-core setting. Source code and documentation can be downloaded from https://github.com/X-DataInitiative/tick

stat.ML