Searcharxiv⌕ Search

arXiv subjects

Emmanuel Bacry

Publications and source records attributed to Emmanuel Bacry.

36 records · Page 2Linked to original sources

Tick: a Python library for statistical learning, with a particular emphasis on time-dependent modelling

Tick is a statistical learning library for Python~3, with a particular emphasis on time-dependent models, such as point processes, and tools for generalized linear models and survival analysis. The core of the library is an optimization module providing model computational classes, solvers and proximal operators for regularization. tick relies on a C++ implementation and state-of-the-art optimization algorithms to provide very fast computations in a single node multi-core setting. Source code and documentation can be downloaded from https://github.com/X-DataInitiative/tick

stat.ML↗

ConvSCCS: convolutional self-controlled case series model for lagged adverse event detection

With the increased availability of large databases of electronic health records (EHRs) comes the chance of enhancing health risks screening. Most post-marketing detections of adverse drug reaction (ADR) rely on physicians' spontaneous reports, leading to under reporting. To take up this challenge, we develop a scalable model to estimate the effect of multiple longitudinal features (drug exposures) on a rare longitudinal outcome. Our procedure is based on a conditional Poisson model also known as self-controlled case series (SCCS). We model the intensity of outcomes using a convolution between exposures and step functions, that are penalized using a combination of group-Lasso and total-variation. This approach does not require the specification of precise risk periods, and allows to study in the same model several exposures at the same time. We illustrate the fact that this approach improves the state-of-the-art for the estimation of the relative risks both on simulations and on a cohort of diabetic patients, extracted from the large French national health insurance database (SNIIRAM), a SQL database built around medical reimbursements of more than 65 million people. This work has been done in the context of a research partnership between Ecole Polytechnique and CNAMTS (in charge of SNIIRAM).

stat.AP↗

Analysis of order book flows using a nonparametric estimation of the branching ratio matrix

We introduce a new non parametric method that allows for a direct, fast and efficient estimation of the matrix of kernel norms of a multivariate Hawkes process, also called branching ratio matrix. We demonstrate the capabilities of this method by applying it to high-frequency order book data from the EUREX exchange. We show that it is able to uncover (or recover) various relationships between all the first level order book events associated with some asset when mapped to a 12-dimensional process. We then scale up the model so as to account for events on two assets simultaneously and we discuss the joint high-frequency dynamics.

q-fin.TR↗

Uncovering Causality from Multivariate Hawkes Integrated Cumulants

We design a new nonparametric method that allows one to estimate the matrix of integrated kernels of a multivariate Hawkes process. This matrix not only encodes the mutual influences of each nodes of the process, but also disentangles the causality relationships between them. Our approach is the first that leads to an estimation of this matrix without any parametric modeling and estimation of the kernels themselves. A consequence is that it can give an estimation of causality relationships between nodes (or users), based on their activity timestamps (on a social network for instance), without knowing or estimating the shape of the activities lifetime. For that purpose, we introduce a moment matching method that fits the third-order integrated cumulants of the process. We show on numerical experiments that our approach is indeed very robust to the shape of the kernels, and gives appealing results on the MemeTracker database.

stat.ML↗

SGD with Variance Reduction beyond Empirical Risk Minimization

We introduce a doubly stochastic proximal gradient algorithm for optimizing a finite average of smooth convex functions, whose gradients depend on numerically expensive expectations. Our main motivation is the acceleration of the optimization of the regularized Cox partial-likelihood (the core model used in survival analysis), but our algorithm can be used in different settings as well. The proposed algorithm is doubly stochastic in the sense that gradient steps are done using stochastic gradient descent (SGD) with variance reduction, where the inner expectations are approximated by a Monte-Carlo Markov-Chain (MCMC) algorithm. We derive conditions on the MCMC number of iterations guaranteeing convergence, and obtain a linear rate of convergence under strong convexity and a sublinear rate without this assumption. We illustrate the fact that our algorithm improves the state-of-the-art solver for regularized Cox partial-likelihood on several datasets from survival analysis.

stat.ML↗

Concentration for matrix martingales in continuous time and microscopic activity of social networks

This paper gives new concentration inequalities for the spectral norm of a wide class of matrix martingales in continuous time. These results extend previously established Freedman and Bernstein inequalities for series of random matrices to the class of continuous time processes. Our analysis relies on a new supermartingale property of the trace exponential proved within the framework of stochastic calculus. We provide also several examples that illustrate the fact that our results allow us to recover easily several formerly obtained sharp bounds for discrete time matrix martingales.

math.PR↗

The role of volume in order book dynamics: a multivariate Hawkes process analysis

We show that multivariate Hawkes processes coupled with the nonparametric estimation procedure first proposed in Bacry and Muzy (2015) can be successfully used to study complex interactions between the time of arrival of orders and their size, observed in a limit order book market. We apply this methodology to high-frequency order book data of futures traded at EUREX. Specifically, we demonstrate how this approach is amenable not only to analyze interplay between different order types (market orders, limit orders, cancellations) but also to include other relevant quantities, such as the order size, into the analysis, showing also that simple models assuming the independence between volume and time are not suitable to describe the data.

q-fin.TR↗

Mean-field inference of Hawkes point processes

We propose a fast and efficient estimation method that is able to accurately recover the parameters of a d-dimensional Hawkes point-process from a set of observations. We exploit a mean-field approximation that is valid when the fluctuations of the stochastic intensity are small. We show that this is notably the case in situations when interactions are sufficiently weak, when the dimension of the system is high or when the fluctuations are self-averaging due to the large number of past events they involve. In such a regime the estimation of a Hawkes process can be mapped on a least-squares problem for which we provide an analytic solution. Though this estimator is biased, we show that its precision can be comparable to the one of the Maximum Likelihood Estimator while its computation speed is shown to be improved considerably. We give a theoretical control on the accuracy of our new approach and illustrate its efficiency using synthetic datasets, in order to assess the statistical estimation error of the parameters.

cs.LG↗

Hawkes processes in finance

In this paper we propose an overview of the recent academic literature devoted to the applications of Hawkes processes in finance. Hawkes processes constitute a particular class of multivariate point processes that has become very popular in empirical high frequency finance this last decade. After a reminder of the main definitions and properties that characterize Hawkes processes, we review their main empirical applications to address many different problems in high frequency finance. Because of their great flexibility and versatility, we show that they have been successfully involved in issues as diverse as estimating the volatility at the level of transaction data, estimating the market stability, accounting for systemic risk contagion, devising optimal execution strategies or capturing the dynamics of the full order book.

q-fin.TR↗

Intermittent process analysis with scattering moments

Scattering moments provide nonparametric models of random processes with stationary increments. They are expected values of random variables computed with a nonexpansive operator, obtained by iteratively applying wavelet transforms and modulus nonlinearities, which preserves the variance. First- and second-order scattering moments are shown to characterize intermittency and self-similarity properties of multiscale processes. Scattering moments of Poisson processes, fractional Brownian motions, Lévy processes and multifractal random walks are shown to have characteristic decay. The Generalized Method of Simulated Moments is applied to scattering moments to estimate data generating models. Numerical applications are shown on financial time-series and on energy dissipation of turbulent flows.

stat.ME↗

Second order statistics characterization of Hawkes processes and non-parametric estimation

We show that the jumps correlation matrix of a multivariate Hawkes process is related to the Hawkes kernel matrix through a system of Wiener-Hopf integral equations. A Wiener-Hopf argument allows one to prove that this system (in which the kernel matrix is the unknown) possesses a unique causal solution and consequently that the second-order properties fully characterize a Hawkes process. The numerical inversion of this system of integral equations allows us to propose a fast and efficient method, which main principles were initially sketched in [Bacry and Muzy, 2013], to perform a non-parametric estimation of the Hawkes kernel matrix. In this paper, we perform a systematic study of this non-parametric estimation procedure in the general framework of marked Hawkes processes. We describe precisely this procedure step by step. We discuss the estimation error and explain how the values for the main parameters should be chosen. Various numerical examples are given in order to illustrate the broad possibilities of this estimation procedure ranging from 1-dimensional (power-law or non positive kernels) up to 3-dimensional (circular dependence) processes. A comparison to other non-parametric estimation procedures is made. Applications to high frequency trading events in financial markets and to earthquakes occurrence dynamics are finally considered.

stat.ME↗

Estimation of slowly decreasing Hawkes kernels: Application to high frequency order book modelling

We present a modified version of the non parametric Hawkes kernel estimation procedure studied in arXiv:1401.0903 that is adapted to slowly decreasing kernels. We show on numerical simulations involving a reasonable number of events that this method allows us to estimate faithfully a power-law decreasing kernel over at least 6 decades. We then propose a 8-dimensional Hawkes model for all events associated with the first level of some asset order book. Applying our estimation procedure to this model, allows us to uncover the main properties of the coupled dynamics of trade, limit and cancel orders in relationship with the mid-price variations.

q-fin.ST↗

Linear processes in high-dimension: phase space and critical properties

In this work we investigate the generic properties of a stochastic linear model in the regime of high-dimensionality. We consider in particular the Vector AutoRegressive model (VAR) and the multivariate Hawkes process. We analyze both deterministic and random versions of these models, showing the existence of a stable and an unstable phase. We find that along the transition region separating the two regimes, the correlations of the process decay slowly, and we characterize the conditions under which these slow correlations are expected to become power-laws. We check our findings with numerical simulations showing remarkable agreement with our predictions. We finally argue that real systems with a strong degree of self-interaction are naturally characterized by this type of slow relaxation of the correlations.

cond-mat.stat-mech↗

Market impacts and the life cycle of investors orders

In this paper, we use a database of around 400,000 metaorders issued by investors and electronically traded on European markets in 2010 in order to study market impact at different scales. At the intraday scale we confirm a square root temporary impact in the daily participation, and we shed light on a duration factor in $1/T^γ$ with $γ\simeq 0.25$. Including this factor in the fits reinforces the square root shape of impact. We observe a power-law for the transient impact with an exponent between $0.5$ (for long metaorders) and $0.8$ (for shorter ones). Moreover we show that the market does not anticipate the size of the meta-orders. The intraday decay seems to exhibit two regimes (though hard to identify precisely): a "slow" regime right after the execution of the meta-order followed by a faster one. At the daily time scale, we show price moves after a metaorder can be split between realizations of expected returns that have triggered the investing decision and an idiosynchratic impact that slowly decays to zero. Moreover we propose a class of toy models based on Hawkes processes (the Hawkes Impact Models, HIM) to illustrate our reasoning. We show how the Impulsive-HIM model, despite its simplicity, embeds appealing features like transience and decay of impact. The latter is parametrized by a parameter $C$ having a macroscopic interpretation: the ratio of contrarian reaction (i.e. impact decay) and of the "herding" reaction (i.e. impact amplification).

q-fin.TR↗

Scaling limits for Hawkes processes and application to financial statistics

We prove a law of large numbers and a functional central limit theorem for multivariate Hawkes processes observed over a time interval $[0,T]$ in the limit $T \rightarrow \infty$. We further exhibit the asymptotic behaviour of the covariation of the increments of the components of a multivariate Hawkes process, when the observations are imposed by a discrete scheme with mesh $Δ$ over $[0,T]$ up to some further time shift $τ$. The behaviour of this functional depends on the relative size of $Δ$ and $τ$ with respect to $T$ and enables to give a full account of the second-order structure. As an application, we develop our results in the context of financial statistics. We introduced in a previous work a microscopic stochastic model for the variations of a multivariate financial asset, based on Hawkes processes and that is confined to live on a tick grid. We derive and characterise the exact macroscopic diffusion limit of this model and show in particular its ability to reproduce important empirical stylised fact such as the Epps effect and the lead-lag effect. Moreover, our approach enable to track these effects across scales in rigorous mathematical terms.

math.PR↗

The nature of price returns during periods of high market activity

By studying all the trades and best bids/asks of ultra high frequency snapshots recorded from the order books of a basket of 10 futures assets, we bring qualitative empirical evidence that the impact of a single trade depends on the intertrade time lags. We find that when the trading rate becomes faster, the return variance per trade or the impact, as measured by the price variation in the direction of the trade, strongly increases. We provide evidence that these properties persist at coarser time scales. We also show that the spread value is an increasing function of the activity. This suggests that order books are more likely empty when the trading rate is high.

q-fin.TR↗

Multifractal analysis in a mixed asymptotic framework

Multifractal analysis of multiplicative random cascades is revisited within the framework of {\em mixed asymptotics}. In this new framework, statistics are estimated over a sample which size increases as the resolution scale (or the sampling period) becomes finer. This allows one to continuously interpolate between the situation where one studies a single cascade sample at arbitrary fine scales and where at fixed scale, the sample length (number of cascades realizations) becomes infinite. We show that scaling exponents of ''mixed'' partitions functions i.e., the estimator of the cumulant generating function of the cascade generator distribution, depends on some ``mixed asymptotic'' exponent $χ$ respectively above and beyond two critical value $p_χ^-$ and $p_χ^+$. We study the convergence properties of partition functions in mixed asymtotics regime and establish a central limit theorem. These results are shown to remain valid within a general wavelet analysis framework. Their interpretation in terms of Besov frontier are discussed. Moreover, within the mixed asymptotic framework, we establish a ``box-counting'' multifractal formalism that can be seen as a rigorous formulation of Mandelbrot's negative dimension theory. Numerical illustrations of our purpose on specific examples are also provided.

math.PR↗

Extreme values and fat tails of multifractal fluctuations

In this paper we discuss the problem of the estimation of extreme event occurrence probability for data drawn from some multifractal process. We also study the heavy (power-law) tail behavior of probability density function associated with such data. We show that because of strong correlations, standard extreme value approach is not valid and classical tail exponent estimators should be interpreted cautiously. Extreme statistics associated with multifractal random processes turn out to be characterized by non self-averaging properties. Our considerations rely upon some analogy between random multiplicative cascades and the physics of disordered systems and also on recent mathematical results about the so-called multifractal formalism. Applied to financial time series, our findings allow us to propose an unified framemork that accounts for the observed multiscaling properties of return fluctuations, the volatility clustering phenomenon and the observed ``inverse cubic law'' of the return pdf tails.

cond-mat.stat-mech↗