SearcharxivSearch

arXiv subjects

Jonathan Donier

Publications and source records attributed to Jonathan Donier.

15 recordsLinked to original sources

Enabling Uncertainty Estimation in Iterative Neural Networks

Turning pass-through network architectures into iterative ones, which use their own output as input, is a well-known approach for boosting performance. In this paper, we argue that such architectures offer an additional benefit: The convergence rate of their successive outputs is highly correlated with the accuracy of the value to which they converge. Thus, we can use the convergence rate as a useful proxy for uncertainty. This results in an approach to uncertainty estimation that provides state-of-the-art estimates at a much lower computational cost than techniques like Ensembles, and without requiring any modifications to the original iterative model. We demonstrate its practical value by embedding it in two application domains: road detection in aerial images and the estimation of aerodynamic properties of 2D and 3D shapes.

cs.AI

DEBOSH: Deep Bayesian Shape Optimization

Graph Neural Networks (GNNs) can predict the performance of an industrial design quickly and accurately and be used to optimize its shape effectively. However, to fully explore the shape space, one must often consider shapes deviating significantly from the training set. For these, GNN predictions become unreliable, something that is often ignored. For optimization techniques relying on Gaussian Processes, Bayesian Optimization (BO) addresses this issue by exploiting their ability to assess their own accuracy. Unfortunately, this is harder to do when using neural networks because standard approaches to estimating their uncertainty can entail high computational loads and reduced model accuracy. Hence, we propose a novel uncertainty-based method tailored to shape optimization. It enables effective BO and increases the quality of the resulting shapes beyond that of state-of-the-art approaches.

cs.LG

HybridSDF: Combining Deep Implicit Shapes and Geometric Primitives for 3D Shape Representation and Manipulation

Deep implicit surfaces excel at modeling generic shapes but do not always capture the regularities present in manufactured objects, which is something simple geometric primitives are particularly good at. In this paper, we propose a representation combining latent and explicit parameters that can be decoded into a set of deep implicit and geometric shapes that are consistent with each other. As a result, we can effectively model both complex and highly regular shapes that coexist in manufactured objects. This enables our approach to manipulate 3D shapes in an efficient and precise manner.

cs.CV

The universality of skipping behaviours on music streaming platforms

A recent study of skipping behaviour on music streaming platforms has shown that the skip profile for a given song -- i.e. the measure of the skipping rate as a function of the time in the song -- can be seen as some intrinsic characteristic of the song, in the sense that it is both very specific and highly stable over time and geographical regions. In this paper, we take this analysis one step further by introducing a simple model of skip behaviours, in which the skip profile for a given song is viewed as the response to a small number of events that happen within it. In particular, it allows us to identify accurately the timing of the events that trigger skip responses, as well as the fraction of users who skip following each these events. Strikingly, the responses triggered by individual events appears to follow a temporal profile that is consistent across songs, genres, devices and listening contexts, suggesting that people react to musical surprises in a universal way.

cs.SI

Ensemble-based cover song detection

Audio-based cover song detection has received much attention in the MIR community in the recent years. To date, the most popular formulation of the problem has been to compare the audio signals of two tracks and to make a binary decision based on this information only. However, leveraging additional signals might be key if one wants to solve the problem at an industrial scale. In this paper, we introduce an ensemble-based method that approaches the problem from a many-to-many perspective. Instead of considering pairs of tracks in isolation, we consider larger sets of potential versions for a given composition, and create and exploit the graph of relationships between these tracks. We show that this can result in a significant improvement in performance, in particular when the number of existing versions of a given composition is large.

cs.SD

Scaling up deep neural networks: a capacity allocation perspective

Following the recent work on capacity allocation, we formulate the conjecture that the shattering problem in deep neural networks can only be avoided if the capacity propagation through layers has a non-degenerate continuous limit when the number of layers tends to infinity. This allows us to study a number of commonly used architectures and determine which scaling relations should be enforced in practice as the number of layers grows large. In particular, we recover the conditions of Xavier initialization in the multi-channel case, and we find that weights and biases should be scaled down as the inverse square root of the number of layers for deep residual networks and as the inverse square root of the desired memory length for recurrent networks.

cs.LG

Capacity allocation through neural network layers

Capacity analysis has been recently introduced as a way to analyze how linear models distribute their modelling capacity across the input space. In this paper, we extend the notion of capacity allocation to the case of neural networks with non-linear layers. We show that under some hypotheses the problem is equivalent to linear capacity allocation, within some extended input space that factors in the non-linearities. We introduce the notion of layer decoupling, which quantifies the degree to which a non-linear activation decouples its outputs, and show that it plays a central role in capacity allocation through layers. In the highly non-linear limit where decoupling is total, we show that the propagation of capacity throughout the layers follows a simple markovian rule, which turns into a diffusion PDE in the limit of deep networks with residual layers. This allows us to recover some known results about deep neural networks, such as the size of the effective receptive field, or why ResNets avoid the shattering problem.

cs.LG

Capacity allocation analysis of neural networks: A tool for principled architecture design

Designing neural network architectures is a task that lies somewhere between science and art. For a given task, some architectures are eventually preferred over others, based on a mix of intuition, experience, experimentation and luck. For many tasks, the final word is attributed to the loss function, while for some others a further perceptual evaluation is necessary to assess and compare performance across models. In this paper, we introduce the concept of capacity allocation analysis, with the aim of shedding some light on what network architectures focus their modelling capacity on, when used on a given task. We focus more particularly on spatial capacity allocation, which analyzes a posteriori the effective number of parameters that a given model has allocated for modelling dependencies on a given point or region in the input space, in linear settings. We use this framework to perform a quantitative comparison between some classical architectures on various synthetic tasks. Finally, we consider how capacity allocation might translate in non-linear settings.

cs.LG

Unravelling the trading invariance hypothesis

We confirm and substantially extend the recent empirical result of Andersen et al. \cite{Andersen2015}, where it is shown that the amount of risk $W$ exchanged in the E-mini S\&P futures market (i.e. price times volume times volatility) scales like the 3/2 power of the number of trades $N$. We show that this 3/2-law holds very precisely across 12 futures contracts and 300 single US stocks, and across a wide range of time scales. However, we find that the "trading invariant" $I=W/N^{3/2}$ proposed by Kyle and Obizhaeva is in fact quite different for different contracts, in particular between futures and single stocks. Our analysis suggests $I/{\cal C}$ as a more natural candidate, where $\cal C$ is the average spread cost of a trade, defined as the average of the trade size times the bid-ask spread. We also establish two more complex scaling laws for the volatility $σ$ and the traded volume $V$ as a function of $N$, that reveal the existence of a characteristic number of trades $N_0$ above which the expected behaviour $σ\sim \sqrt{N}$ and $V \sim N$ hold, but below which strong deviations appear, induced by the size of the~tick.

q-fin.TR

Why Do Markets Crash? Bitcoin Data Offers Unprecedented Insights

Crashes have fascinated and baffled many canny observers of financial markets. In the strict orthodoxy of the efficient market theory, crashes must be due to sudden changes of the fundamental valuation of assets. However, detailed empirical studies suggest that large price jumps cannot be explained by news and are the result of endogenous feedback loops. Although plausible, a clear-cut empirical evidence for such a scenario is still lacking. Here we show how crashes are conditioned by the market liquidity, for which we propose a new measure inspired by recent theories of market impact and based on readily available, public information. Our results open the possibility of a dynamical evaluation of liquidity risk and early warning signs of market instabilities, and could lead to a quantitative description of the mechanisms leading to market crashes.

q-fin.TR

Quadratic Hawkes processes for financial prices

We introduce and establish the main properties of QHawkes ("Quadratic" Hawkes) models. QHawkes models generalize the Hawkes price models introduced in E. Bacry et al. (2014), by allowing all feedback effects in the jump intensity that are linear and quadratic in past returns. A non-parametric fit on NYSE stock data shows that the off-diagonal component of the quadratic kernel indeed has a structure that standard Hawkes models fail to reproduce. Our model exhibits two main properties, that we believe are crucial in the modelling and the understanding of the volatility process: first, the model is time-reversal asymmetric, similar to financial markets whose time evolution has a preferred direction. Second, it generates a multiplicative, fat-tailed volatility process, that we characterize in detail in the case of exponentially decaying kernels, and which is linked to Pearson diffusions in the continuous limit. Several other interesting properties of QHawkes processes are discussed, in particular the fact that they can generate long memory without necessarily be at the critical point. Finally, we provide numerical simulations of our calibrated QHawkes model, which is indeed seen to reproduce, with only a small amount of quadratic non-linearity, the correct magnitude of fat-tails and time reversal asymmetry seen in empirical time series.

q-fin.TR

A Million Metaorder Analysis of Market Impact on the Bitcoin

We present a thorough empirical analysis of market impact on the Bitcoin/USD exchange market using a complete dataset that allows us to reconstruct more than one million metaorders. We empirically confirm the "square-root law'' for market impact, which holds on four decades in spite of the quasi-absence of statistical arbitrage and market marking strategies. We show that the square-root impact holds during the whole trajectory of a metaorder and not only for the final execution price. We also attempt to decompose the order flow into an "informed'' and "uninformed'' component, the latter leading to an almost complete long-term decay of impact. This study sheds light on the hypotheses and predictions of several market impact models recently proposed in the literature and promotes heterogeneous agent models as promising candidates to explain price impact on the Bitcoin market -- and, we believe, on other markets as well.

q-fin.TR

From Walras' auctioneer to continuous time double auctions: A general dynamic theory of supply and demand

In standard Walrasian auctions, the price of a good is defined as the point where the supply and demand curves intersect. Since both curves are generically regular, the response to small perturbations is linearly small. However, a crucial ingredient is absent of the theory, namely transactions themselves. What happens after they occur? To answer the question, we develop a dynamic theory for supply and demand based on agents with heterogeneous beliefs. When the inter-auction time is infinitely long, the Walrasian mechanism is recovered. When transactions are allowed to happen in continuous time, a peculiar property emerges: close to the price, supply and demand vanish quadratically, which we empirically confirm on the Bitcoin. This explains why price impact in financial markets is universally observed to behave as the square root of the excess volume. The consequences are important, as they imply that the very fact of clearing the market makes prices hypersensitive to small fluctuations.

q-fin.TR

A fully consistent, minimal model for non-linear market impact

We propose a minimal theory of non-linear price impact based on a linear (latent) order book approximation, inspired by diffusion-reaction models and general arguments. Our framework allows one to compute the average price trajectory in the presence of a meta-order, that consistently generalizes previously proposed propagator models. We account for the universally observed square-root impact law, and predict non-trivial trajectories when trading is interrupted or reversed. We prove that our framework is free of price manipulation, and that prices can be made diffusive (albeit with a generic short-term mean-reverting contribution). Our model suggests that prices can be decomposed into a transient "mechanical" impact component and a permanent "informational" component.

q-fin.TR

Market Impact with Autocorrelated Order Flow under Perfect Competition

Our goal in this paper is to study the market impact in a market in which the order flow is autocorrelated. We build a model which explains qualitatively and quantitatively the empirical facts observed so far concerning market impact. We define different notions of market impact, and show how they lead to the different price paths observed in the literature. For each one, under the assumption of perfect competition and information, we derive and explain the relationships between the correlations in the order flow, the shape of the market impact function while a meta-order is being executed, and the expected price after the completion. We also derive an expression for the decay of market impact after a trade, and show how it can result in a better liquidation strategy for an informed trader. We show how, in spite of auto-correlation in order-flow, prices can be martingales, and how price manipulation is ruled out even though the bare impact function is concave. We finally assess the cost of market impact and try to make a step towards optimal strategies.

q-fin.TR