SearcharxivSearch

arXiv subjects

Anna Bonnet

Publications and source records attributed to Anna Bonnet.

16 recordsLinked to original sources

Independent Component Discovery in Temporal Count Data

Advances in data collection are producing growing volumes of temporal count observations, making adapted modeling increasingly necessary. In this work, we introduce a generative framework for independent component analysis of temporal count data, combining regime-adaptive dynamics with Poisson log-normal emissions. The model identifies disentangled components with regime-dependent contributions, enabling representation learning and perturbations analysis. Notably, we establish the identifiability of the model, supporting principled interpretation. To learn the parameters, we propose an efficient amortized variational inference procedure. Experiments on simulated data evaluate recovery of the mixing function and latent sources across diverse settings, while real-world applications to gut microbiome and climate datasets reveal co-variation patterns and regime shifts consistent with domain-specific knowledge.

stat.ME

Hawkes process with a diffusion-driven baseline: long-run behavior, inference, statistical tests

Event-driven systems in fields such as neuroscience, social networks, and finance often exhibit dynamics influenced by continuously evolving external covariates. Motivated by these applications, we introduce a new class of multivariate Hawkes processes, in which the spontaneous rate of events is modulated by a diffusion process. This framework allows the point process to adapt dynamically to continuously evolving covariates, capturing both intrinsic self-excitation and external influences. In this article, we establish the probabilistic properties of the coupled process, proving stability and ergodicity under moderate assumptions. Classical functional results, including law of large numbers and mixing properties, are extended to this diffusion-driven setting. Building on these results, we study parametric inference for the Hawkes component: we derive consistency and asymptotic normality of the maximum likelihood estimator in the long-time regime, and derive stronger convergence results under additional assumptions on the covariate process. We further propose hypothesis testing procedures to assess the statistical relevance of the covariate. Simulation studies illustrate the validity of the asymptotic results and the effectiveness of the proposed inference methods. Overall, this work provides theoretical and practical foundations for diffusion-driven Hawkes models.

math.ST

Hawkes Processes with Variable Length Memory: Existence, Inference and Application to Neuronal Activity

Multivariate Hawkes processes are past-dependant point processes originally introduced to model excitation effects, later extended to a nonlinear framework to account for the opposite effect, known as inhibition. Motivated by applications in neuroscience, where the memory of a neuron may reset upon firing, we introduce a new class of nonlinear Hawkes processes with variable length memory. Our model generalises classical Hawkes processes, with or without inhibition, describing the situation where the probability of an event occurring within a given subprocess may depend differently on the history before and after its last event. In particular, if the subprocess does not depend on the history before its last event, it is said to have a variable length memory. Our main contributions are to prove existence of such processes, and to derive a workable likelihood maximisation method, capable of identifying both classical and variable memory dynamics. We demonstrate the effectiveness of our approach both on synthetic data, and on a neuronal activity dataset.

stat.ME

A Markov switching discrete-time Hawkes process: application to the monitoring of bats behavior

Over the past few decades, the Hawkes process has become a popular framework for modeling temporal events thanks to its flexibility to capture different dependency structures. The objective of this work is to model call sequences emitted by bats for echolocation, whose patterns are known to change depending on the animal's activity. The novelty of the model lies in the combination of a Hawkes-type dependency from past events, as well as a latent variable that encodes changes in bat behavior. More precisely, we consider a discrete-time version of the Hawkes process, with an exponential kernel, where the immigration term varies according to a latent Markov chain. We prove that this model is identifiable and can be reformulated in terms of a Hidden Markov Model, with Poisson emissions. Based on these properties, we show that maximum likelihood inference of the model parameters can be performed using an EM algorithm, which involves a recursive M-step. A simulation study demonstrates the performance of our approach method for estimating the parameters, recovering the number of hidden states and classifying each bin of the trajectory. Finally, we illustrate the use of the proposed modeling to distinguish different behaviors of bats, based on the recording of their cries.

stat.ME

TaxaPLN: a taxonomy-aware augmentation strategy for microbiome-trait classification including metadata

The gut microbiome plays a crucial role in human health, making it a corner stone of modern biomedical research. To study its structure and dynamics, machine learning models are increasingly used to identify key microbial patterns associated with disease and environmental factors. However, microbiome data present unique challenges due to their compositionality, high-dimensionality, sparsity, and high variability, which can obscure meaningful signals. Besides, the effectiveness of machine learning models is often constrained by limited sample sizes, as microbiome data collection remains costly and time consuming. In this context, data augmentation has emerged as a promising strategy to enhance model robustness and predictive performance by generating artificial microbiome data. The aim of this study is to improve predictive modeling from microbiome data by introducing a model-based data augmentation approach that incorporates both taxonomic relationships and covariate information. To that end, we propose TaxaPLN, a data augmentation method built on PLN-Tree generative models, which leverages the taxonomy and a data-driven sampler to generate realistic synthetic microbiome compositions. We further introduce a conditional extension based on feature-wise linear modulation, enabling covariate-aware generation. Experiments on high-quality curated microbiome datasets show that TaxaPLN preserves ecological properties and generally improves or maintains predictive performances, particularly with non-linear classifiers, outperforming state-of-the-art baselines. Besides, TaxaPLN conditional augmentation establishes a novel benchmark for covariate-aware microbiome augmentation. The MIT-licensed source code is available at https://github.com/ AlexandreChaussard/PLNTree-package along with the datasets used in our experiments.

stat.AP

Nonparametric estimation of Hawkes processes with RKHSs

This paper addresses nonparametric estimation of nonlinear multivariate Hawkes processes, where the interaction functions are assumed to lie in a reproducing kernel Hilbert space (RKHS). Motivated by applications in neuroscience, the model allows complex interaction functions, in order to express exciting and inhibiting effects, but also a combination of both (which is particularly interesting to model the refractory period of neurons), and considers in return that conditional intensities are rectified by the ReLU function. The latter feature incurs several methodological challenges, for which workarounds are proposed in this paper. In particular, it is shown that a representer theorem can be obtained for approximated versions of the log-likelihood and the least-squares criteria. Based on it, we propose an estimation method, that relies on two common approximations (of the ReLU function and of the integral operator). We provide a bound that controls the impact of these approximations. Numerical results on synthetic data confirm this fact as well as the good asymptotic behavior of the proposed estimator. It also shows that our method achieves a better performance compared to related nonparametric estimation techniques and suits neuronal applications.

stat.ML

Testing procedures based on maximum likelihood estimation for Marked Hawkes processes

The Hawkes model is a past-dependent point process, widely used in various fields for modeling temporal clustering of events. Extending this framework, the multidimensional marked Hawkes process incorporates multiple interacting event types and additional marks, enhancing its capability to model complex dependencies in multivariate time series data. However, increasing the complexity of the model also increases the computational cost of the associated estimation methods and may induce an overfitting of the model. Therefore, it is essential to find a trade-off between accuracy and artificial complexity of the model. In order to find the appropriate version of Hawkes processes, we address, in this paper, the tasks of model fit evaluation and parameter testing for marked Hawkes processes. This article focuses on parametric Hawkes processes with exponential memory kernels, a popular variant for its theoretical and practical advantages. Our work introduces robust testing methodologies for assessing model parameters and complexity, building upon and extending previous theoretical frameworks. We then validate the practical robustness of these tests through comprehensive numerical studies, especially in scenarios where theoretical guarantees remains incomplete.

stat.ME

Tree-based variational inference for Poisson log-normal models

When studying ecosystems, hierarchical trees are often used to organize entities based on proximity criteria, such as the taxonomy in microbiology, social classes in geography, or product types in retail businesses, offering valuable insights into entity relationships. Despite their significance, current count-data models do not leverage this structured information. In particular, the widely used Poisson log-normal (PLN) model, known for its ability to model interactions between entities from count data, lacks the possibility to incorporate such hierarchical tree structures, limiting its applicability in domains characterized by such complexities. To address this matter, we introduce the PLN-Tree model as an extension of the PLN model, specifically designed for modeling hierarchical count data. By integrating structured variational inference techniques, we propose an adapted training procedure and establish identifiability results, enhancing both theoretical foundations and practical interpretability. Experiments on synthetic datasets and human gut microbiome data highlight generative improvements when using PLN-Tree, demonstrating the practical interest of knowledge graphs like the taxonomy in microbiome modeling. Additionally, we present a proof-of-concept implication of the identifiability results by illustrating the practical benefits of using identifiable features for classification tasks, showcasing the versatility of the framework.

stat.ME

Spectral analysis for the inference of noisy Hawkes processes

Classic estimation methods for Hawkes processes rely on the assumption that observed event times are indeed a realisation of a Hawkes process, without considering any potential perturbation of the model. However, in practice, observations are often altered by some noise, the form of which depends on the context. It is then required to model the alteration mechanism in order to infer accurately such a noisy Hawkes process. While several models exist, we consider, in this work, the observations to be the indistinguishable union of event times coming from a Hawkes process and from an independent Poisson process. Since standard inference methods (such as maximum likelihood or Expectation-Maximisation) are either unworkable or numerically prohibitive in this context, we propose an estimation procedure based on the spectral analysis of second order properties of the noisy Hawkes process. Novel results include sufficient conditions for identifiability of the ensuing statistical model with exponential interaction functions for both univariate and bivariate processes, along with consistency and asymptotic normality guarantees of our estimator in the univariate case. Although we mainly focus on the exponential scenario, other types of kernels are investigated and discussed. A new estimator based on maximising the spectral log-likelihood is then described, and its behaviour is numerically illustrated on both synthetic data and neuronal data. Besides being free from knowing the source of each observed time (Hawkes or Poisson process), the proposed estimator is shown to perform accurately in estimating both processes.

stat.ME

Inference of multivariate exponential Hawkes processes with inhibition and application to neuronal activity

The multivariate Hawkes process is a past-dependent point process used to model the relationship of event occurrences between different phenomena.Although the Hawkes process was originally introduced to describe excitation effects, which means that one event increases the chances of another occurring, there has been a growing interest in modelling the opposite effect, known as inhibition.In this paper, we focus on how to infer the parameters of a multidimensional exponential Hawkes process with both excitation and inhibition effects. Our first result is to prove the identifiability of this model under a few sufficient assumptions. Then we propose a maximum likelihood approach to estimate the interaction functions, which is, to the best of our knowledge, the first exact inference procedure in the frequentist framework.Our method includes a variable selection step in order to recover the support of interactions and therefore to infer the connectivity graph.A benefit of our method is to provide an explicit computation of the log-likelihood, which enables in addition to perform a goodness-of-fit test for assessing the quality of estimations.We compare our method to standard approaches, which were developed in the linear framework and are not specifically designed for handling inhibiting effects.We show that the proposed estimator performs better on synthetic data than alternative approaches. We also illustrate the application of our procedure to a neuronal activity dataset, which highlights the presence of both exciting and inhibiting effects between neurons.

stat.ME

Neuronal Network Inference and Membrane Potential Model using Multivariate Hawkes Processes

In this work, we propose to catch the complexity of the membrane potential's dynamic of a motoneuron between its spikes, taking into account the spikes from other neurons around. Our approach relies on two types of data: extracellular recordings of multiple spikes trains and intracellular recordings of the membrane potential of a central neuron. Our main contribution is to provide a unified framework and a complete pipeline to analyze neuronal activity from data extraction to statistical inference. The first step of the procedure is to select a subnetwork of neurons impacting the central neuron: we use a multivariate Hawkes process to model the spike trains of all neurons and compare two sparse inference procedures to identify the connectivity graph. Then we infer a jump-diffusion dynamic in which jumps are driven from a Hawkes process, the occurrences of which correspond to the spike trains of the aforementioned subset of neurons that interact with the central neuron. We validate the Hawkes model with a goodness-of-fit test and we show that taking into account the information from the connectivity graph improves the inference of the jump-diffusion process. The entire code has been developed and is freely available on GitHub.

math.ST

Maximum Likelihood Estimation for Hawkes Processes with self-excitation or inhibition

In this paper, we present a maximum likelihood method for estimating the parameters of a univariate Hawkes process with self-excitation or inhibition. Our work generalizes techniques and results that were restricted to the self-exciting scenario. The proposed estimator is implemented for the classical exponential kernel and we show that, in the inhibition context, our procedure provides more accurate estimations than current alternative approaches.

math.ST

Uniform Deconvolution for Poisson Point Processes

We focus on the estimation of the intensity of a Poisson process in the presence of a uniform noise. We propose a kernel-based procedure fully calibrated in theory and practice. We show that our adaptive estimator is optimal from the oracle and minimax points of view, and provide new lower bounds when the intensity belongs to a Sobolev ball. By developing the Goldenshluger-Lepski methodology in the case of deconvolution for Poisson processes, we propose an optimal data-driven selection of the kernel bandwidth. Our method is illustrated on the spatial distribution of replication origins and sequence motifs along the human genome.

stat.ME

Heritability estimation of diseases in case-control studies

In the field of genetics, the concept of heritability refers to the proportion of variations of a biological trait or disease that can be explained by genetic factors. Quantifying the heritability of a disease is a fundamental challenge in human genetics, especially when the causes are plural and not clearly identified. Although the literature regarding heritability estimation for binary traits is less rich than for quantitative traits, several methods have been proposed to estimate the heritability of complex diseases. However, to the best of our knowledge, the existing methods are not supported by theoretical grounds. Moreover, most of the methodologies do not take into account a major specificity of the data coming from medical studies, which is the oversampling of the number of patients compared to controls. We propose in this paper to investigate the theoretical properties of the method developed by Golan et al. (2014), which is very efficient in practice, despite the oversampling of patients. Our main result is the proof of the consistency of this estimator. We also provide a numerical study to compare two approximations leading to two heritability estimators.

math.ST

Improving heritability estimation by a variable selection approach in sparse high dimensional linear mixed models

Motivated by applications in neuroanatomy, we propose a novel methodology for estimating the heritability which corresponds to the proportion of phenotypic variance which can be explained by genetic factors. Estimating this quantity for neuroanatomical features is a fundamental challenge in psychiatric disease research. Since the phenotypic variations may only be due to a small fraction of the available genetic information, we propose an estimator of the heritability that can be used in high dimensional sparse linear mixed models. Our method consists of three steps. Firstly, a variable selection stage is performed in order to recover the support of the genetic effects -- also called causal variants -- that is to find the genetic effects which really explain the phenotypic variations. Secondly, we propose a maximum likelihood strategy for estimating the heritability which only takes into account the causal genetic effects found in the first step. Thirdly, we compute the standard error and the 95% confidence interval associated to our heritability estimator thanks to a nonparametric bootsrap approach. Our main contribution consists in providing an estimation of the heritability with standard errors substantially smaller than methods without variable selection when the genetic effects are very sparse. Since the real genetic architecture is in general unknown in practice, we also propose an empirical criterion which allows the user to decide whether it is relevant to apply a variable selection based approach or not. We illustrate the performance of our methodology on synthetic and real neuroanatomic data coming from the Imagen project. We also show that our approach has a very low computational burden and is very efficient from a statistical point of view.

math.ST

Heritability estimation in high dimensional linear mixed models

Motivated by applications in genetic fields, we propose to estimate the heritability in high dimensional sparse linear mixed models. The heritability determines how the variance is shared between the different random components of a linear mixed model. The main novelty of our approach is to consider that the random effects can be sparse, that is may contain null components, but we do not know neither their proportion nor their positions. The estimator that we consider is strongly inspired by the one proposed by Pirinen et al. (2013), and is based on a maximum likelihood approach. We also study the theoretical properties of our estimator, namely we establish that our estimator of the heritability is $\sqrt{n}$-consistent when both the number of observations $n$ and the number of random effects $N$ tend to infinity under mild assumptions. We also prove that our estimator of the heritability satisfies a central limit theorem which gives as a byproduct a confidence interval for the heritability. Some Monte-Carlo experiments are also conducted in order to show the finite sample performances of our estimator.

math.ST