Searcharxiv⌕ Search

arXiv subjects

Romain Azaïs

Publications and source records attributed to Romain Azaïs.

At least 19 recordsLinked to original sources

Asymptotic Analysis and Practical Evaluation of Jump Rate Estimators in Piecewise-Deterministic Markov Processes

Piecewise-deterministic Markov processes (PDMPs) offer a powerful stochastic modeling framework that combines deterministic trajectories with random perturbations at random times. Estimating their local characteristics (particularly the jump rate) is an important yet challenging task. In recent years, non-parametric methods for jump rate inference have been developed, but these approaches often rely on distinct theoretical frameworks, complicating direct comparisons. In this paper, we propose a unified framework to standardize and consolidate state-of-the-art approaches. We establish new results on consistency and asymptotic normality within this framework, enabling rigorous theoretical comparisons of convergence rates and asymptotic variances. Notably, we demonstrate that no single method uniformly outperforms the others, even within the same model. These theoretical insights are validated through numerical simulations using a representative PDMP application: the TCP model. Furthermore, we extend the comparison to real-world data, focusing on cell growth and division dynamics in Escherichia coli. This work enhances the theoretical understanding of PDMP inference while offering practical insights into the relative strengths and limitations of existing methods.

stat.ME↗

Macroscopic Activity-Based Modeling of Urban Active Mobility

This paper develops a macroscopic, activity-based model of urban active mobility using nonintrusive sensor data. It introduces attendance functions to describe spatio-temporal travel patterns between activities and formulates the disaggregation of aggregated counts as a statistical inference problem. Counts are modeled as Poisson variables, and unknown subpopulation sizes are estimated via maximum likelihood, with theoretical guarantees and an efficient EM algorithm for computation. Grounded in a microscopic stochastic model, the framework offers a scalable and privacy-preserving approach to analyzing urban soft mobility dynamics.

stat.ME↗

Maximum likelihood estimation for spinal-structured trees

We investigate some aspects of the problem of the estimation of birth distributions (BD) in multi-type Galton-Watson trees (MGW) with unobserved types. More precisely, we consider two-type MGW called spinal-structured trees. This kind of tree is characterized by a spine of special individuals whose BD $ν$ is different from the other individuals in the tree (called normal whose BD is denoted $μ$). In this work, we show that even in such a very structured two-type population, our ability to distinguish the two types and estimate $μ$ and $ν$ is constrained by a trade-off between the growth-rate of the population and the similarity of $μ$ and $ν$. Indeed, if the growth-rate is too large, large deviations events are likely to be observed in the sampling of the normal individuals preventing us to distinguish them from special ones. Roughly speaking, our approach succeeds if $r<\mathfrak{D}(μ,ν)$ where $r$ is the exponential growth-rate of the population and $\mathfrak{D}$ is a divergence measuring the dissimilarity between $μ$ and $ν$.

math.ST↗

Characterization of random walks on space of unordered trees using efficient metric simulation

The simple random walk on $\mathbb{Z}^p$ shows two drastically different behaviours depending on the value of $p$: it is recurrent when $p\in\{1,2\}$ while it escapes (with a rate increasing with $p$) as soon as $p\geq3$. This classical example illustrates that the asymptotic properties of a random walk provides some information on the structure of its state space. This paper aims to explore analogous questions on space made up of combinatorial objects with no algebraic structure. We take as a model for this problem the space of unordered unlabeled rooted trees endowed with Zhang edit distance. To this end, it defines the canonical unbiased random walk on the space of trees and provides an efficient algorithm to evaluate its escape rate. Compared to Zhang algorithm, it is incremental and computes the edit distance along the random walk approximately 100 times faster on trees of size $500$ on average. The escape rate of the random walk on trees is precisely estimated using intensive numerical simulations, out of reasonable reach without the incremental algorithm.

cs.DS↗

Detection of Common Subtrees with Identical Label Distribution

Frequent pattern mining is a relevant method to analyse structured data, like sequences, trees or graphs. It consists in identifying characteristic substructures of a dataset. This paper deals with a new type of patterns for tree data: common subtrees with identical label distribution. Their detection is far from obvious since the underlying isomorphism problem is graph isomorphism complete. An elaborated search algorithm is developed and analysed from both theoretical and numerical perspectives. Based on this, the enumeration of patterns is performed through a new lossless compression scheme for trees, called DAG-RW, whose complexity is investigated as well. The method shows very good properties, both in terms of computation times and analysis of real datasets from the literature. Compared to other substructures like topological subtrees and labelled subtrees for which the isomorphism problem is linear, the patterns found provide a more parsimonious representation of the data.

cs.DS↗

Enumeration of Irredundant Forests

Reverse search is a convenient method for enumerating structured objects, that can be used both to address theoretical issues and to solve data mining problems. This method has already been successfully developed to handle unordered trees. If the literature proposes solutions to enumerate singletons of trees, we study in this article a more general problem, the enumeration of sets of trees -- forests. Specifically, we mainly study irredundant forests, i.e., where no tree is a subtree of another. By compressing each such forest into a Directed Acyclic Graph (DAG), we develop a reverse search like method to enumerate DAGs compressing irredundant forests. Remarkably, we prove that these DAGs are in bijection with the row-Fishburn matrices, a well-studied class of combinatorial objects. In a second step, we derive our irredundant forest enumeration to provide algorithms for tackling related problems: (i) enumeration of forests in their classical sense (where redundancy is allowed); (ii) the enumeration of "subforests" of a forest, and (iii) the frequent "subforest" mining problem. All the methods presented in this article enumerate each item uniquely, up to isomorphism.

cs.DM↗

Isomorphic unordered labeled trees up to substitution ciphering

Given two messages - as linear sequences of letters, it is immediate to determine whether one can be transformed into the other by simple substitution cipher of the letters. On the other hand, if the letters are carried as labels on nodes of topologically isomorphic unordered trees, determining if a substitution exists is referred to as marked tree isomorphism problem in the literature and has been show to be as hard as graph isomorphism. While the left-to-right direction provides the cipher of letters in the case of linear messages, if the messages are carried by unordered trees, the cipher is given by a tree isomorphism. The number of isomorphisms between two trees is roughly exponential in the size of the trees, which makes the problem of finding a cipher difficult by exhaustive search. This paper presents a method that aims to break the combinatorics of the isomorphisms search space. We show that in a linear time (in the size of the trees), we reduce the cardinality of this space by an exponential factor on average.

cs.DM↗

Bipartite non-locality beyond Bell inequalities by means of multivariable correlations

We give a set of necessary conditions for locality in bipartite systems, which include and generalize known Bell's inequalities. Each condition corresponds to a specific order of the expansion of random variables defined on graphs, in terms of genuinely n-variable correlation functions. The first non-trivial order leads to known Bell inequalities, while higher orders produce additional, non-equivalent conditions. In particular, in CHSH settings, we obtain at least two additional tight conditions which do not reduce to the CHSH inequality. This shows that tight Bell inequalities are sufficient but not necessary conditions for the non-locality of bipartite quantum correlations.

quant-ph↗

The Weight Function in the Subtree Kernel is Decisive

Tree data are ubiquitous because they model a large variety of situations, e.g., the architecture of plants, the secondary structure of RNA, or the hierarchy of XML files. Nevertheless, the analysis of these non-Euclidean data is difficult per se. In this paper, we focus on the subtree kernel that is a convolution kernel for tree data introduced by Vishwanathan and Smola in the early 2000's. More precisely, we investigate the influence of the weight function from a theoretical perspective and in real data applications. We establish on a 2-classes stochastic model that the performance of the subtree kernel is improved when the weight of leaves vanishes, which motivates the definition of a new weight function, learned from the data and not fixed by the user as usually done. To this end, we define a unified framework for computing the subtree kernel from ordered or unordered trees, that is particularly suitable for tuning parameters. We show through eight real data classification problems the great efficiency of our approach, in particular for small datasets, which also states the high importance of the weight function. Finally, a visualization tool of the significant features is derived.

stat.ML↗

Nearest Embedded and Embedding Self-Nested Trees

Self-nested trees present a systematic form of redundancy in their subtrees and thus achieve optimal compression rates by DAG compression. A method for quantifying the degree of self-similarity of plants through self-nested trees has been introduced by Godin and Ferraro in 2010. The procedure consists in computing a self-nested approximation, called the nearest embedding self-nested tree, that both embeds the plant and is the closest to it. In this paper, we propose a new algorithm that computes the nearest embedding self-nested tree with a smaller overall complexity, but also the nearest embedded self-nested tree. We show from simulations that the latter is mostly the closest to the initial data, which suggests that this better approximation should be used as a privileged measure of the degree of self-similarity of plants.

cs.DS↗

Inference for conditioned Galton-Watson trees from their Harris path

Tree-structured data naturally appear in various fields, particularly in biology where plants and blood vessels may be described by trees, but also in computer science because XML documents form a tree structure. This paper is devoted to the estimation of the relative scale parameter of conditioned Galton-Watson trees. New estimators are introduced and their consistency is stated. A comparison is made with an existing approach of the literature. A simulation study shows the good behavior of our procedure on finite-sample sizes and from missing or noisy data. An application to the analysis of revisions of Wikipedia articles is also considered through real data.

math.ST↗

Approximation of trees by self-nested trees

The class of self-nested trees presents remarkable compression properties because of the systematic repetition of subtrees in their structure. In this paper, we provide a better combinatorial characterization of this specific family of trees. In particular, we show from both theoretical and practical viewpoints that complex queries can be quickly answered in self-nested trees compared to general trees. We also present an approximation algorithm of a tree by a self-nested one that can be used in fast prediction of edit distance between two trees.

cs.DS↗

Estimation of the average number of continuous crossings for non-stationary non-diffusion processes

Assume that you observe trajectories of a non-diffusive non-stationary process and that you are interested in the average number of times where the process crosses some threshold (in dimension $d=1$) or hypersurface (in dimension $d\geq2$). Of course, you can actually estimate this quantity by its empirical version counting the number of observed crossings. But is there a better way? In this paper, for a wide class of piecewise smooth processes, we propose estimators of the average number of continuous crossings of an hypersurface based on Kac-Rice formulae. We revisit these formulae in the uni- and multivariate framework in order to be able to handle non-stationary processes. Our statistical method is tested on both simulated and real data.

stat.ME↗

Integral estimation based on Markovian design

Suppose that a mobile sensor describes a Markovian trajectory in the ambient space. At each time the sensor measures an attribute of interest, e.g., the temperature. Using only the location history of the sensor and the associated measurements, the aim is to estimate the average value of the attribute over the space. In contrast to classical probabilistic integration methods, e.g., Monte Carlo, the proposed approach does not require any knowledge on the distribution of the sensor trajectory. Probabilistic bounds on the convergence rates of the estimator are established. These rates are better than the traditional "root n"-rate, where n is the sample size, attached to other probabilistic integration methods. For finite sample sizes, the good behaviour of the procedure is demonstrated through simulations and an application to the evaluation of the average temperature of oceans is considered.

math.ST↗

A new characterization of the jump rate for piecewise-deterministic Markov processes with discrete transitions

Piecewise-deterministic Markov processes form a general class of non-diffusion stochastic models that involve both deterministic trajectories and random jumps at random times. In this paper, we state a new characterization of the jump rate of such a process with discrete transitions. We deduce from this result a nonparametric technique for estimating this feature of interest. We state the uniform convergence in probability of the estimator. The methodology is illustrated on a numerical example.

stat.ME↗

Optimal choice among a class of nonparametric estimators of the jump rate for piecewise-deterministic Markov processes

A piecewise-deterministic Markov process is a stochastic process whose behavior is governed by an ordinary differential equation punctuated by random jumps occurring at random times. We focus on the nonparametric estimation problem of the jump rate for such a stochastic model observed within a long time interval under an ergodicity condition. We introduce an uncountable class (indexed by the deterministic flow) of recursive kernel estimates of the jump rate and we establish their strong pointwise consistency as well as their asymptotic normality. We propose to choose among this class the estimator with the minimal variance, which is unfortunately unknown and thus remains to be estimated. We also discuss the choice of the bandwidth parameters by cross-validation methods.

math.ST↗

Piecewise deterministic Markov process - recent results

We give a short overview of recent results on a specific class of Markov process: the Piecewise Deterministic Markov Processes (PDMPs). We first recall the definition of these processes and give some general results. On more specific cases such as the TCP model or a model of switched vector fields, better results can be proved, especially as regards long time behaviour. We continue our review with an infinite dimensional example of neuronal activity. From the statistical point of view, these models provide specific challenges: we illustrate this point with the example of the estimation of the distribution of the inter-jumping times. We conclude with a short overview on numerical methods used for simulating PDMPs.

math.ST↗

A recursive nonparametric estimator for the transition kernel of a piecewise-deterministic Markov process

In this paper, we investigate a nonparametric approach to provide a recursive estimator of the transition density of a non-stationary piecewise-deterministic Markov process, from only one observation of the path within a long time. In this framework, we do not observe a Markov chain with transition kernel of interest. Fortunately, one may write the transition density of interest as the ratio of the invariant distributions of two embedded chains of the process. Our method consists in estimating these invariant measures. We state a result of consistency and a central limit theorem under some general assumptions about the main features of the process. A simulation study illustrates the well asymptotic behavior of our estimator.

math.ST↗