Searcharxiv⌕ Search

arXiv subjects

Jean-Jil Duchamps

Publications and source records attributed to Jean-Jil Duchamps.

16 recordsLinked to original sources

Sharp $L \log L$ condition for supercritical Galton-Watson processes with countable types

We investigate Kesten-Stigum-like results for multi-type Galton-Watson processes with a countable number of types in a general setting, allowing us in particular to consider processes with an infinite total population at each generation. Specifically, a sharp $L\log L$ condition is found under the only assumption that the mean reproduction matrix is positive recurrent in the sense of Vere-Jones (1967). The type distribution is shown to always converge in probability in the recurrent case, and under conditions covering many cases it is shown to converge almost surely. We finally apply these results to the study of a directed network built from countably many configuration models as a model of epidemics.

math.PR↗

A functional central limit theorem for kernel gradient flow and infinitesimal gradient boosting

Building on the large-sample analysis of infinitesimal gradient boosting (Dombry and Duchamps, 2024b), we study the fluctuations of the process around its deterministic limit and establish a functional central limit theorem: the rescaled deviations converge in distribution to a Gaussian process. The analysis is carried out in a reproducing kernel Hilbert space (RKHS) naturally associated with the softmax gradient tree base learner, in which the boosting process is characterized as the solution of an autonomous ordinary differential equation (ODE). The proof rests on a general stochastic perturbation analysis of ODEs in Banach spaces, which is of independent interest: whenever a sequence of vector fields converges and satisfies a central limit theorem, so does the associated ODE solution. We first illustrate this perturbation approach in the simpler setting of kernel gradient flow, where the Gaussian limit admits an explicit characterization, and then consider the more complicated tree-based gradient boosting setting.

math.PR↗

A parameterized family of balance indices for phylogenetic networks

We introduce a new family of balance indices for phylogenetic networks: the $H_α$ indices, where $α$ is a positive real number. This family includes the $B_2$ index as a special case ($α= 1$) and provides a natural extension of the Sackin index to phylogenetic networks. We show that the $H_α$ indices share many structural properties with the $B_2$ index, most notably a "grafting property" that makes it possible to express the $H_α$ index of a network in terms of the $H_α$ indices of its biconnected components. These properties allow us to identify networks that minimize / maximize $H_α$ for various classes of phylogenetic networks, and to study its distribution for several models of random trees and networks (in particular, Galton-Watson trees and binary Markov branching trees, with a focus on the Yule and PDA models). Finally, we show how local limits can be used to analyze the asymptotic behavior of $H_α$ for large trees and networks, and we obtain general results for the moments of $H_α$ for a broad class of random phylogenetic networks known as blowups of Galton-Watson trees.

math.CO↗

The $B_2$ index of galled trees

In recent years, there has been an effort to extend the classical notion of phylogenetic balance, originally defined in the context of trees, to networks. One of the most natural ways to do this is with the so-called $B_2$ index. In this paper, we study the $B_2$ index for a prominent class of phylogenetic networks: galled trees. We show that the $B_2$ index of a uniform leaf-labeled galled tree converges in distribution as the network becomes large. We characterize the corresponding limiting distribution, and provide a way to compute its moments. This is the first time that a balance index has been studied to this level of detail for a random phylogenetic network. One specificity of this work is that we use two different and independent approaches, each with its advantages: analytic combinatorics, and local limits. The analytic combinatorics approach is more direct, as it relies on standard tools; but it involves slightly more complex calculations. Because it has not previously been used to study such questions, the local limit approach requires developing an extensive framework beforehand; however, this framework is interesting in itself and can be used to tackle other similar problems.

q-bio.PE↗

An RKHS Perspective on Tree Ensembles

Random Forests and Gradient Boosting are among the most effective algorithms for supervised learning on tabular data. Both belong to the class of tree-based ensemble methods, where predictions are obtained by aggregating many randomized regression trees. In this paper, we develop a theoretical framework for analyzing such methods through Reproducing Kernel Hilbert Spaces (RKHSs) constructed on tree ensembles -- more precisely, on the random partitions generated by randomized regression trees. We establish fundamental analytical properties of the resulting Random Forest kernel, including boundedness, continuity, and universality, and show that a Random Forest predictor can be characterized as the unique minimizer of a penalized empirical risk functional in this RKHS, providing a variational interpretation of ensemble learning. We further extend this perspective to the continuous-time formulation of Gradient Boosting introduced by Dombry and Duchamps, and demonstrate that it corresponds to a gradient flow on a Hilbert manifold induced by the Random Forest RKHS. A key feature of this framework is that both the kernel and the RKHS geometry are data-dependent, offering a theoretical explanation for the strong empirical performance of tree-based ensembles. Finally, we illustrate the practical potential of this approach by introducing a kernel principal component analysis built on the Random Forest kernel, which enhances the interpretability of ensemble models, as well as GVI, a new geometric variable importance criterion.

stat.ML↗

A branching process with coalescence to model random phylogenetic networks

We introduce a biologically natural, mathematically tractable model of random phylogenetic network to describe evolution in the presence of hybridization. One of the features of this model is that the hybridization rate of the lineages correlates negatively with their phylogenetic distance. We give formulas / characterizations for quantities of biological interest that make them straightforward to compute in practice. We show that the appropriately rescaled network, seen as a metric space, converges to the Brownian continuum random tree, and that the uniformly rooted network has a local weak limit, which we describe explicitly.

math.PR↗

A large sample theory for infinitesimal gradient boosting

Infinitesimal gradient boosting (Dombry and Duchamps, 2021) is defined as the vanishing-learning-rate limit of the popular tree-based gradient boosting algorithm from machine learning. It is characterized as the solution of a nonlinear ordinary differential equation in a infinite-dimensional function space where the infinitesimal boosting operator driving the dynamics depends on the training sample. We consider the asymptotic behavior of the model in the large sample limit and prove its convergence to a deterministic process. This population limit is again characterized by a differential equation that depends on the population distribution. We explore some properties of this population limit: we prove that the dynamics makes the test error decrease and we consider its long time behavior.

stat.ML↗

General epidemiological models: Law of large numbers and contact tracing

We study a class of individual-based, fixed-population size epidemic models under general assumptions, e.g., heterogeneous contact rates encapsulating changes in behavior and/or enforcement of control measures. We show that the large-population dynamics are deterministic and relate to the Kermack-McKendrick PDE. Our assumptions are minimalistic in the sense that the only important requirement is that the basic reproduction number of the epidemic $R_0$ be finite, and allow us to tackle both Markovian and non-Markovian dynamics. The novelty of our approach is to study the "infection graph" of the population. We show local convergence of this random graph to a Poisson (Galton-Watson) marked tree, recovering Markovian backward-in-time dynamics in the limit as we trace back the transmission chain leading to a focal infection. This effectively models the process of contact tracing in a large population. It is expressed in terms of the Doob $h$-transform of a certain renewal process encoding the time of infection along the chain. Our results provide a mathematical formulation relating a fundamental epidemiological quantity, the generation time distribution, to the successive time of infections along this transmission chain.

math.PR↗

Infinitesimal gradient boosting

We define infinitesimal gradient boosting as a limit of the popular tree-based gradient boosting algorithm from machine learning. The limit is considered in the vanishing-learning-rate asymptotic, that is when the learning rate tends to zero and the number of gradient trees is rescaled accordingly. For this purpose, we introduce a new class of randomized regression trees bridging totally randomized trees and Extra Trees and using a softmax distribution for binary splitting. Our main result is the convergence of the associated stochastic algorithm and the characterization of the limiting procedure as the unique solution of a nonlinear ordinary differential equation in a infinite dimensional function space. Infinitesimal gradient boosting defines a smooth path in the space of continuous functions along which the training error decreases, the residuals remain centered and the total variation is well controlled.

stat.ML↗

From individual-based epidemic models to McKendrick-von Foerster PDEs: A guide to modeling and inferring COVID-19 dynamics

We present a unifying, tractable approach for studying the spread of viruses causing complex diseases requiring to be modeled using a large number of types (e.g., infective stage, clinical state, risk factor class). We show that recording each infected individual's infection age, i.e., the time elapsed since infection, has three benefits. First, regardless of the number of types, the age distribution of the population can be described by means of a first-order, one-dimensional partial differential equation (PDE) known as the McKendrick-von Foerster equation. The frequency of type $i$ is simply obtained by integrating the probability of being in state $i$ at a given age against the age distribution. This representation induces a simple methodology based on the additional assumption of Poisson sampling to infer and forecast the epidemic. We illustrate this technique using French data from the COVID-19 epidemic. Second, our approach generalizes and simplifies standard compartmental models using high-dimensional systems of ordinary differential equations (ODEs) to account for disease complexity. We show that such models can always be rewritten in our framework, thus, providing a low-dimensional yet equivalent representation of these complex models. Third, beyond the simplicity of the approach, we show that our population model naturally appears as a universal scaling limit of a large class of fully stochastic individual-based epidemic models, where the initial condition of the PDE emerges as the limiting age structure of an exponentially growing population starting from a single individual.

q-bio.PE↗

The Moran forest

Starting from any graph on $\{1, \ldots, n\}$, consider the Markov chain where at each time-step a uniformly chosen vertex is disconnected from all of its neighbors and reconnected to another uniformly chosen vertex. This Markov chain has a stationary distribution whose support is the set of non-empty forests on $\{1, \ldots, n\}$. The random forest corresponding to this stationary distribution has interesting connections with the uniform rooted labeled tree and the uniform attachment tree. We fully characterize its degree distribution, the distribution of its number of trees, and the limit distribution of the size of a tree sampled uniformly. We also show that the size of the largest tree is asymptotically $α\log n$, where $α= (1 - \log(e - 1))^{-1} \approx 2.18$, and that the degree of the most connected vertex is asymptotically $\log n / \log\log n$.

math.PR↗

Fragmentations with self-similar branching speeds

We consider fragmentation processes with values in the space of marked partitions of $\mathbb{N}$, i.e. partitions where each block is decorated with a nonnegative real number. Assuming that the marks on distinct blocks evolve as independent positive self-similar Markov processes and determine the speed at which their blocks fragment, we get a natural generalization of the self-similar fragmentations of Bertoin (2002). Our main result is the characterization of these generalized fragmentation processes: a Lévy-Khinchin representation is obtained, using techniques from positive self-similar Markov processes and from classical fragmentation processes. We then give sufficient conditions for their absorption in finite time to a frozen state, and for the genealogical tree of the process to have finite total length.

math.PR↗

Renewal sequences and record chains related to multiple zeta sums

For the random interval partition of $[0,1]$ generated by the uniform stick-breaking scheme known as GEM$(1)$, let $u_k$ be the probability that the first $k$ intervals created by the stick-breaking scheme are also the first $k$ intervals to be discovered in a process of uniform random sampling of points from $[0,1]$. Then $u_k$ is a renewal sequence. We prove that $u_k$ is a rational linear combination of the real numbers $1, ζ(2), \ldots, ζ(k)$ where $ζ$ is the Riemann zeta function, and show that $u_k$ has limit $1/3$ as $k \to \infty$. Related results provide probabilistic interpretations of some multiple zeta values in terms of a Markov chain derived from the interval partition. This Markov chain has the structure of a weak record chain. Similar results are given for the GEM$(θ)$ model, with beta$(1,θ)$ instead of uniform stick-breaking factors, and for another more algebraic derivation of renewal sequences from the Riemann zeta function.

math.PR↗

Trees within trees: Simple nested coalescents

We consider the compact space of pairs of nested partitions of $\mathbb N$, where by analogy with models used in molecular evolution, we call "gene partition" the finer partition and "species partition" the coarser one. We introduce the class of nondecreasing processes valued in nested partitions, assumed Markovian and with exchangeable semigroup. These processes are said simple when each partition only undergoes one coalescence event at a time (but possibly the same time). Simple nested exchangeable coalescent (SNEC) processes can be seen as the extension of $Λ$-coalescents to nested partitions. We characterize the law of SNEC processes as follows. In the absence of gene coalescences, species blocks undergo $Λ$-coalescent type events and in the absence of species coalescences, gene blocks lying in the same species block undergo i.i.d. $Λ$-coalescents. Simultaneous coalescence of the gene and species partitions are governed by an intensity measure $ν_s$ on $(0,1]\times {\mathcal M}_1 ([0,1])$ providing the frequency of species merging and the law in which are drawn (independently) the frequencies of genes merging in each coalescing species block. As an application, we also study the conditions under which a SNEC process comes down from infinity.

math.PR↗

Trees within trees II: Nested Fragmentations

Similarly as in (Blancas et al. 2018) where nested coalescent processes are studied, we generalize the definition of partition-valued homogeneous Markov fragmentation processes to the setting of nested partitions, i.e. pairs of partitions $(ζ,ξ)$ where $ζ$ is finer than $ξ$. As in the classical univariate setting, under exchangeability and branching assumptions, we characterize the jump measure of nested fragmentation processes, in terms of erosion coefficients and dislocation measures. Among the possible jumps of a nested fragmentation, three forms of erosion and two forms of dislocation are identified - one of which being specific to the nested setting and relating to a bivariate paintbox process.

math.PR↗

Mutations on a Random Binary Tree with Measured Boundary

Consider a random real tree whose leaf set, or boundary, is endowed with a finite mass measure. Each element of the tree is further given a type, or allele, inherited from the most recent atom of a random point measure (infinitely-many-allele model) on the skeleton of the tree. The partition of the boundary into distinct alleles is the so-called allelic partition. In this paper, we are interested in the infinite trees generated by supercritical, possibly time-inhomogeneous, binary branching processes, and in their boundary, which is the set of particles `co-existing at infinity'. We prove that any such tree can be mapped to a random, compact ultrametric tree called coalescent point process, endowed with a `uniform' measure on its boundary which is the limit as $t\to\infty$ of the properly rescaled counting measure of the population at time $t$. We prove that the clonal (i.e., carrying the same allele as the root) part of the boundary is a regenerative set that we characterize. We then study the allelic partition of the boundary through the measures of its blocks. We also study the dynamics of the clonal subtree, which is a Markovian increasing tree process as mutations are removed.

math.PR↗