SearcharxivSearch

arXiv subjects

Clement Dombry

Publications and source records attributed to Clement Dombry.

8 recordsLinked to original sources

An RKHS Perspective on Tree Ensembles

Random Forests and Gradient Boosting are among the most effective algorithms for supervised learning on tabular data. Both belong to the class of tree-based ensemble methods, where predictions are obtained by aggregating many randomized regression trees. In this paper, we develop a theoretical framework for analyzing such methods through Reproducing Kernel Hilbert Spaces (RKHSs) constructed on tree ensembles -- more precisely, on the random partitions generated by randomized regression trees. We establish fundamental analytical properties of the resulting Random Forest kernel, including boundedness, continuity, and universality, and show that a Random Forest predictor can be characterized as the unique minimizer of a penalized empirical risk functional in this RKHS, providing a variational interpretation of ensemble learning. We further extend this perspective to the continuous-time formulation of Gradient Boosting introduced by Dombry and Duchamps, and demonstrate that it corresponds to a gradient flow on a Hilbert manifold induced by the Random Forest RKHS. A key feature of this framework is that both the kernel and the RKHS geometry are data-dependent, offering a theoretical explanation for the strong empirical performance of tree-based ensembles. Finally, we illustrate the practical potential of this approach by introducing a kernel principal component analysis built on the Random Forest kernel, which enhances the interpretability of ensemble models, as well as GVI, a new geometric variable importance criterion.

stat.ML

Pareto processes for threshold exceedances in spatial extremes

We review some recent development in the theory of spatial extremes related to Pareto Processes and modeling of threshold exceedances. We provide theoretical background, methodology for modeling, simulation and inference as well as an illustration to wave height modelling. This preprint is an author version of a chapter to appear in a collaborative book.

math.ST

A large sample theory for infinitesimal gradient boosting

Infinitesimal gradient boosting (Dombry and Duchamps, 2021) is defined as the vanishing-learning-rate limit of the popular tree-based gradient boosting algorithm from machine learning. It is characterized as the solution of a nonlinear ordinary differential equation in a infinite-dimensional function space where the infinitesimal boosting operator driving the dynamics depends on the training sample. We consider the asymptotic behavior of the model in the large sample limit and prove its convergence to a deterministic process. This population limit is again characterized by a differential equation that depends on the population distribution. We explore some properties of this population limit: we prove that the dynamics makes the test error decrease and we consider its long time behavior.

stat.ML

Simple models for multivariate regular variations and the Hüsler-Reiss Pareto distribution

We revisit multivariate extreme value theory modeling by emphasizing multivariate regular variations and the multivariate Breiman Lemma. This allows us to recover in a simple framework the most popular multivariate extreme value distributions, such as the logistic, negative logistic, Dirichlet, extremal-$t$ and Hüsler-Reiss models. In a second part of the paper, we focus on the Hüsler-Reiss Pareto model and its surprising exponential family property. After a thorough study of this exponential family structure, we focus on maximum likelihood estimation. We also consider the generalized Hüsler-Reiss Pareto model with different tail indices and a likelihood ratio test for discriminating constant tail index versus varying tail indices.

stat.ME

Bayesian inference for multivariate extreme value distributions

Statistical modeling of multivariate and spatial extreme events has attracted broad attention in various areas of science. Max-stable distributions and processes are the natural class of models for this purpose, and many parametric families have been developed and successfully applied. Due to complicated likelihoods, the efficient statistical inference is still an active area of research, and usually composite likelihood methods based on bivariate densities only are used. Thibaud et al. (2016, Ann. Appl. Stat., to appear) use a Bayesian approach to fit a Brown--Resnick process to extreme temperatures. In this paper, we extend this idea to a methodology that is applicable to general max-stable distributions and that uses full likelihoods. We further provide simple conditions for the asymptotic normality of the median of the posterior distribution and verify them for the commonly used models in multivariate and spatial extreme value statistics. A simulation study shows that this point estimator is considerably more efficient than the composite likelihood estimator in a frequentist framework. From a Bayesian perspective, our approach opens the way for new techniques such as Bayesian model comparison in multivariate and spatial extremes.

stat.ME

Asymptotic properties of the maximum likelihood estimator for multivariate extreme value distributions

Max-stable distributions and processes are important models for extreme events and the assessment of tail risks. The full, multivariate likelihood of a parametric max-stable distribution is complicated and only recent advances enable its use. The asymptotic properties of the maximum likelihood estimator in multivariate extremes are mostly unknown. In this paper we provide natural conditions on the exponent function and the angular measure of the max-stable distribution that ensure asymptotic normality of the estimator. We show the effectiveness of this result by applying it to popular parametric models in multivariate extreme value statistics and to the most commonly used families of spatial max-stable processes.

math.ST

Functional macroscopic behavior of weighted random ball model

We consider a generalization of the weighted random ball model. The model is driven by a random Poisson measure with a product heavy tailed intensity measure. Such a model typically represents the transmission of a network of stations with a fading effect. In a previous article, the authors proved the convergence of the finite-dimensional distributions of related generalized random fields under various scalings and in the particular case when the fading function is the indicator function of the unit ball. In this paper, tightness and functional convergence are investigated. Using suitable moment estimates, we prove functional convergences for some parametric classes of configurations under the so-called large ball scaling and intermediate ball scaling. Convergence in the space of distributions is also discussed.

math.PR

The Curie-Weiss model with dynamical external field

We study a Curie-Weiss model with a random external field generated by a dynamical system. Probabilistic limit theorems (weak law of large numbers, central limit theorems) are proven for the corresponding magnetization.

math.PR