Searcharxiv⌕ Search

arXiv subjects

Clément Dombry

Publications and source records attributed to Clément Dombry.

At least 19 recordsLinked to original sources

A functional central limit theorem for kernel gradient flow and infinitesimal gradient boosting

Building on the large-sample analysis of infinitesimal gradient boosting (Dombry and Duchamps, 2024b), we study the fluctuations of the process around its deterministic limit and establish a functional central limit theorem: the rescaled deviations converge in distribution to a Gaussian process. The analysis is carried out in a reproducing kernel Hilbert space (RKHS) naturally associated with the softmax gradient tree base learner, in which the boosting process is characterized as the solution of an autonomous ordinary differential equation (ODE). The proof rests on a general stochastic perturbation analysis of ODEs in Banach spaces, which is of independent interest: whenever a sequence of vector fields converges and satisfies a central limit theorem, so does the associated ODE solution. We first illustrate this perturbation approach in the simpler setting of kernel gradient flow, where the Gaussian limit admits an explicit characterization, and then consider the more complicated tree-based gradient boosting setting.

math.PR↗

Out-of-Distribution generalization of quantile regression with heavy tailed inputs: an SVM approach

We study quantile regression in an extrapolation regime where the covariate takes unusually large values. Under regular variation assumptions, extreme observations can be effectively characterized through their angular components, enabling learning strategies that focus on the angle of the most extreme observations. This approach is formalized through the minimization of an asymptotic conditional risk that localizes learning in the tail of the covariate distribution. We propose a novel Support Vector Machine (SVM) framework for extreme quantile regression, leveraging reproducing kernel Hilbert spaces to handle high-dimensional and nonlinear settings. Our method also accommodates unbounded response variables and avoids restrictive transformations. We establish finite-sample learning guarantees under mild regularity assumptions. The proposed framework unifies ideas from statistical learning and multivariate extremes, providing a tractable and theoretically grounded approach to extrapolation. We complement our theoretical findings with an empirical study on river flow data from the Danube, demonstrating the practical relevance of our methods.

stat.ML↗

Infinite random forests for imbalanced classification tasks

We study predictive probability inference in classification tasks using random forests under class imbalance. We focus on two simplified variants of Breiman's algorithm, namely subsampling Infinite Random Forests (IRFs) and under-sampling IRFs, and establish their asymptotic normality. In the under-sampling setting, training data from both classes are resampled to achieve balance, which enhances minority class representation but introduces a biased model. To correct this, we propose a debiasing procedure based on Importance Sampling (IS) using odds ratios. We instantiate our results using 1-Nearest Neighbor (1-NN) classifiers as base learners in the IRFs and prove the nearly minimax optimality of the approach for Lipschitz continuous objectives. We also show that the IS bagged 1-NN estimator matches the convergence rate of its subsampled counterpart while attaining lower asymptotic variance in most cases. Our theoretical findings are supported by simulation studies, highlighting the empirical benefits of the proposed approach.

math.ST↗

Asymptotic theory for Bayesian inference and prediction: from the ordinary to a conditional Peaks-Over-Threshold method

The Peaks Over Threshold (POT) method is the most popular statistical method for the analysis of univariate extremes. Even though there is a rich applied literature on Bayesian inference for the POT, the asymptotic theory for such proposals is missing. Even more importantly, the ambitious and challenging problem of predicting future extreme events according to a proper predictive statistical approach has received no attention to date. In this paper we fill this gap by developing the asymptotic theory of posterior distributions (consistency, contraction rates, asymptotic normality and asymptotic coverage of credible intervals) and prediction within the Bayesian framework in the POT context. We extend this asymptotic theory to account for cases where the focus is on the tail properties of the conditional distribution of a response variable given a vector of random covariates. To enable accurate predictions of extreme events more severe than those previously observed, we derive the posterior predictive distribution as an estimator of the conditional distribution of an out-of-sample random variable, given that it exceeds a sufficiently high threshold. We establish Wasserstein consistency of the posterior predictive distribution under both the unconditional and covariate-conditional approaches and derive its contraction rates. Simulations show the good performances of the proposed Bayesian inferential methods. The analysis of the change in the frequency of financial crises over time shows the utility of our methodology.

math.ST↗

Distributional regression with reject option

Selective prediction, where a model has the option to abstain from making a decision, is crucial for machine learning applications in which mistakes are costly. In this work, we focus on distributional regression and introduce a framework that enables the model to abstain from estimation in situations of high uncertainty. We refer to this approach as distributional regression with reject option, inspired by similar concepts in classification and regression with reject option. We study the scenario where the rejection rate is fixed. We derive a closed-form expression for the optimal rule, which relies on thresholding the entropy function of the Continuous Ranked Probability Score (CRPS). We propose a semi-supervised estimation procedure for the optimal rule, using two datasets: the first, labeled, is used to estimate both the conditional distribution function and the entropy function of the CRPS, while the second, unlabeled, is employed to calibrate the desired rejection rate. Notably, the control of the rejection rate is distribution-free. Under mild conditions, we show that our procedure is asymptotically as effective as the optimal rule, both in terms of error rate and rejection rate. Additionally, we establish rates of convergence for our approach based on distributional k-nearest neighbor. A numerical analysis on real-world datasets demonstrates the strong performance of our procedure

math.ST↗

Distributional regression: CRPS-error bounds for model fitting, model selection and convex aggregation

Distributional regression aims at estimating the conditional distribution of a targetvariable given explanatory co-variates. It is a crucial tool for forecasting whena precise uncertainty quantification is required. A popular methodology consistsin fitting a parametric model via empirical risk minimization where the risk ismeasured by the Continuous Rank Probability Score (CRPS). For independentand identically distributed observations, we provide a concentration result for theestimation error and an upper bound for its expectation. Furthermore, we considermodel selection performed by minimization of the validation error and provide aconcentration bound for the regret. A similar result is proved for convex aggregationof models. Finally, we show that our results may be applied to various models suchas Ensemble Model Output Statistics (EMOS), distributional regression networks,distributional nearest neighbors or distributional random forests and we illustrateour findings on two data sets (QSAR aquatic toxicity and Airfoil self-noise).

math.ST↗

Distributional Regression U-Nets for the Postprocessing of Precipitation Ensemble Forecasts

Accurate precipitation forecasts have a high socio-economic value due to their role in decision-making in various fields such as transport networks and farming. We propose a global statistical postprocessing method for grid-based precipitation ensemble forecasts. This U-Net-based distributional regression method predicts marginal distributions in the form of parametric distributions inferred by scoring rule minimization. Distributional regression U-Nets are compared to state-of-the-art postprocessing methods for daily 21-h forecasts of 3-h accumulated precipitation over the South of France. Training data comes from the Météo-France weather model AROME-EPS and spans 3 years. A practical challenge appears when consistent data or reforecasts are not available. Distributional regression U-Nets compete favorably with the raw ensemble. In terms of continuous ranked probability score, they reach a performance comparable to quantile regression forests (QRF). However, they are unable to provide calibrated forecasts in areas associated with high climatological precipitation. In terms of predictive power for heavy precipitation events, they outperform both QRF and semi-parametric QRF with tail extensions.

stat.ML↗

Proper Scoring Rules for Multivariate Probabilistic Forecasts based on Aggregation and Transformation

Proper scoring rules are an essential tool to assess the predictive performance of probabilistic forecasts. However, propriety alone does not ensure an informative characterization of predictive performance and it is recommended to compare forecasts using multiple scoring rules. With that in mind, interpretable scoring rules providing complementary information are necessary. We formalize a framework based on aggregation and transformation to build interpretable multivariate proper scoring rules. Aggregation-and-transformation-based scoring rules are able to target specific features of the probabilistic forecasts; which improves the characterization of the predictive performance. This framework is illustrated through examples taken from the literature and studied using numerical experiments showcasing its benefits. In particular, it is shown that it can help bridge the gap between proper scoring rules and spatial verification tools.

stat.ME↗

A new methodology to predict the oncotype scores based on clinico-pathological data with similar tumor profiles

Introduction: The Oncotype DX (ODX) test is a commercially available molecular test for breast cancer assay that provides prognostic and predictive breast cancer recurrence information for hormone positive, HER2-negative patients. The aim of this study is to propose a novel methodology to assist physicians in their decision-making. Methods: A retrospective study between 2012 and 2020 with 333 cases that underwent an ODX assay from three hospitals in Bourgogne Franche-Comt{é} was conducted. Clinical and pathological reports were used to collect the data. A methodology based on distributional random forest was developed using 9 clinico-pathological characteristics. This methodology can be used particularly to identify the patients of the training cohort that share similarities with the new patient and to predict an estimate of the distribution of the ODX score. Results: The mean age of participants id 56.9 years old. We have correctly classified 92% of patients in low risk and 40.2% of patients in high risk. The overall accuracy is 79.3%. The proportion of low risk correct predicted value (PPV) is 82%. The percentage of high risk correct predicted value (NPV) is approximately 62.3%. The F1-score and the Area Under Curve (AUC) are of 0.87 and 0.759, respectively. Conclusion: The proposed methodology makes it possible to predict the distribution of the ODX score for a patient and provides an explanation of the predicted score. The use of the methodology with the pathologist's expertise on the different histological and immunohistochemical characteristics has a clinical impact to help oncologist in decision-making regarding breast cancer therapy.

stat.AP↗

Stone's theorem for distributional regression in Wasserstein distance

We extend the celebrated Stone's theorem to the framework of distributional regression. More precisely, we prove that weighted empirical distribution with local probability weights satisfying the conditions of Stone's theorem provide universally consistent estimates of the conditional distributions, where the error is measured by the Wasserstein distance of order p $\ge$ 1. Furthermore, for p = 1, we determine the minimax rates of convergence on specific classes of distributions. We finally provide some applications of these results, including the estimation of conditional tail expectation or probability weighted moment.

math.ST↗

Infinitesimal gradient boosting

We define infinitesimal gradient boosting as a limit of the popular tree-based gradient boosting algorithm from machine learning. The limit is considered in the vanishing-learning-rate asymptotic, that is when the learning rate tends to zero and the number of gradient trees is rescaled accordingly. For this purpose, we introduce a new class of randomized regression trees bridging totally randomized trees and Extra Trees and using a softmax distribution for binary splitting. Our main result is the convergence of the associated stochastic algorithm and the characterization of the limiting procedure as the unique solution of a nonlinear ordinary differential equation in a infinite dimensional function space. Infinitesimal gradient boosting defines a smooth path in the space of continuous functions along which the training error decreases, the residuals remain centered and the total variation is well controlled.

stat.ML↗

Gradient boosting for extreme quantile regression

Extreme quantile regression provides estimates of conditional quantiles outside the range of the data. Classical quantile regression performs poorly in such cases since data in the tail region are too scarce. Extreme value theory is used for extrapolation beyond the range of observed values and estimation of conditional extreme quantiles. Based on the peaks-over-threshold approach, the conditional distribution above a high threshold is approximated by a generalized Pareto distribution with covariate dependent parameters. We propose a gradient boosting procedure to estimate a conditional generalized Pareto distribution by minimizing its deviance. Cross-validation is used for the choice of tuning parameters such as the number of trees and the tree depths. We discuss diagnostic plots such as variable importance and partial dependence plots, which help to interpret the fitted models. In simulation studies we show that our gradient boosting procedure outperforms classical methods from quantile regression and extreme value theory, especially for high-dimensional predictor spaces and complex parameter response surfaces. An application to statistical post-processing of weather forecasts with precipitation data in the Netherlands is proposed.

stat.ME↗

Distributional regression and its evaluation with the CRPS: Bounds and convergence of the minimax risk

The theoretical advances on the properties of scoring rules over the past decades have broadened the use of scoring rules in probabilistic forecasting. In meteorological forecasting, statistical postprocessing techniques are essential to improve the forecasts made by deterministic physical models. Numerous state-of-the-art statistical postprocessing techniques are based on distributional regression evaluated with the Continuous Ranked Probability Score (CRPS). However, theoretical properties of such evaluation with the CRPS have solely considered the unconditional framework (i.e. without covariates) and infinite sample sizes. We extend these results and study the rate of convergence in terms of CRPS of distributional regression methods. We find the optimal minimax rate of convergence for a given class of distributions and show that the k-nearest neighbor method and the kernel method reach this optimal minimax rate.

math.ST↗

Hidden regular variation for point processes and the single/multiple large point heuristic

We consider regular variation for marked point processes with independent heavy-tailed marks and prove a single large point heuristic: the limit measure is concentrated on the cone of point measures with one single point. We then investigate successive hidden regular variation removing the cone of point measures with at most $k$ points, $k\geq 1$, and prove a multiple large point phenomenon: the limit measure is concentrated on the cone of point measures with $k+1$ points. We show how these results imply hidden regular variation in Skorokhod space of the associated risk process. Finally, we provide an application to risk theory in a reinsurance model where the $k$ largest claims are covered and we study the asymptotic behavior of the residual risk.

math.PR↗

Behavior of linear L2-boosting algorithms in the vanishing learning rate asymptotic

We investigate the asymptotic behaviour of gradient boosting algorithms when the learning rate converges to zero and the number of iterations is rescaled accordingly. We mostly consider L2-boosting for regression with linear base learner as studied in B{ü}hlmann and Yu (2003) and analyze also a stochastic version of the model where subsampling is used at each step (Friedman 2002). We prove a deterministic limit in the vanishing learning rate asymptotic and characterize the limit as the unique solution of a linear differential equation in an infinite dimensional function space. Besides, the training and test error of the limiting procedure are thoroughly analyzed. We finally illustrate and discuss our result on a simple numerical experiment where the linear L2-boosting operator is interpreted as a smoothed projection and time is related to its number of degrees of freedom.

stat.ML↗

Extreme quantile regression in a proportional tail framework

We revisit the model of heteroscedastic extremes initially introduced by Einmahl et al. (JRSSB, 2016) to describe the evolution of a non stationary sequence whose extremes evolve over time and adapt it into a general extreme quantile regression framework. We provide estimates for the extreme value index and the integrated skedasis function and prove their asymptotic normality. Our results are quite similar to those developed for heteroscedastic extremes but with a different proof approach emphasizing coupling arguments. We also propose a pointwise estimator of the skedasis function and a Weissman estimator of the conditional extreme quantile and prove the asymptotic normality of both estimators.

math.ST↗

The coupling method in extreme value theory

A coupling method is developed for univariate extreme value theory , providing an alternative to the use of the tail empirical/quantile processes. Emphasizing the Peak-over-Threshold approach that approximates the distribution above high threshold by the Generalized Pareto distribution, we compare the empirical distribution of exceedances and the empirical distribution associated to the limit Generalized Pareto model and provide sharp bounds for their Wasser-stein distance in the second order Wasserstein space. As an application , we recover standard results on the asymptotic behavior of the Hill estimator, the Weissman extreme quantile estimator or the probability weighted moment estimators, shedding some new light on the theory.

math.ST↗

A central limit theorem for functions of stationary max-stable random fields on $\mathbb{R}^d$

Max-stable random fields are very appropriate for the statistical modelling of spatial extremes. Hence, integrals of functions of max-stable random fields over a given region can play a key role in the assessment of the risk of natural disasters, meaning that it is relevant to improve our understanding of their probabilistic behaviour. For this purpose, in this paper, we propose a general central limit theorem for functions of stationary max-stable random fields on $\mathbb{R}^d$. Then, we show that appropriate functions of the Brown-Resnick random field with a power variogram and of the Smith random field satisfy the central limit theorem. Another strong motivation for our work lies in the fact that central limit theorems for random fields on $\mathbb{R}^d$ have been barely considered in the literature. As an application, we briefly show the usefulness of our results in a risk assessment context.

math.PR↗