SearcharxivSearch

arXiv subjects

Charles Tillier

Publications and source records attributed to Charles Tillier.

7 recordsLinked to original sources

Asymptotic Normality of Infinite Centered Random Forests -Application to Imbalanced Classification

Many classification tasks involve imbalanced data, in which a class is largely underrepresented. Several techniques consists in creating a rebalanced dataset on which a classifier is trained. In this paper, we study theoretically such a procedure, when the classifier is a Centered Random Forests (CRF). We establish a Central Limit Theorem (CLT) on the infinite CRF with explicit rates and exact constant. We then prove that the CRF trained on the rebalanced dataset exhibits a bias, which can be removed with appropriate techniques. Based on an importance sampling (IS) approach, the resulting debiased estimator, called IS-ICRF, satisfies a CLT centered at the prediction function value. For high imbalance settings, we prove that the IS-ICRF estimator enjoys a variance reduction compared to the ICRF trained on the original data. Therefore, our theoretical analysis highlights the benefits of training random forests on a rebalanced dataset (followed by a debiasing procedure) compared to using the original data. Our theoretical results, especially the variance rates and the variance reduction, appear to be valid for Breiman's random forests in our experiments.

stat.ML

Infinite random forests for imbalanced classification tasks

We study predictive probability inference in classification tasks using random forests under class imbalance. We focus on two simplified variants of Breiman's algorithm, namely subsampling Infinite Random Forests (IRFs) and under-sampling IRFs, and establish their asymptotic normality. In the under-sampling setting, training data from both classes are resampled to achieve balance, which enhances minority class representation but introduces a biased model. To correct this, we propose a debiasing procedure based on Importance Sampling (IS) using odds ratios. We instantiate our results using 1-Nearest Neighbor (1-NN) classifiers as base learners in the IRFs and prove the nearly minimax optimality of the approach for Lipschitz continuous objectives. We also show that the IS bagged 1-NN estimator matches the convergence rate of its subsampled counterpart while attaining lower asymptotic variance in most cases. Our theoretical findings are supported by simulation studies, highlighting the empirical benefits of the proposed approach.

math.ST

Hidden regular variation for point processes and the single/multiple large point heuristic

We consider regular variation for marked point processes with independent heavy-tailed marks and prove a single large point heuristic: the limit measure is concentrated on the cone of point measures with one single point. We then investigate successive hidden regular variation removing the cone of point measures with at most $k$ points, $k\geq 1$, and prove a multiple large point phenomenon: the limit measure is concentrated on the cone of point measures with $k+1$ points. We show how these results imply hidden regular variation in Skorokhod space of the associated risk process. Finally, we provide an application to risk theory in a reinsurance model where the $k$ largest claims are covered and we study the asymptotic behavior of the residual risk.

math.PR

Multivariate boundary regression models

In this work, we consider a multivariate regression model with one-sided errors. We assume for the regression function to lie in a general Hölder class and estimate it via a nonparametric local polynomial approach that consists of minimization of the local integral of a polynomial approximation lying above the data points. While the consideration of multivariate covariates offers an undeniable opportunity from an application-oriented standpoint, it requires a new method of proof to replace the established ones for the univariate case. The main purpose of this paper is to show the uniform consistency and to provide the rates of convergence of the considered nonparametric estimator for both multivariate random covariates and multivariate deterministic design points. To demonstrate the performance of the estimators, the small sample behavior is investigated in a simulation study in dimension two and three.

math.ST

Weighted Empirical Risk Minimization: Sample Selection Bias Correction based on Importance Sampling

We consider statistical learning problems, when the distribution $P'$ of the training observations $Z'_1,\; \ldots,\; Z'_n$ differs from the distribution $P$ involved in the risk one seeks to minimize (referred to as the test distribution) but is still defined on the same measurable space as $P$ and dominates it. In the unrealistic case where the likelihood ratio $Φ(z)=dP/dP'(z)$ is known, one may straightforwardly extends the Empirical Risk Minimization (ERM) approach to this specific transfer learning setup using the same idea as that behind Importance Sampling, by minimizing a weighted version of the empirical risk functional computed from the 'biased' training data $Z'_i$ with weights $Φ(Z'_i)$. Although the importance function $Φ(z)$ is generally unknown in practice, we show that, in various situations frequently encountered in practice, it takes a simple form and can be directly estimated from the $Z'_i$'s and some auxiliary information on the statistical population $P$. By means of linearization techniques, we then prove that the generalization capacity of the approach aforementioned is preserved when plugging the resulting estimates of the $Φ(Z'_i)$'s into the weighted empirical risk. Beyond these theoretical guarantees, numerical results provide strong empirical evidence of the relevance of the approach promoted in this article.

stat.ML

Semi-parametric transformation boundary regression models

In the context of nonparametric regression models with one-sided errors, we consider parametric transformations of the response variable in order to obtain independence between the errors and the covariates. We focus in this paper on stritcly increasing and continuous transformations. In view of estimating the tranformation parameter, we use a minimum distance approach and show the uniform consistency of the estimator under mild conditions. The boundary curve, i.e. the regression function, is estimated applying a smoothed version of a local constant approximation for which we also prove the uniform consistency. We deal with both cases of random covariates and deterministic (fixed) design points. To highlight the applicability of the procedures and to demonstrate their performance, the small sample behavior is investigated in a simulation study using the so-called Yeo-Johnson transformations.

math.ST

Regular variation of a random length sequence of random variables and application to risk assessment

When assessing risks on a finite-time horizon, the problem can often be reduced to the study of a random sequence $C(N)=(C_1,\ldots,C_N)$ of random length $N$, where $C(N)$ comes from the product of a matrix $A(N)$ of random size $N \times N$ and a random sequence $X(N)$ of random length $N$. Our aim is to build a regular variation framework for such random sequences of random length, to study their spectral properties and, subsequently, to develop risk measures. In several applications, many risk indicators can be expressed from the asymptotic behavior of $\vert \vert C(N)\vert\vert$, for some norm $\Vert \cdot \Vert$. We propose a generalization of Breiman Lemma that gives way to an asymptotic equivalent to $\Vert C(N) \Vert$ and provides risk indicators such as the ruin probability and the tail index for Shot Noise Processes on a finite-time horizon. Lastly, we apply our final result to a model used in dietary risk assessment and in non-life insurance mathematics to illustrate the applicability of our method.

math.PR