Searcharxiv⌕ Search

arXiv subjects

Thomas Laloë

Publications and source records attributed to Thomas Laloë.

6 recordsLinked to original sources

High-dimensional variable clustering based on maxima of a weakly dependent random process

We propose a new class of models for variable clustering called Asymptotic Independent block (AI-block) models, which defines population-level clusters based on the independence of the maxima of a multivariate stationary mixing random process among clusters. This class of models is identifiable, meaning that there exists a maximal element with a partial order between partitions, allowing for statistical inference. We also present an algorithm depending on a tuning parameter that recovers the clusters of variables without specifying the number of clusters \emph{a priori}. Our work provides some theoretical insights into the consistency of our algorithm, demonstrating that under certain conditions it can effectively identify clusters in the data with a computational complexity that is polynomial in the dimension. A data-driven selection method for the tuning parameter is also proposed. To further illustrate the significance of our work, we applied our method to neuroscience and environmental real-datasets. These applications highlight the potential and versatility of the proposed approach.

math.ST↗

Identifying regions of concomitant compound precipitation and wind speed extremes over Europe

The task of simplifying the complex spatio-temporal variables associated with climate modeling is of utmost importance and comes with significant challenges. In this research, our primary objective is to tailor clustering techniques to handle compound extreme events within gridded climate data across Europe. Specifically, we intend to identify subregions that display asymptotic independence concerning compound precipitation and wind speed extremes. To achieve this, we utilise daily precipitation sums and daily maximum wind speed data derived from the ERA5 reanalysis dataset spanning from 1979 to 2022. Our approach hinges on a tuning parameter and the application of a divergence measure to spotlight disparities in extremal dependence structures without relying on specific parametric assumptions. We propose a data-driven approach to determine the tuning parameter. This enables us to generate clusters that are spatially concentrated, which can provide more insightful information about the regional distribution of compound precipitation and wind speed extremes. In the process, we aim to elucidate the respective roles of extreme precipitation and wind speed in the resulting clusters. The proposed method is able to extract valuable information about extreme compound events while also significantly reducing the size of the dataset within reasonable computational timeframes.

stat.AP↗

Estimation of extreme $L^1$-multivariate expectiles with functional covariates

The present article is devoted to the semi-parametric estimation of multivariate expectiles for extreme levels. The considered multivariate risk measures also include the possible conditioning with respect to a functional covariate, belonging to an infinite-dimensional space. By using the first order optimality condition, we interpret these expectiles as solutions of a multidimensional nonlinear optimum problem. Then the inference is based on a minimization algorithm of gradient descent type, coupled with consistent kernel estimations of our key statistical quantities such as conditional quantiles, conditional tail index and conditional tail dependence functions. The method is valid for equivalently heavy-tailed marginals and under a multivariate regular variation condition on the underlying unknown random vector with arbitrary dependence structure. Our main result establishes the consistency in probability of the optimum approximated solution vectors with a speed rate. This allows us to estimate the global computational cost of the whole procedure according to the data sample size.

math.ST↗

Non-parametric estimator of a multivariate madogram for missing-data and extreme value framework

The modeling of dependence between maxima is an important subject in several applications in risk analysis. To this aim, the extreme value copula function, characterised via the madogram, can be used as a margin-free description of the dependence structure. From a practical point of view, the family of extreme value distributions is very rich and arise naturally as the limiting distribution of properly normalised component-wise maxima. In this paper, we investigate the nonparametric estimation of the madogram where data are completely missing at random. We provide the functional central limit theorem for the considered multivariate madrogram correctly normalized, towards a tight Gaussian process for which the covariance function depends on the probabilities of missing. Explicit formula for the asymptotic variance is also given. Our results are illustrated in a finite sample setting with a simulation study.

math.ST↗

Detection of dependence patterns with delay

The Unitary Events (UE) method is a popular and efficient method used this last decade to detect dependence patterns of joint spike activity among simultaneously recorded neurons. The first introduced method is based on binned coincidence count \citep{Grun1996} and can be applied on two or more simultaneously recorded neurons. Among the improvements of the methods, a transposition to the continuous framework has recently been proposed in \citep{muino2014frequent} and fully investigated in \citep{MTGAUE} for two neurons. The goal of the present paper is to extend this study to more than two neurons. The main result is the determination of the limit distribution of the coincidence count. This leads to the construction of an independence test between $L\geq 2$ neurons. Finally we propose a multiple test procedure via a Benjamini and Hochberg approach \citep{Benjamini1995}. All the theoretical results are illustrated by a simulation study, and compared to the UE method proposed in \citep{Grun2002}. Furthermore our method is applied on real data.

math.ST↗

Estimating level sets of a distribution function using a plug-in method: a multidimensional extension

This paper deals with the problem of estimating the level sets $L(c)= \{F(x) \geq c \}$, with $c \in (0,1)$, of an unknown distribution function $F$ on \mathbb{R}^d_+$. A plug-in approach is followed. That is, given a consistent estimator $F_n$ of $F$, we estimate $L(c)$ by $L_n(c)= \{F_n(x) \geq c \}$. We state consistency results with respect to the Hausdorff distance and the volume of the symmetric difference. These results can be considered as generalizations of results previously obtained, in a bivariate framework, in Di Bernardino et al. (2011). Finally we investigate the effects of scaling data on our consistency results.

math.ST↗