SearcharxivSearch

arXiv subjects

Nicolas Meyer

Publications and source records attributed to Nicolas Meyer.

7 recordsLinked to original sources

Spatio-temporal modeling of urban extreme rainfall events at high resolution

Modeling precipitation and its accumulation over time and space is essential for flood risk assessment. In this paper, we analyze rainfall data collected over several years through a micro-scale precipitation sensor network in Montpellier, France. A novel spatio-temporal stochastic model is proposed for high-resolution urban extreme rainfall and combines realistic marginal behaviour and flexible dependence structure. Marginally, rainfall intensities are described by the Extended Generalized Pareto Distribution (EGPD), capturing both moderate and extreme events without threshold selection. Based on peaks-over-threshold theory for spatial processes, dependence during extreme episodes is modeled by an r-Pareto process with a non-separable variogram allowing for episode-specific advection, such that the displacement of rainfall cells is represented explicitly. Based on a catalog of extreme space-time episodes extracted from observations, parameters are estimated by a new composite likelihood based on joint exceedance indicators. Empirical advection velocities are derived beforehand from a radar reanalysis dataset. We show that the model accurately reproduces the spatio-temporal structure of extreme rainfall observed in the Montpellier OMSEV network and enables realistic stochastic scenario generation for flood risk assessment.

stat.AP

Kartezio: Evolutionary Design of Explainable Pipelines for Biomedical Image Analysis

An unresolved issue in contemporary biomedicine is the overwhelming number and diversity of complex images that require annotation, analysis and interpretation. Recent advances in Deep Learning have revolutionized the field of computer vision, creating algorithms that compete with human experts in image segmentation tasks. Crucially however, these frameworks require large human-annotated datasets for training and the resulting models are difficult to interpret. In this study, we introduce Kartezio, a modular Cartesian Genetic Programming based computational strategy that generates transparent and easily interpretable image processing pipelines by iteratively assembling and parameterizing computer vision functions. The pipelines thus generated exhibit comparable precision to state-of-the-art Deep Learning approaches on instance segmentation tasks, while requiring drastically smaller training datasets, a feature which confers tremendous flexibility, speed, and functionality to this approach. We also deployed Kartezio to solve semantic and instance segmentation problems in four real-world Use Cases, and showcase its utility in imaging contexts ranging from high-resolution microscopy to clinical pathology. By successfully implementing Kartezio on a portfolio of images ranging from subcellular structures to tumoral tissue, we demonstrated the flexibility, robustness and practical utility of this fully explicable evolutionary designer for semantic and instance segmentation.

cs.CV

Multivariate sparse clustering for extremes

Identifying directions where extreme events occur is a major challenge in multivariate extreme value analysis. In this paper, we use the concept of sparse regular variation introduced by Meyer and Wintenberger (2021)} to infer the tail dependence of a random vector X. This approach relies on the Euclidean projection onto the simplex which better exhibits the sparsity structure of the tail of X than the standard methods. Our procedure based on a rigorous methodology aims at capturing clusters of extremal coordinates of X. It also includes the identification of the threshold above which the values taken by X are considered as extreme. We provide an efficient and scalable algorithm called MUSCLE and apply it on numerical examples to highlight the relevance of our findings. Finally we illustrate our approach with financial return data.

math.ST

Bayesian two-interval test

The null hypothesis test (NHT) is widely used for validating scientific hypotheses but is actually highly criticized. Although Bayesian tests overcome several criticisms, some limits remain. We propose a Bayesian two-interval test (2IT) in which two hypotheses on an effect being present or absent are expressed as prespecified joint or disjoint intervals and their posterior probabilities are computed. The same formalism can be applied for superiority, non-inferiority, or equivalence tests. The 2IT was studied for three real examples and three sets of simulations (comparison of a proportion and a mean to a reference and comparison of two proportions). Several scenarios were created (with different sample sizes), and simulations were conducted to compute the probabilities of the parameter of interest being in the interval corresponding to either hypothesis given the data generated under one of the hypotheses. Posterior estimates were obtained using conjugacy with a low-informative prior. Bias was also estimated. The probability of accepting a hypothesis when that hypothesis is true progressively increases the sample size, tending towards 1, while the probability of accepting the other hypothesis is always very low (less than 5%) and tends towards 0. The speed of convergence varies with the gap between the hypotheses and with their width. In the case of a mean, the bias is low and rapidly becomes negligible. We propose a Bayesian test that follows a scientifically sound process, in which two interval hypotheses are explicitly used and tested. The proposed test has almost none of the limitations of the NHT and suggests new features, such as a rationale for serendipity or a justification for a "trend in data". The conceptual framework of the 2-IT also allows the calculation of a sample size and the use of sequential methods in numerous contexts.

stat.ME

Determining the Number of Components in PLS Regression on Incomplete Data

Partial least squares regression---or PLS---is a multivariate method in which models are estimated using either the SIMPLS or NIPALS algorithm. PLS regression has been extensively used in applied research because of its effectiveness in analysing relationships between an outcome and one or several components. Note that the NIPALS algorithm is able to provide estimates on incomplete data. Selection of the number of components used to build a representative model in PLS regression is an important problem. However, how to deal with missing data when using PLS regression remains a matter of debate. Several approaches have been proposed in the literature, including the $Q^2$ criterion, and the AIC and BIC criteria. Here we study the behavior of the NIPALS algorithm when used to fit a PLS regression for various proportions of missing data and for different types of missingness. We compare criteria for selecting the number of components for a PLS regression on incomplete data and on imputed datasets using three imputation methods: multiple imputation by chained equations, k-nearest neighbor imputation, and singular value decomposition imputation. Various criteria were tested with different proportions of missing data (ranging from 5% to 50%) under different missingness assumptions. Q2-leave-one-out component selection methods gave more reliable results than AIC and BIC-based ones.

stat.ME

New developments in Sparse PLS regression

Methods based on partial least squares (PLS) regression, which has recently gained much attention in the analysis of high-dimensional genomic datasets, have been developed since the early 2000s for performing variable selection. Most of these techniques rely on tuning parameters that are often determined by cross-validation (CV) based methods, which raises important stability issues. To overcome this, we have developed a new dynamic bootstrapbased method for significant predictor selection, suitable for both PLS regression and its incorporation into generalized linear models (GPLS). It relies on the establishment of bootstrap confidence intervals, that allows testing of the significance of predictors at preset type I risk $α$, and avoids the use of CV. We have also developed adapted versions of sparse PLS (SPLS) and sparse GPLS regression (SGPLS), using a recently introduced non-parametric bootstrap-based technique for the determination of the numbers of components. We compare their variable selection reliability and stability concerning tuning parameters determination, as well as their predictive ability, using simulated data for PLS and real microarray gene expression data for PLS-logistic classification. We observe that our new dynamic bootstrapbased method has the property of best separating random noise in y from the relevant information with respect to other methods, leading to better accuracy and predictive abilities, especially for non-negligible noise levels. Keywords: Variable selection, PLS, GPLS, Bootstrap, Stability

stat.ME

A new Universal Resample Stable Bootstrap-based Stopping Criterion in PLS Components Construction

We develop a new robust stopping criterion in Partial Least Squares Regressions (PLSR) components construction characterised by a high level of stability. This new criterion is defined as a universal one since it is suitable both for PLSR and its extension to Generalized Linear Regressions (PLSGLR). This criterion is based on a non-parametric bootstrap process and has to be computed algorithmically. It allows to test each successive components on a preset significant level alpha. In order to assess its performances and robustness with respect to different noise levels, we perform intensive datasets simulations, with a preset and known number of components to extract, both in the case n>p (n being the number of subjects and p the number of original predictors), and for datasets with n<p. We then use t-tests to compare the performance of our approach to some others classical criteria. The property of robustness is particularly tested through resampling processes on a real allelotyping dataset. Our conclusion is that our criterion presents also better global predictive performances, both in the PLSR and PLSGLR (Logistic and Poisson) frameworks.

stat.ME