Searcharxiv⌕ Search

arXiv subjects

Stéphane Girard

Publications and source records attributed to Stéphane Girard.

At least 19 recordsLinked to original sources

Functional Extreme-PLS

We propose an extreme dimension reduction method extending the Extreme-PLS approach to the discretized functional framework, where the covariate lies in the infinite-dimensional Hilbert space $L^2([0,1])$ but is partially observed on a dense time grid. The ideas are partly borrowed from both Partial Least-Squares (PLS) and Sliced Inverse Regression (SIR) techniques. Accordingly, the method relies on the projection of the covariate onto a subspace and maximizes the covariance between its projection and the response conditionally on an extreme event capturing the tail-information. The covariate and the heavy-tailed response are supposed to be linked through a non-linear inverse single-index model and our goal is to infer the index in this regression framework. We propose a new family of estimators and show its asymptotic consistency with convergence rates under the model. Assuming mild conditions on the noise, most of the assumptions are stated in terms of regular variation unlike the standard literature on SIR and single-index regression. In addition, we expand the theoretical analysis with a model-free almost sure consistency result for the empirical tail-moments in a general separable Hilbert space. Finally, our results are illustrated on a finite-sample study with synthetic functional data as well as high-frequency financial data, highlighting the effectiveness of the dimension reduction for capturing tail dependence and for extreme risk management.

math.ST↗

Extreme-PLS with missing data under weak dependence

This paper develops a theoretical framework for Extreme Partial Least Squares (EPLS) dimension reduction in the presence of missing data and weak temporal dependence. Building upon the recent EPLS methodology for modeling extremal dependence between a response variable and high-dimensional covariates, we extend the approach to more realistic data settings where both serial correlation and missing-ness occur. Specifically, we consider a single-index inverse regression model under heavy-tailed conditions and introduce a Missing-at-Random (MAR) mechanism acting on the covariates, whose probability depends on the extremeness of the response. The asymptotic behavior of the proposed estimator is established within an alpha-mixing framework, leading to consistency results under regularly varying tails. Extensive Monte-Carlo experiments covering eleven dependence schemes (including ARMA, GARCH, and nonlinear ESTAR processes) demonstrate that the method performs robustly across a wide range of heavy-tailed and dependent scenarios, even when substantial portions of data are missing. A real-world application to environmental data further confirms the method's capacity to recover meaningful tail directions.

stat.ME↗

Shrinkage for Extreme Partial Least-Squares

This work focuses on dimension-reduction techniques for modelling conditional extreme values. Specifically, we investigate the idea that extreme values of a response variable can be explained by nonlinear functions derived from linear projections of an input random vector. In this context, the estimation of projection directions is examined, as approached by the Extreme Partial Least Squares (EPLS) method--an adaptation of the original Partial Least Squares (PLS) method tailored to the extreme-value framework. Further, a novel interpretation of EPLS directions as maximum likelihood estimators is introduced, utilizing the von Mises-Fisher distribution applied to hyperballs. The dimension reduction process is enhanced through the Bayesian paradigm, enabling the incorporation of prior information into the projection direction estimation. The maximum a posteriori estimator is derived in two specific cases, elucidating it as a regularization or shrinkage of the EPLS estimator. We also establish its asymptotic behavior as the sample size approaches infinity. A simulation data study is conducted in order to assess the practical utility of our proposed method. This clearly demonstrates its effectiveness even in moderate data problems within high-dimensional settings. Furthermore, we provide an illustrative example of the method's applicability using French farm income data, highlighting its efficacy in real-world scenarios.

stat.ME↗

Reparameterization of extreme value framework for improved Bayesian workflow

Using Bayesian methods for extreme value analysis offers an alternative to frequentist ones, with several advantages such as easily dealing with parametric uncertainty or studying irregular models. However, computations can be challenging and the efficiency of algorithms can be altered by poor parametrization choices. The focus is on the Poisson process characterization of univariate extremes and outline two key benefits of an orthogonal parameterization. First, Markov chain Monte Carlo convergence is improved when applied on orthogonal parameters. This analysis relies on convergence diagnostics computed on several simulations. Second, orthogonalization also helps deriving Jeffreys and penalized complexity priors, and establishing posterior propriety thereof. The proposed framework is applied to return level estimation of Garonne flow data (France).

stat.ME↗

On the use of a local $\hat{R}$ to improve MCMC convergence diagnostic

Diagnosing convergence of Markov chain Monte Carlo is crucial and remains an essentially unsolved problem. Among the most popular methods, the potential scale reduction factor, commonly named $\hat{R}$, is an indicator that monitors the convergence of output chains to a target distribution, based on a comparison of the between- and within-variances. Several improvements have been suggested since its introduction in the 90s. Here, we aim at better understanding the $\hat{R}$ behavior by proposing a localized version that focuses on quantiles of the target distribution. This new version relies on key theoretical properties of the associated population value. It naturally leads to proposing a new indicator $\hat{R}_\infty$, which is shown to allow both for localizing the Markov chain Monte Carlo convergence in different quantiles of the target distribution, and at the same time for handling some convergence issues not detected by other $\hat{R}$ versions.

math.ST↗

Dependence between Bayesian neural network units

The connection between Bayesian neural networks and Gaussian processes gained a lot of attention in the last few years, with the flagship result that hidden units converge to a Gaussian process limit when the layers width tends to infinity. Underpinning this result is the fact that hidden units become independent in the infinite-width limit. Our aim is to shed some light on hidden units dependence properties in practical finite-width Bayesian neural networks. In addition to theoretical results, we assess empirically the depth and width impacts on hidden units dependence properties.

stat.ML↗

Bayesian neural network unit priors and generalized Weibull-tail property

The connection between Bayesian neural networks and Gaussian processes gained a lot of attention in the last few years. Hidden units are proven to follow a Gaussian process limit when the layer width tends to infinity. Recent work has suggested that finite Bayesian neural networks may outperform their infinite counterparts because they adapt their internal representations flexibly. To establish solid ground for future research on finite-width neural networks, our goal is to study the prior induced on hidden units. Our main result is an accurate description of hidden units tails which shows that unit priors become heavier-tailed going deeper, thanks to the introduced notion of generalized Weibull-tail. This finding sheds light on the behavior of hidden units of finite Bayesian neural networks.

stat.ML↗

Dependence properties and Bayesian inference for asymmetric multivariate copulas

We study a broad class of asymmetric copulas introduced by Liebscher (2008) as a combination of multiple - usually symmetric - copulas. The main thrust of the paper is to provide new theoretical properties including exact tail dependence expressions and stability properties. A subclass of Liebscher copulas obtained by combining Fréchet copulas is studied in more details. We establish further dependence properties for copulas of this class and show that they are characterized by an arbitrary number of singular components. Furthermore, we introduce a novel iterative representation for general Liebscher copulas which de facto insures uniform margins, thus relaxing a constraint of Liebscher's original construction. Besides, we show that this iterative construction proves useful for inference by developing an Approximate Bayesian computation sampling scheme. This inferential procedure is demonstrated on simulated data.

math.ST↗

On kernel smoothing for extremal quantile regression

Nonparametric regression quantiles obtained by inverting a kernel estimator of the conditional distribution of the response are long established in statistics. Attention has been, however, restricted to ordinary quantiles staying away from the tails of the conditional distribution. The purpose of this paper is to extend their asymptotic theory far enough into the tails. We focus on extremal quantile regression estimators of a response variable given a vector of covariates in the general setting, whether the conditional extreme-value index is positive, negative, or zero. Specifically, we elucidate their limit distributions when they are located in the range of the data or near and even beyond the sample boundary, under technical conditions that link the speed of convergence of their (intermediate or extreme) order with the oscillations of the quantile function and a von-Mises property of the conditional distribution. A simulation experiment and an illustration on real data were presented. The real data are the American electric data where the estimation of conditional extremes is found to be of genuine interest.

math.ST↗

Functional kernel estimators of large conditional quantiles

We address the estimation of conditional quantiles when the covariate is functional and when the order of the quantiles converges to one as the sample size increases. In a first time, we investigate to what extent these large conditional quantiles can still be estimated through a functional kernel estimator of the conditional survival function. Sufficient conditions on the rate of convergence of their order to one are provided to obtain asymptotically Gaussian distributed estimators. In a second time, basing on these result, a functional Weissman estimator is derived, permitting to estimate large conditional quantiles of arbitrary large order. These results are illustrated on finite sample situations.

math.ST↗

A note on extreme values and kernel estimators of sample boundaries

In a previous paper, we studied a kernel estimate of the upper edge of a two-dimensional bounded set, based upon the extreme values of a Poisson point process. The initial paper "Geffroy J. (1964) Sur un problème d'estimation géométrique.Publications de l'Institut de Statistique de l'Université de Paris, XIII, 191-200" on the subject treats the frontier as the boundary of the support set for a density and the points as a random sample. We claimed in"Girard, S. and Jacob, P. (2004) Extreme values and kernel estimates of point processes boundaries.ESAIM: Probability and Statistics, 8, 150-168" that we are able to deduce the random sample case fr om the point process case. The present note gives some essential indications to this end, including a method which can be of general interest.

math.ST↗

Kernel discriminant analysis and clustering with parsimonious Gaussian process models

This work presents a family of parsimonious Gaussian process models which allow to build, from a finite sample, a model-based classifier in an infinite dimensional space. The proposed parsimonious models are obtained by constraining the eigen-decomposition of the Gaussian processes modeling each class. This allows in particular to use non-linear mapping functions which project the observations into infinite dimensional spaces. It is also demonstrated that the building of the classifier can be directly done from the observation space through a kernel function. The proposed classification method is thus able to classify data of various types such as categorical data, functional data or networks. Furthermore, it is possible to classify mixed data by combining different kernels. The methodology is as well extended to the unsupervised classification case. Experimental results on various data sets demonstrate the effectiveness of the proposed method.

stat.ME↗

Comparison of Weibull tail-coefficient estimators

We address the problem of estimating the Weibull tail-coefficient which is the regular variation exponent of the inverse failure rate function. We propose a family of estimators of this coefficient and an associate extreme quantile estimator. Their asymptotic normality are established and their asymptotic mean-square errors are compared. The results are illustrated on some finite sample situations.

stat.ME↗

Frontier estimation with local polynomials and high power-transformed data

We present a new method for estimating the frontier of a sample. The estimator is based on a local polynomial regression on the power-transformed data. We assume that the exponent of the transformation goes to infinity while the bandwidth goes to zero. We give conditions on these two parameters to obtain almost complete convergence. The asymptotic conditional bias and variance of the estimator are provided and its good performance is illustrated on some finite sample situations.

stat.ME↗

A Note on Sliced Inverse Regression with Regularizations

In "Li, L. and Yin, X. (2008). Sliced Inverse Regression with Regularizations. Biometrics, 64(1):124--131" a ridge SIR estimator is introduced as the solution of a minimization problem and computed thanks to an alternating least-squares algorithm. This methodology reveals good performance in practice. In this note, we focus on the theoretical properties of the estimator. Is it shown that the minimization problem is degenerated in the sense that only two situations can occur: Either the ridge SIR estimator does not exist or it is zero.

math.ST↗

Estimation procedures for a semiparametric family of bivariate copulas

In this paper, we propose simple estimation methods dedicated to a semiparametric family of bivariate copulas. These copulas can be simply estimated through the estimation of their univariate generating function. We take profit of this result to estimate the associated measures of association as well as the high probability regions of the copula. These procedures are illustrated on simulations and on real data.

stat.ME↗

Auto-associative models, nonlinear Principal component analysis, manifolds and projection pursuit

In this paper, auto-associative models are proposed as candidates to the generalization of Principal Component Analysis. We show that these models are dedicated to the approximation of the dataset by a manifold. Here, the word "manifold" refers to the topology properties of the structure. The approximating manifold is built by a projection pursuit algorithm. At each step of the algorithm, the dimension of the manifold is incremented. Some theoretical properties are provided. In particular, we can show that, at each step of the algorithm, the mean residuals norm is not increased. Moreover, it is also established that the algorithm converges in a finite number of steps. Some particular auto-associative models are exhibited and compared to the classical PCA and some neural networks models. Implementation aspects are discussed. We show that, in numerous cases, no optimization procedure is required. Some illustrations on simulated and real data are presented.

stat.ML↗