SearcharxivSearch

arXiv subjects

Gilles Stupfler

Publications and source records attributed to Gilles Stupfler.

14 recordsLinked to original sources

Central limit theory for serial tail dependence estimators in heavy-tailed long memory linear time series

We prove multiple central limit theorems for serial tail dependence estimators in heavy-tailed long memory linear time series. The main theoretical tools are two novel multivariate reduction principles for partial sums of heavy-tailed long memory linear time series, subordinated over sliding windows and above a threshold growing with sample size. This requires addressing several substantial difficulties, including handling a nonlinear, sample-size dependent, and multivariate subordination mechanism, the dependence between several overlapping linear processes, and the lack of higher-order moments of the marginal distribution. Despite these obstacles, our assumptions are mild and, in particular, the innovation process is allowed to have infinite variance. A key feature of our theory is that our second reduction principle holds uniformly in the threshold, allowing central limit theory for empirical extremograms with sample quantiles as thresholds. This question has received little attention in the literature on serial extremal dependence estimation even though the version of empirical extremograms with random thresholds is ubiquitous in practice. We compare our results in several respects with those that may be obtained under short-range dependence, thereby discovering markedly different convergence rates and limit laws in our long memory setting.

math.ST

Logistic lasso regression with nearest neighbors for gradient-based dimension reduction

This paper investigates a new approach to estimate the gradient of the conditional probability given the covariates in the binary classification framework. The proposed approach consists of fitting a localized nearest-neighbor logistic model with $\ell_1$-penalty in order to cope with possibly high-dimensional covariates. Our theoretical analysis shows that the pointwise convergence rate of the gradient estimator is optimal under very mild assumptions. Moreover, using an outer product of such gradient estimates at several points in the covariate space, we provide a new method for estimating the central subspace, a well-known object allowing to carry out dimension reduction within the covariate space. Our implementation uses cross-validation on the misclassification rate to estimate the dimension of this subspace. We find that the proposed approach outperforms existing competitors in synthetic and real data applications.

math.ST

Central limit theory for Peaks-over-Threshold partial sums of long memory linear time series

Over the last 30 years, extensive work has been devoted to developing central limit theory for partial sums of subordinated long memory linear time series. A much less studied problem, motivated by questions that are ubiquitous in extreme value theory, is the asymptotic behavior of such partial sums when the subordination mechanism has a threshold depending on sample size, so as to focus on the right tail of the time series. This article substantially extends longstanding asymptotic techniques by allowing the subordination mechanism to depend on the sample size in this way and to grow at a polynomial rate, while permitting the innovation process to have infinite variance. The cornerstone of our theoretical approach is a tailored reduction principle, which enables the use of classical results on partial sums of long memory linear processes. In this way we obtain asymptotic theory for certain Peaks-over-Threshold estimators with deterministic or random thresholds. Applications cover both heavy- and light-tailed regimes, yielding unexpected results which, to the best of our knowledge, are new to the literature. A simulation study illustrates the relevance of our findings in finite samples.

math.PR

Regularized geometric quantiles and universal linear distribution functionals

Geometric quantiles are popular location functionals to build rank-based statistical procedures in multivariate settings. They are obtained through the minimization of a non-smooth convex objective function. As a result, the singularity of the directional derivatives leads to numerical instabilities and poor sample properties as well as surprising `phase transitions' from empirical to population distributions. To solve these issues, we introduce a regularized version of geometric distribution functions and quantiles that are provably close to the usual geometric concepts and share their qualitative properties, both in the empirical and continuous case, while allowing for a much broader applicability of asymptotic results without any moment condition. We also show that any linear assignment of probability measures (such as the univariate distribution function), that is also translation- and orthogonal-equivariant, necessarily coincides with one of our regularized geometric distribution functions.

math.ST

Concentration and excess risk bounds for imbalanced classification with synthetic oversampling

Synthetic oversampling of minority examples using SMOTE and its variants is a leading strategy for addressing imbalanced classification problems. Despite the success of this approach in practice, its theoretical foundations remain underexplored. We develop a theoretical framework to analyze the behavior of SMOTE and related methods when classifiers are trained on synthetic data. We first derive a uniform concentration bound on the discrepancy between the empirical risk over synthetic minority samples and the population risk on the true minority distribution. We then provide a nonparametric excess risk guarantee for kernel-based classifiers trained using such synthetic data. These results lead to practical guidelines for better parameter tuning of both SMOTE and the downstream learning algorithm. Numerical experiments are provided to illustrate and support the theoretical findings

stat.ML

Asymptotic Properties of Generalized Shortfall Risk Measures for Heavy-tailed Risks

We study a general risk measure called the generalized shortfall risk measure, which was first introduced in Mao and Cai (2018). It is proposed under the rank-dependent expected utility framework, or equivalently induced from the cumulative prospect theory. This risk measure can be flexibly designed to capture the decision maker's behavior toward risks and wealth when measuring risk. In this paper, we derive the first- and second-order asymptotic expansions for the generalized shortfall risk measure. Our asymptotic results can be viewed as unifying theory for, among others, distortion risk measures and utility-based shortfall risk measures. They also provide a blueprint for the estimation of these measures at extreme levels, and we illustrate this principle by constructing and studying a quantile-based estimator in a special case. The accuracy of the asymptotic expansions and of the estimator is assessed on several numerical examples.

q-fin.RM

Extreme expectile estimation for short-tailed data, with an application to market risk assessment

The use of expectiles in risk management has recently gathered remarkable momentum due to their excellent axiomatic and probabilistic properties. In particular, the class of elicitable law-invariant coherent risk measures only consists of expectiles. While the theory of expectile estimation at central levels is substantial, tail estimation at extreme levels has so far only been considered when the tail of the underlying distribution is heavy. This article is the first work to handle the short-tailed setting where the loss (e.g. negative log-returns) distribution of interest is bounded to the right and the corresponding extreme value index is negative. We derive an asymptotic expansion of tail expectiles in this challenging context under a general second-order extreme value condition, which allows to come up with two semiparametric estimators of extreme expectiles, and with their asymptotic properties in a general model of strictly stationary but weakly dependent observations. A simulation study and a real data analysis from a forecasting perspective are performed to verify and compare the proposed competing estimation procedures.

math.ST

Optimal pooling and distributed inference for the tail index and extreme quantiles

This paper investigates pooling strategies for tail index and extreme quantile estimation from heavy-tailed data. To fully exploit the information contained in several samples, we present general weighted pooled Hill estimators of the tail index and weighted pooled Weissman estimators of extreme quantiles calculated through a nonstandard geometric averaging scheme. We develop their large-sample asymptotic theory across a fixed number of samples, covering the general framework of heterogeneous sample sizes with different and asymptotically dependent distributions. Our results include optimal choices of pooling weights based on asymptotic variance and MSE minimization. In the important application of distributed inference, we prove that the variance-optimal distributed estimators are asymptotically equivalent to the benchmark Hill and Weissman estimators based on the unfeasible combination of subsamples, while the AMSE-optimal distributed estimators enjoy a smaller AMSE than the benchmarks in the case of large bias. We consider additional scenarios where the number of subsamples grows with the total sample size and effective subsample sizes can be low. We extend our methodology to handle serial dependence and the presence of covariates. Simulations confirm that our pooled estimators perform virtually as well as the benchmark estimators. Two applications to real weather and insurance data are showcased.

math.ST

Tail risk inference via expectiles in heavy-tailed time series

Expectiles define the only law-invariant, coherent and elicitable risk measure apart from the expectation. The popularity of expectile-based risk measures is steadily growing and their properties have been studied for independent data, but further results are needed to use extreme expectiles with dependent time series such as financial data. In this paper we establish a basis for inference on extreme expectiles and expectile-based marginal expected shortfall in a general $β$-mixing context that encompasses ARMA, ARCH and GARCH models with heavy-tailed innovations. Simulations and applications to financial returns show that the new estimators and confidence intervals greatly improve on existing ones when the data are dependent.

stat.ME

GARCH-UGH: A bias-reduced approach for dynamic extreme Value-at-Risk estimation in financial time series

The Value-at-Risk (VaR) is a widely used instrument in financial risk management. The question of estimating the VaR of loss return distributions at extreme levels is an important question in financial applications, both from operational and regulatory perspectives; in particular, the dynamic estimation of extreme VaR given the recent past has received substantial attention. We propose here a two-step bias-reduced estimation methodology called GARCH-UGH (Unbiased Gomes-de Haan), whereby financial returns are first filtered using an AR-GARCH model, and then a bias-reduced estimator of extreme quantiles is applied to the standardized residuals to estimate one-step ahead dynamic extreme VaR. Our results indicate that the GARCH-UGH estimates are more accurate than those obtained by combining conventional AR-GARCH filtering and extreme value estimates from the perspective of in-sample and out-of-sample backtestings of historical daily returns on several financial time series.

stat.AP

Joint inference on extreme expectiles for multivariate heavy-tailed distributions

The notion of expectiles, originally introduced in the context of testing for homoscedasticity and conditional symmetry of the error distribution in linear regression, induces a law-invariant, coherent and elicitable risk measure that has received a significant amount of attention in actuarial and financial risk management contexts. A number of recent papers have focused on the behaviour and estimation of extreme expectile-based risk measures and their potential for risk management. Joint inference of several extreme expectiles has however been left untouched; in fact, even the inference of a marginal extreme expectile turns out to be a difficult problem in finite samples. We investigate the simultaneous estimation of several extreme marginal expectiles of a random vector with heavy-tailed marginal distributions. This is done in a general extremal dependence model where the emphasis is on pairwise dependence between the margins. We use our results to derive accurate confidence regions for extreme expectiles, as well as a test for the equality of several extreme expectiles. Our methods are showcased in a finite-sample simulation study and on real financial data.

stat.ME

On a class of norms generated by nonnegative integrable distributions

We show that any distribution function on $\mathbb{R}^d$ with nonnegative, nonzero and integrable marginal distributions can be characterized by a norm on $\mathbb{R}^{d+1}$, called $F$-norm. We characterize the set of $F$-norms and prove that pointwise convergence of a sequence of $F$-norms to an $F$-norm is equivalent to convergence of the pertaining distribution functions in the Wasserstein metric. On the statistical side, an $F$-norm can easily be estimated by an empirical $F$-norm, whose consistency and weak convergence we establish. The concept of $F$-norms can be extended to arbitrary random vectors under suitable integrability conditions fulfilled by, for instance, normal distributions. The set of $F$-norms is endowed with a semigroup operation which, in this context, corresponds to ordinary convolution of the underlying distributions. Limiting results such as the central limit theorem can then be formulated in terms of pointwise convergence of products of $F$-norms. We conclude by showing how, using the geometry of $F$-norms, we may characterize nonnegative integrable distributions in $\mathbb{R}^d$ by simple compact sets in $\mathbb{R}^{d+1}$. We then relate convergence of those distributions in the Wasserstein metric to convergence of these characteristic sets with respect to Hausdorff distances.

math.PR

An Offspring of Multivariate Extreme-Value Theory: The Max-Characteristic Function

This paper introduces max-characteristic functions (max-CFs), which are an offspring of multivariate extreme-value theory. A max-CF characterizes the distribution of a random vector in R^d , whose components are nonnegative and have finite expectation. Pointwise convergence of max-CFs is shown to be equivalent with convergence with respect to the Wasserstein metric. The space of max-CFs is not closed in the sense of pointwise convergence. An inversion formula for max-CFs is established.

math.PR

Uniform strong consistency of a frontier estimator using kernel regression on high order moments

We consider the high order moments estimator of the frontier of a random pair introduced by Girard, S., Guillou, A., Stupfler, G. (2012). {\it Frontier estimation with kernel regression on high order moments}. In the present paper, we show that this estimator is strongly uniformly consistent on compact sets and its rate of convergence is given when the conditional cumulative distribution function belongs to the Hall class of distribution functions.

math.ST