SearcharxivSearch

arXiv subjects

Liujun Chen

Publications and source records attributed to Liujun Chen.

10 recordsLinked to original sources

Generalized Linear Models for Extremes: Estimation and Inference in High Dimensions

We propose a regression model for the extreme tail of a response variable, in which covariates rescale the tail without changing its shape. A single covariate-dependent function then characterizes the entire conditional tail, in contrast to extreme quantile regression, which targets a quantile at a pre-specified level. The tail shape itself is left unrestricted: heavy-, light- and short-tailed responses are covered by the same framework. We specify the function through a link function and a linear combination of the covariates, which is in the spirit of a generalized linear model. In estimation, we match the parametric specification to the underlying tail function under a Bregman divergence, over a region localized at the largest observations. The resulting loss is convex, and an $\ell_1$-penalty allows the number of covariates to exceed the effective sample size. The tail localization makes the asymptotic theory deviate from that for classical penalized generalized linear models. Only the tail observations selected by a random threshold are used in the statistical analysis, making them dependent. We derive the convergence rate of the penalized estimator and propose a debiased estimator that is asymptotically normal, yielding confidence intervals for individual coefficients. Its asymptotic variance is determined by the covariance of the score, which under tail localization differs from the Hessian and must be estimated separately. We apply the method to automobile insurance claims data.

stat.ME

High Dimensional Mean Test for Shrinking Random Variables with Applications to Backtesting

We propose a high dimensional mean test framework for shrinking random variables, where the underlying random variables shrink to zero as the sample size increases. By pooling observations across overlapping subsets of dimensions, we estimate subsets means and test whether the maximum absolute mean deviates from zero. This approach overcomes cancellations that occur in simple averaging and remains valid even when marginal asymptotic normality fails. We establish theoretical properties of the test statistic and develop a multiplier bootstrap procedure to approximate its distribution. The method provides a flexible and powerful tool for the validation and comparative backtesting of value-at-risk. Simulations show superior performance in high-dimensional settings, and a real-data application demonstrates its practical effectiveness in backtesting.

stat.ME

Clustering Tails in High Dimension

One potential solution to combat the scarcity of tail observations in extreme value analysis is to integrate information from multiple datasets sharing similar tail properties, for instance, a common extreme value index. In other words, for a multivariate dataset, we intend to group dimensions into clusters first, before applying any pooling techniques. This paper addresses the clustering problem for a high dimensional dataset, according to their extreme value indices. We propose an iterative clustering procedure that sequentially partitions the variables into groups, ordered from the heaviest-tailed to the lightesttailed distributions. At each step, our method identifies and extracts a group of variables that share the highest extreme value index among the remaining ones. This approach differs fundamentally from conventional clustering methods such as using pre-estimated extreme value indices in a two-step clustering method. We show the consistency property of the proposed algorithm and demonstrate its finite-sample performance using a simulation study and a real data application.

stat.ME

Max-Linear Tail Regression

The relationship between a response variable and its covariates can vary significantly, especially in scenarios where covariates take on extremely high or low values. This paper introduces a max-linear tail regression model specifically designed to capture such extreme relationships. To estimate the regression coefficients within this framework, we propose a novel M-estimator based on extreme value theory. The consistency and asymptotic normality of our proposed estimator are rigorously established under mild conditions. Simulation results demonstrate that our estimation method outperforms the conditional least squares approach. We validate the practical applicability of our model through two case studies: one using financial data and the other using rainfall data.

stat.ME

High dimensional inference for extreme value indices

When applying multivariate extreme value statistics to analyze tail risk in compound events defined by a multivariate random vector, one often assumes that all dimensions share the same extreme value index. While such an assumption can be tested using a Wald-type test, the performance of such a test deteriorates as the dimensionality increases. This paper introduces novel tests for comparing extreme value indices in highdimensional settings, under both weak and general cross-sectional tail dependence. We establish the asymptotic behavior of the proposed tests. The proposed tests significantly outperform existing methods in high-dimensional scenarios in simulations. We demonstrate real-life applications of the proposed tests for two datasets previously assumed to have identical extreme value indices across all dimensions.

stat.ME

Tail Gini Functional under Asymptotic Independence

Tail Gini functional is a measure of tail risk variability for systemic risks, and has many applications in banking, finance and insurance. Meanwhile, there is growing attention on aymptotic independent pairs in quantitative risk management. This paper addresses the estimation of the tail Gini functional under asymptotic independence. We first estimate the tail Gini functional at an intermediate level and then extrapolate it to the extreme tails. The asymptotic normalities of both the intermediate and extreme estimators are established. The simulation study shows that our estimator performs comparatively well in view of both bias and variance. The application to measure the tail variability of weekly loss of individual stocks given the occurence of extreme events in the market index in Hong Kong Stock Exchange provides meaningful results, and leads to new insights in risk management.

stat.ME

Adapting the Hill estimator to distributed inference: dealing with the bias

The distributed Hill estimator is a divide-and-conquer algorithm for estimating the extreme value index when data are stored in multiple machines. In applications, estimates based on the distributed Hill estimator can be sensitive to the choice of the number of the exceedance ratios used in each machine. Even when choosing the number at a low level, a high asymptotic bias may arise. We overcome this potential drawback by designing a bias correction procedure for the distributed Hill estimator, which adheres to the setup of distributed inference. The asymptotically unbiased distributed estimator we obtained, on the one hand, is applicable to distributed stored data, on the other hand, inherits all known advantages of bias correction methods in extreme value statistics.

stat.ME

Distributed Inference for Tail Risk

For measuring tail risk with scarce extreme events, extreme value analysis is often invoked as the statistical tool to extrapolate to the tail of a distribution. The presence of large datasets benefits tail risk analysis by providing more observations for conducting extreme value analysis. However, large datasets can be stored distributedly preventing the possibility of directly analyzing them. In this paper, we introduce a comprehensive set of tools for examining the asymptotic behavior of tail empirical and quantile processes in the setting where data is distributed across multiple sources, for instance, when data are stored on multiple machines. Utilizing these tools, one can establish the oracle property for most distributed estimators in extreme value statistics in a straightforward way. The main theoretical challenge arises when the number of machines diverges to infinity. The number of machines resembles the role of dimensionality in high dimensional statistics. We provide various examples to demonstrate the practicality and value of our proposed toolkit.

stat.ME

Estimating Extreme Value Index by Subsampling for Massive Datasets with Heavy-Tailed Distributions

Modern statistical analyses often encounter datasets with massive sizes and heavy-tailed distributions. For datasets with massive sizes, traditional estimation methods can hardly be used to estimate the extreme value index directly. To address the issue, we propose here a subsampling-based method. Specifically, multiple subsamples are drawn from the whole dataset by using the technique of simple random subsampling with replacement. Based on each subsample, an approximate maximum likelihood estimator can be computed. The resulting estimators are then averaged to form a more accurate one. Under appropriate regularity conditions, we show theoretically that the proposed estimator is consistent and asymptotically normal. With the help of the estimated extreme value index, we can estimate high-level quantiles and tail probabilities of a heavy-tailed random variable consistently. Extensive simulation experiments are provided to demonstrate the promising performance of our method. A real data analysis is also presented for illustration purpose.

stat.ME

Multiple Hybrid Phase Transition: Bootstrap Percolation on Complex Networks with Communities

Bootstrap percolation is a well-known model to study the spreading of rumors, new products or innovations on social networks. The empirical studies show that community structure is ubiquitous among various social networks. Thus, studying the bootstrap percolation on the complex networks with communities can bring us new and important insights of the spreading dynamics on social networks. It attracts a lot of scientists' attentions recently. In this letter, we study the bootstrap percolation on Erd\H{o}s-R\'{e}nyi networks with communities and observed second order, hybrid (both second and first order) and multiple hybrid phase transitions, which is rare in natural system. Moreover, we have analytically solved this system and obtained the phase diagram, which is further justified well by the corresponding simulations.

physics.soc-ph