SearcharxivSearch

arXiv subjects

Yanxi Hou

Publications and source records attributed to Yanxi Hou.

10 recordsLinked to original sources

Nonparametric Inference for Extreme CoVaR and CoES

Systemic risk measures quantify the potential risk to an individual financial constituent arising from the distress of entire financial system. As a generalization of two widely applied risk measures, Value-at-Risk and Expected Shortfall, the Conditional Value-at-Risk (CoVaR) and Conditional Expected Shortfall (CoES) have recently been receiving growing attention on applications in economics and finance, since they serve as crucial metrics for systemic risk measurement. However, existing approaches confront some challenges in statistical inference and asymptotic theories when estimating CoES, particularly at high risk levels. In this paper, within a framework of upper tail dependence, we propose several extrapolative methods to estimate both extreme CoVaR and CoES nonparametrically via an adjustment factor, which are intimately related to the nonparametric modelling of the tail dependence function. In addition, we study the asymptotic theories of all proposed extrapolative methods based on multivariate extreme value theory. Finally, some simulations and real data analyses are conducted to demonstrate the empirical performances of our methods.

stat.ME

Factorized Tail Volatility Model: Augmenting Excess-over-Threshold Method for High-Dimensional Hevay-Tailed Data

Ecess-over-Threshold method is a crucial technique in extreme value analysis, which approximately models larger observations over a threshold using a Generalized Pareto Distribution. This paper presents a comprehensive framework for analyzing tail risk in high-dimensional data by introducing the Factorized Tail Volatility Model (FTVM) and integrating it with central quantile models through the EoT method. This integrated framework is termed the FTVM-EoT method. In this framework, a quantile-related high-dimensional data model is employed to select an appropriate threshold at the central quantile for the EoT method, while the FTVM captures heteroscedastic tail volatility by decomposing tail quantiles into a low-rank linear factor structure and a heavy-tailed idiosyncratic component. The FTVM-EoT method is highly flexible, allowing for the joint modeling of central, intermediate, and extreme quantiles of high-dimensional data, thereby providing a holistic approach to tail risk analysis. In addition, we develop an iterative estimation algorithm for the FTVM-EoT method and establish the asymptotic properties of the estimators for latent factors, loadings, intermediate quantiles, and extreme quantiles. A validation procedure is introduced, and an information criterion is proposed for optimal factor selection. Simulation studies demonstrate that the FTVM-EoT method consistently outperforms existing methods at intermediate and extreme quantiles.

stat.ME

Tail Risk Equivalent Level Transition and Its Application for Estimating Extreme $L_p$-quantiles

$L_p$-quantile has recently been receiving growing attention in risk management since it has desirable properties as a risk measure and is a generalization of two widely applied risk measures, Value-at-Risk and Expectile. The statistical methodology for $L_p$-quantile is not only feasible but also straightforward to implement as it represents a specific form of M-quantile using $p$-power loss function. In this paper, we introduce the concept of Tail Risk Equivalent Level Transition (TRELT) to capture changes in tail risk when we make a risk transition between two $L_p$-quantiles. TRELT is motivated by PELVE in Li and Wang (2023) but for tail risk. As it remains unknown in theory how this transition works, we investigate the existence, uniqueness, and asymptotic properties of TRELT (as well as dual TRELT) for $L_p$-quantiles. In addition, we study the inference methods for TRELT and extreme $L_p$-quantiles by using this risk transition, which turns out to be a novel extrapolation method in extreme value theory. The asymptotic properties of the proposed estimators are established, and both simulation studies and real data analysis are conducted to demonstrate their empirical performance.

stat.ME

Bootstrap-based Inference for Bivariate Heteroscedastic Extremes with a Changing Tail Copula

This paper introduces a copula-based model for independent but non-identically distributed data with heteroscedastic extremes marginal and changing tail dependence structures. We establish a unified framework for inference by proving the weak convergence of the bivariate sequential tail empirical process and its empirical bootstrap counterpart. We derive the asymptotic properties of several estimators on the tail, including the quasi-tail copula, integrated scedasis function, and Hill estimator, treating them as functionals of the bivariate sequential tail empirical process. This process-centric approach enables the development of bootstrap-based methods and ensures the theoretical validity of the derived statistics. As an application of our inference method, we propose bootstrap-based tests for the equivalence of extreme value indices, the equivalence of scedasis functions, and non-changing tail dependence when marginal scedasis functions are identical. Our simulations validate the robustness and efficiency of the bootstrap-based tests.

stat.ME

Combining Structural and Unstructured Data: A Topic-based Finite Mixture Model for Insurance Claim Prediction

Modeling insurance claim amounts and classifying claims into different risk levels are critical yet challenging tasks. Traditional predictive models for insurance claims often overlook the valuable information embedded in claim descriptions. This paper introduces a novel approach by developing a joint mixture model that integrates both claim descriptions and claim amounts. Our method establishes a probabilistic link between textual descriptions and loss amounts, enhancing the accuracy of claims clustering and prediction. In our proposed model, the latent topic/component indicator serves as a proxy for both the thematic content of the claim description and the component of loss distributions. Specifically, conditioned on the topic/component indicator, the claim description follows a multinomial distribution, while the claim amount follows a component loss distribution. We propose two methods for model calibration: an EM algorithm for maximum a posteriori estimates, and an MH-within-Gibbs sampler algorithm for the posterior distribution. The empirical study demonstrates that the proposed methods work effectively, providing interpretable claims clustering and prediction.

stat.AP

Learning to Simulate: Generative Metamodeling via Quantile Regression

Stochastic simulation models effectively capture complex system dynamics but are often too slow for real-time decision-making. Traditional metamodeling techniques learn relationships between simulator inputs and a single output summary statistic, such as the mean or median. These techniques enable real-time predictions without additional simulations. However, they require prior selection of one appropriate output summary statistic, limiting their flexibility in practical applications. We propose a new concept: generative metamodeling. It aims to construct a "fast simulator of the simulator," generating random outputs significantly faster than the original simulator while preserving approximately equal conditional distributions. Generative metamodels enable rapid generation of numerous random outputs upon input specification, facilitating immediate computation of any summary statistic for real-time decision-making. We introduce a new algorithm, quantile-regression-based generative metamodeling (QRGMM), and establish its distributional convergence and convergence rate. Extensive numerical experiments demonstrate QRGMM's efficacy compared to other state-of-the-art generative algorithms in practical real-time decision-making scenarios.

cs.LG

Online Prediction of Extreme Conditional Quantiles via B-Spline Interpolation

Extreme quantiles are critical for understanding the behavior of data in the tail region of a distribution. It is challenging to estimate extreme quantiles, particularly when dealing with limited data in the tail. In such cases, extreme value theory offers a solution by approximating the tail distribution using the Generalized Pareto Distribution (GPD). This allows for the extrapolation beyond the range of observed data, making it a valuable tool for various applications. However, when it comes to conditional cases, where estimation relies on covariates, existing methods may require computationally expensive GPD fitting for different observations. This computational burden becomes even more problematic as the volume of observations increases, sometimes approaching infinity. To address this issue, we propose an interpolation-based algorithm named EMI. EMI facilitates the online prediction of extreme conditional quantiles with finite offline observations. Combining quantile regression and GPD-based extrapolation, EMI formulates as a bilevel programming problem, efficiently solvable using classic optimization methods. Once estimates for offline observations are obtained, EMI employs B-spline interpolation for covariate-dependent variables, enabling estimation for online observations with finite GPD fitting. Simulations and real data analysis demonstrate the effectiveness of EMI across various scenarios.

stat.ME

EVIboost for the Estimation of Extreme Value Index under Heterogeneous Extremes

Modeling heterogeneity on heavy-tailed distributions under a regression framework is challenging, and classical statistical methodologies usually place conditions on the distribution models to facilitate the learning procedure. However, these conditions are likely to overlook the complex dependence structure between the heaviness of tails and the covariates. Moreover, data sparsity on tail regions also makes the inference method less stable, leading to largely biased estimates for extreme-related quantities. This paper proposes a gradient boosting algorithm to estimate a functional extreme value index with heterogeneous extremes. Our proposed algorithm is a data-driven procedure that captures complex and dynamic structures in tail distributions. We also conduct extensive simulation studies to show the prediction accuracy of the proposed algorithm. In addition, we apply our method to a real-world data set to illustrate the state-dependent and time-varying properties of heavy-tail phenomena in the financial industry.

stat.ME

Clustering by the Probability Distributions from Extreme Value Theory

Clustering is an essential task to unsupervised learning. It tries to automatically separate instances into coherent subsets. As one of the most well-known clustering algorithms, k-means assigns sample points at the boundary to a unique cluster, while it does not utilize the information of sample distribution or density. Comparably, it would potentially be more beneficial to consider the probability of each sample in a possible cluster. To this end, this paper generalizes k-means to model the distribution of clusters. Our novel clustering algorithm thus models the distributions of distances to centroids over a threshold by Generalized Pareto Distribution (GPD) in Extreme Value Theory (EVT). Notably, we propose the concept of centroid margin distance, use GPD to establish a probability model for each cluster, and perform a clustering algorithm based on the covering probability function derived from GPD. Such a GPD k-means thus enables the clustering algorithm from the probabilistic perspective. Correspondingly, we also introduce a naive baseline, dubbed as Generalized Extreme Value (GEV) k-means. GEV fits the distribution of the block maxima. In contrast, the GPD fits the distribution of distance to the centroid exceeding a sufficiently large threshold, leading to a more stable performance of GPD k-means. Notably, GEV k-means can also estimate cluster structure and thus perform reasonably well over classical k-means. Thus, extensive experiments on synthetic datasets and real datasets demonstrate that GPD k-means outperforms competitors. The github codes are released in https://github.com/sixiaozheng/EVT-K-means.

cs.LG

Incrementally Zero-Shot Detection by an Extreme Value Analyzer

Human beings not only have the ability to recognize novel unseen classes, but also can incrementally incorporate the new classes to existing knowledge preserved. However, zero-shot learning models assume that all seen classes should be known beforehand, while incremental learning models cannot recognize unseen classes. This paper introduces a novel and challenging task of Incrementally Zero-Shot Detection (IZSD), a practical strategy for both zero-shot learning and class-incremental learning in real-world object detection. An innovative end-to-end model -- IZSD-EVer was proposed to tackle this task that requires incrementally detecting new classes and detecting the classes that have never been seen. Specifically, we propose a novel extreme value analyzer to detect objects from old seen, new seen, and unseen classes, simultaneously. Additionally and technically, we propose two innovative losses, i.e., background-foreground mean squared error loss alleviating the extreme imbalance of the background and foreground of images, and projection distance loss aligning the visual space and semantic spaces of old seen classes. Experiments demonstrate the efficacy of our model in detecting objects from both the seen and unseen classes, outperforming the alternative models on Pascal VOC and MSCOCO datasets.

cs.CV