Searcharxiv⌕ Search

arXiv subjects

Juan-Juan Cai

Publications and source records attributed to Juan-Juan Cai.

7 recordsLinked to original sources

Causal Inference for Heterogeneous Extreme Quantiles with Heavy-Tailed Outcomes

We propose a framework for estimating conditional extreme quantile treatment effects (CEQTEs) in observational studies with heavy-tailed outcomes. Our procedure first estimates intermediate conditional quantiles using inverse-probability-weighted (IPW) quantile regression and then extrapolates them to extreme levels using extreme value theory. Under a linear conditional quantile model, we show that the conditional and marginal distributions of each potential outcome share a common extreme value index (EVI), motivating two complementary Hill-type EVI estimators based on conditional and marginal information, respectively. On the theoretical front, we introduce an IPW tail quantile score process that bridges regression quantile score processes and uniform tail empirical processes while accounting for treatment assignment. We establish its functional weak convergence under mild regularity conditions, without requiring a max-domain-of-attraction condition. This result provides the probabilistic foundation for the asymptotic analysis of the proposed CEQTE estimators. Simulation studies demonstrate favorable finite-sample performance, and an application to NLSY79 data reveals substantial heterogeneity in the effect of college education on extremely high hourly wages across confounder-defined subpopulations.

stat.ME↗

Gradient boosting for extreme quantile regression

Extreme quantile regression provides estimates of conditional quantiles outside the range of the data. Classical quantile regression performs poorly in such cases since data in the tail region are too scarce. Extreme value theory is used for extrapolation beyond the range of observed values and estimation of conditional extreme quantiles. Based on the peaks-over-threshold approach, the conditional distribution above a high threshold is approximated by a generalized Pareto distribution with covariate dependent parameters. We propose a gradient boosting procedure to estimate a conditional generalized Pareto distribution by minimizing its deviance. Cross-validation is used for the choice of tuning parameters such as the number of trees and the tree depths. We discuss diagnostic plots such as variable importance and partial dependence plots, which help to interpret the fitted models. In simulation studies we show that our gradient boosting procedure outperforms classical methods from quantile regression and extreme value theory, especially for high-dimensional predictor spaces and complex parameter response surfaces. An application to statistical post-processing of weather forecasts with precipitation data in the Netherlands is proposed.

stat.ME↗

Interpretable random forest models through forward variable selection

Random forest is a popular prediction approach for handling high dimensional covariates. However, it often becomes infeasible to interpret the obtained high dimensional and non-parametric model. Aiming for obtaining an interpretable predictive model, we develop a forward variable selection method using the continuous ranked probability score (CRPS) as the loss function. Our stepwise procedure leads to a smallest set of variables that optimizes the CRPS risk by performing at each step a hypothesis test on a significant decrease in CRPS risk. We provide mathematical motivation for our method by proving that in population sense the method attains the optimal set. Additionally, we show that the test is consistent provided that the random forest estimator of a quantile function is consistent. In a simulation study, we compare the performance of our method with an existing variable selection method, for different sample sizes and different correlation strength of covariates. Our method is observed to have a much lower false positive rate. We also demonstrate an application of our method to statistical post-processing of daily maximum temperature forecasts in the Netherlands. Our method selects about 10% covariates while retaining the same predictive power.

stat.ME↗

Parametric and non-parametric estimation of extreme earthquake event: the joint tail inference for mainshocks and aftershocks

In an earthquake event, the combination of a strong mainshock and damaging aftershocks is often the cause of severe structural damages and/or high death tolls. The objective of this paper is to provide estimation for the probability of such extreme events where the mainshock and the largest aftershocks exceed certain thresholds. Two approaches are illustrated and compared -- a parametric approach based on previously observed stochastic laws in earthquake data, and a non-parametric approach based on bivariate extreme value theory. We analyze the earthquake data from the North Anatolian Fault Zone (NAFZ) in Turkey during 1965-2018 and show that the two approaches provide unifying results.

stat.AP↗

Improving precipitation forecasts using extreme quantile regression

Aiming to estimate extreme precipitation forecast quantiles, we propose a nonparametric regression model that features a constant extreme value index. Using local linear quantile regression and an extrapolation technique from extreme value theory, we develop an estimator for conditional quantiles corresponding to extreme high probability levels. We establish uniform consistency and asymptotic normality of the estimators. In a simulation study, we examine the performance of our estimator on finite samples in comparison with a method assuming linear quantiles. On a precipitation data set in the Netherlands, these estimators have greater predictive skill compared to the upper member of ensemble forecasts provided by a numerical weather prediction model.

stat.ME↗

Estimation of the marginal expected shortfall under asymptotic independence

We study the asymptotic behavior of the marginal expected shortfall when the two random variables are asymptotic independent but positive associated, which is modeled by the so-called tail dependent coefficient. We construct an estimator of the marginal expected shortfall which is shown to be asymptotically normal. The finite sample performance of the estimator is investigated in a small simulation study. The method is also applied to estimate the expected amount of rainfall at a weather station given that there is a once every 100 years rainfall at another weather station nearby.

math.ST↗

Estimation of extreme risk regions under multivariate regular variation

When considering d possibly dependent random variables, one is often interested in extreme risk regions, with very small probability p. We consider risk regions of the form ${\mathbf{z}\in\mathbb{R}^d:f(\mathbf{z})\leqβ}$, where f is the joint density and $β$ a small number. Estimation of such an extreme risk region is difficult since it contains hardly any or no data. Using extreme value theory, we construct a natural estimator of an extreme risk region and prove a refined form of consistency, given a random sample of multivariate regularly varying random vectors. In a detailed simulation and comparison study, the good performance of the procedure is demonstrated. We also apply our estimator to financial data.

math.ST↗