Searcharxiv⌕ Search

arXiv subjects

Jeong-Soo Park

Publications and source records attributed to Jeong-Soo Park.

9 recordsLinked to original sources

Automated selection of r for stationary and nonstationary models for r largest order statistics

In generalized extreme value model for the r largest order statistics, denoted by rGEV, the selection of r is critical. The existing entropy difference test for selecting r is applicable to large sample. Another existing method (the score test with parametric bootstrap) is applicable to small sample, but computationally demanding. To address this problem for small sample, we propose a new method using a sequence of the goodness-of-fit tests based on the conditional cumulative distribution function (CCDF). The proposed CCDF test is easy to implement and computationally fast. The Cram{é}r-von Mises test was employed for the goodness-of-fit purpose. The proposed method is compared via Monte Carlo simulations with existing methods including the spacings, the score, and the entropy difference tests. The proposed CCDF test turned out to perform well for both small and large samples, comparable to the spacings and entropy difference tests. The utility of the proposed method is illustrated by an application to the r largest daily rainfall data in Korea. Additionally, we extended the existing methods and the CCDF test to a nonstationary rGEV model. Wide applicability of the proposed method are discussed.

stat.ME↗

Model averaging with mixed criteria for estimating high quantiles of extreme values: Application to heavy rainfall

Accurately estimating high quantiles beyond the largest observed value is crucial for risk assessment and devising effective adaptation strategies to prevent a greater disaster. The generalized extreme value distribution is widely used for this purpose, with L-moment estimation (LME) and maximum likelihood estimation (MLE) being the primary methods. However, estimating high quantiles with a small sample size becomes challenging when the upper endpoint is unbounded, or equivalently, when there are larger uncertainties involved in extrapolation. This study introduces an improved approach using a model averaging (MA) technique. The proposed method combines MLE and LME to construct candidate submodels and assign weights effectively. The properties of the proposed approach are evaluated through Monte Carlo simulations and an application to maximum daily rainfall data in Korea. In addition, theoretical properties of the MA estimator are examined, including the asymptotic variance with random weights. A surrogate model of MA estimation is also developed and applied for further analysis. Finally, a Bayesian model averaging approach is considered to reduce the estimation bias occurring in the MA methods.

stat.ME↗

Generalized method of L-moment estimation for stationary and nonstationary extreme value models

Precisely estimating out-of-sample upper quantiles is very important in risk assessment and in engineering practice for structural design to prevent a greater disaster. For this purpose, the generalized extreme value (GEV) distribution has been broadly used. To estimate the parameters of GEV distribution, the maximum likelihood estimation (MLE) and L-moment estimation (LME) methods have been primarily employed. For a better estimation using the MLE, several studies considered the generalized MLE (penalized likelihood or Bayesian) methods to cooperate with a penalty function or prior information for parameters. However, a generalized LME method for the same purpose has not been developed yet in the literature. We thus propose the generalized method of L-moment estimation (GLME) to cooperate with a penalty function or prior information. The proposed estimation is based on the generalized L-moment distance and a multivariate normal likelihood approximation. Because the L-moment estimator is more efficient and robust for small samples than the MLE, we reasonably expect the advantages of LME to continue to hold for GLME. The proposed method is applied to the stationary and nonstationary GEV models with two novel (data-adaptive) penalty functions to correct the bias of LME. A simulation study indicates that the biases of LME are considerably corrected by the GLME with slight increases in the standard error. Applications to US flood damage data and maximum rainfall at Phliu Agromet in Thailand illustrate the usefulness of the proposed method. This study may promote further work on penalized or Bayesian inferences based on L-moments.

stat.ME↗

Building nonstationary extreme value model using L-moments

The maximum likelihood estimation for a time-dependent nonstationary (NS) extreme value model is often too sensitive to influential observations, such as large values toward the end of a sample. Thus, alternative methods using L-moments have been developed in NS models to address this problem while retaining the advantages of the stationary L-moment method. However, one method using L-moments displays inferior performance compared to stationary estimation when the data exhibit a positive trend in variance. To address this problem, we propose a new algorithm for efficiently estimating the NS parameters. The proposed method combines L-moments and robust regression, using standardized residuals. A simulation study demonstrates that the proposed method overcomes the mentioned problem. The comparison is conducted using conventional and redefined return level estimates. An application to peak streamflow data in Trehafod in the UK illustrates the usefulness of the proposed method. Additionally, we extend the proposed method to a NS extreme value model in which physical covariates are employed as predictors. Furthermore, we consider a model selection criterion based on the cross-validated generalized L-moment distance as an alternative to the likelihood-based criteria.

stat.ME↗

Modeling climate extremes using the four-parameter kappa distribution for $r$-largest order statistics

Accurate estimation of the T-year return levels of climate extremes using statistical distribution is a critical step in the projection of future climate and in engineering design for disaster response. We show how the estimation of such quantities can be improved by fitting {the four-parameter kappa distribution for $r$-largest order statistics} (rK4D), which was developed in this study. The rK4D is an extension of {the generalized extreme value distribution for $r$-largest order statistics} (rGEVD), similar to the four-parameter kappa distribution (K4D), which is an extension of the generalized extreme value distribution (GEVD). This new distribution (rK4D) can be useful not only for fitting data when three parameters in the GEVD are not sufficient to capture the variability of the extreme observations, but also in reducing the estimation uncertainty by making use of the r-largest extreme observations instead of only the block maxima. We derive a joint probability density function (PDF) of rK4D and the marginal and conditional cumulative distribution functions and PDFs. To estimate the parameters, the maximum likelihood estimation and the maximum penalized likelihood estimation methods were considered. The usefulness and practical effectiveness of the rK4D are illustrated by the Monte Carlo simulation and by an application to the Bangkok extreme rainfall data. A few new distributions for $r$-largest order statistics are also derived as special cases of the rK4D, such as the $r$-largest logistic, the $r$-largest generalized logistic, and the $r$-largest generalized Gumbel distributions. These distributions for $r$-largest order statistics would be useful in modeling extreme values for many research areas, including hydrology and climatology.

stat.ME↗

Generalized logistic model for $r$ largest order statistics, with hydrological application

The effective use of available information in extreme value analysis is critical because extreme values are scarce. Thus, using the $r$ largest order statistics (rLOS) instead of the block maxima is encouraged. Based on the four-parameter kappa model for the rLOS (rK4D), we introduce a new distribution for the rLOS as a special case of the rK4D. That is the generalized logistic model for rLOS (rGLO). This distribution can be useful when the generalized extreme value model for rLOS is no longer efficient to capture the variability of extreme values. Moreover, the rGLO enriches a pool of candidate distributions to determine the best model to yield accurate and robust quantile estimates. We derive a joint probability density function, the marginal and conditional distribution functions of new model. The maximum likelihood estimation, delta method, profile likelihood, order selection by the entropy difference test, cross-validated likelihood criteria, and model averaging were considered for inferences. The usefulness and practical effectiveness of the rGLO are illustrated by the Monte Carlo simulation and an application to extreme streamflow data in Bevern Stream, UK.

stat.AP↗

Penalized Likelihood Approach for the Four-parameter Kappa Distribution

The four-parameter kappa distribution (K4D) is a generalized form of some commonly used distributions such as generalized logistic, generalized Pareto, generalized Gumbel, and generalized extreme value (GEV) distributions. Owing to its flexibility, the K4D is widely applied in modeling in several fields such as hydrology and climatic change. For the estimation of the four parameters, the maximum likelihood approach and the method of L-moments are usually employed. The L-moment estimator (LME) method works well for some parameter spaces, with up to a moderate sample size, but it is sometimes not feasible in terms of computing the appropriate estimates. Meanwhile, the maximum likelihood estimator (MLE) is optimal for large samples and applicable to a very wide range of situations, including non-stationary data. However, using the MLE of K4D with small sample sizes shows substantially poor performance in terms of a large variance of the estimator. We therefore propose a maximum penalized likelihood estimation (MPLE) of K4D by adjusting the existing penalty functions that restrict the parameter space. Eighteen combinations of penalties for two shape parameters are considered and compared. The MPLE retains modeling flexibility and large sample optimality while also improving on small sample properties. The properties of the proposed estimator are verified through a Monte Carlo simulation, and an application case is demonstrated taking Thailand's annual maximum temperature data. Based on this study, we suggest using combinations of penalty functions in general.

stat.ME↗

Iterative Method for Tuning Complex Simulation Code

Tuning a complex simulation code refers to the process of improving the agreement of a code calculation with respect to a set of experimental data by adjusting parameters implemented in the code. This process belongs to the class of inverse problems or model calibration. For this problem, the approximated nonlinear least squares (ANLS) method based on a Gaussian process (GP) metamodel has been employed by some researchers. A potential drawback of the ANLS method is that the metamodel is built only once and not updated thereafter. To address this difficulty, we propose an iterative algorithm in this study. In the proposed algorithm, the parameters of the simulation code and GP metamodel are alternatively re-estimated and updated by maximum likelihood estimation and the ANLS method. This algorithm uses both computer and experimental data repeatedly until convergence. A study using toy-models including inexact computer code with bias terms reveals that the proposed algorithm performs better than the ANLS method and the conditional-likelihood-based approach. Finally, an application to a nuclear fusion simulation code is illustrated.

stat.CO↗

Integration of max-stable processes and Bayesian model averaging to predict extreme climatic events in multi-model ensembles

Projections of changes in extreme climate are sometimes predicted by using multi-model ensemble methods such as Bayesian model averaging (BMA) embedded with the generalized extreme value (GEV) distribution. BMA is a popular method for combining the forecasts of individual simulation models by weighted averaging and characterizing the uncertainty induced by simulating the model structure. This method is referred to as the GEV-embedded BMA. It is, however, based on a point-wise analysis of extreme events, which means it overlooks the spatial dependency between nearby grid cells. Instead of a point-wise model, a spatial extreme model such as the max-stable process (MSP) is often employed to improve precision by considering spatial dependency. We propose an approach that integrates the MSP into BMA, which is referred to as the MSP-BMA herein. The superiority of the proposed method over the GEV-embedded BMA is demonstrated by using extreme rainfall intensity data on the Korean peninsula from Coupled Model Intercomparison Project Phase 5 (CMIP5) multi-models. The reanalysis data called APHRODITE (Asian Precipitation Highly-Resolved Observational Data Integration Towards Evaluation, v1101) and 17 CMIP5 models are examined for 10 grid boxes in Korea. In this example, the MSP-BMA achieves a variance reduction over the GEV-embedded BMA. The bias inflation by MSP-BMA over the GEV-embedded BMA is also discussed. A by-product technical advantage of the MSP-BMA is that tedious `regridding' is not required before and after the analysis while it should be done for the GEV-embedded BMA.

stat.AP↗