SearcharxivSearch

arXiv subjects

Chanseok Park

Publications and source records attributed to Chanseok Park.

14 recordsLinked to original sources

A Note on Location Parameter Estimation using the Weighted Hodges-Lehmann Estimator

Robust design is one of the main tools employed by engineers for the facilitation of the design of high-quality processes. However, most real-world processes invariably contend with external uncontrollable factors, often denoted as outliers or contaminated data, which exert a substantial distorting effect upon the computed sample mean. In pursuit of mitigating the inherent bias entailed by outliers within the dataset, the concept of weight adjustment emerges as a prudent recourse, to make the sample more representative of the statistical population. In this sense, the intricate challenge lies in the judicious application of these diverse weights toward the estimation of an alternative to the robust location estimator. Different from the previous studies, this study proposes two categories of new weighted Hodges-Lehmann (WHL) estimators that incorporate weight factors in the location parameter estimation. To evaluate their robust performances in estimating the location parameter, this study constructs a set of comprehensive simulations to compare various location estimators including mean, weighted mean, weighted median, Hodges-Lehmann estimator, and the proposed WHL estimators. The findings unequivocally manifest that the proposed WHL estimators clearly outperform the traditional methods in terms of their breakdown points, biases, and relative efficiencies.

stat.ME

A goodness-of-fit test for the Birnbaum-Saunders distribution based on the probability plot

In the present paper, we develop a new goodness-of-fit test for the Birnbaum- Saunders distribution based on the probability plot. We utilize the sample correlation coefficient from the Birnbaum-Saunders probability plot as a measure of goodness of fit. Unfortunately, it is impossible or extremely difficult to obtain an explicit distribution of this sample correlation coefficient. To address this challenge, we employ extensive Monte Carlo simulations to obtain the empirical distribution of the sample correlation coefficient from the Birnbaum-Saunders probability plot. This empirical distribution allows us to determine the critical values alongside their corresponding significance levels, thus facilitating the computation of the p-value when the sample correlation coefficient is obtained. Finally, two real-data examples are provided for illustrative purposes.

stat.AP

Development of robust X-bar charts with unequal sample sizes

The traditional variable control charts, such as the X-bar chart, are widely used to monitor variation in a process. They have been shown to perform well for monitoring processes under the general assumptions that the observations are normally distributed without data contamination and that the sample sizes from the process are all equal. However, these two assumptions may not be met and satisfied in many practical applications and thus make them potentially limited for widespread application especially in production processes. In this paper, we alleviate this limitation by providing a novel method for constructing the robust X-bar control charts, which can simultaneously deal with both data contamination and unequal sample sizes. The proposed method for the process parameters is optimal in a sense of the best linear unbiased estimation. Numerical results from extensive Monte Carlo simulations and a real data analysis reveal that traditional control charts seriously underperform for monitoring process in the presence of data contamination and are extremely sensitive to even a single contaminated value, while the proposed robust control charts outperform in a manner that is comparable with the traditional ones, whereas they are far superior when the data are contaminated by outliers.

stat.ME

Robust explicit estimation of the log-logistic distribution with applications

The parameters of the log-logistic distribution are generally estimated based on classical methods such as maximum likelihood estimation, whereas these methods usually result in severe biased estimates when the data contain outliers. In this paper, we consider several alternative estimators, which not only have closed-form expressions, but also are quite robust to a certain level of data contamination. We investigate the robustness property of each estimator in terms of the breakdown point. The finite sample performance and effectiveness of these estimators are evaluated through Monte Carlo simulations and a real-data application. Numerical results demonstrate that the proposed estimators perform favorably in a manner that they are comparable with the maximum likelihood estimator for the data without contamination and that they provide superior performance in the presence of data contamination.

stat.ME

Empirical distributions of the robustified $t$-test statistics

Based on the median and the median absolute deviation estimators, and the Hodges-Lehmann and Shamos estimators, robustified analogues of the conventional $t$-test statistic are proposed. The asymptotic distributions of these statistics are recently provided. However, when the sample size is small, it is not appropriate to use the asymptotic distribution of the robustified $t$-test statistics for making a statistical inference including hypothesis testing, confidence interval, p-value, etc. In this article, through extensive Monte Carlo simulations, we obtain the empirical distributions of the robustified $t$-test statistics and their quantile values. Then these quantile values can be used for making a statistical inference.

stat.AP

A note on the g and h control charts

In this note, we revisit the $g$ and $h$ control charts that are commonly used for monitoring the number of conforming cases between the two consecutive appearances of nonconformities. It is known that the process parameter of these charts is usually unknown and estimated by using the maximum likelihood estimator and the minimum variance unbiased estimator. However, the minimum variance unbiased estimator in the control charts has been inappropriately used in the quality engineering literature. This observation motivates us to provide the correct minimum variance unbiased estimator and investigate theoretical and empirical biases of these estimators under consideration. Given that these charts are developed based on the underlying assumption that samples from the process should be balanced, which is often not satisfied in many practical applications, we propose a method for constructing these charts with unbalanced samples.

stat.AP

Investigation of finite-sample properties of robust location and scale estimators

When the experimental data set is contaminated, we usually employ robust alternatives to common location and scale estimators such as the sample median and Hodges-Lehmann estimators for location and the sample median absolute deviation and Shamos estimators for scale. It is well known that these estimators have high positive asymptotic breakdown points and are Fisher-consistent as the sample size tends to infinity. To the best of our knowledge, the finite-sample properties of these estimators, depending on the sample size, have not well been studied in the literature. In this paper, we fill this gap by providing their closed-form finite-sample breakdown points and calculating the unbiasing factors and relative efficiencies of the robust estimators through the extensive Monte Carlo simulations up to the sample size 100. The numerical study shows that the unbiasing factor improves the finite-sample performance significantly. In addition, we provide the predicted values for the unbiasing factors obtained by using the least squares method which can be used for the case of sample size more than 100.

stat.ME

B2BII - Data conversion from Belle to Belle II

We describe the conversion of simulated and recorded data by the Belle experiment to the Belle~II format with the software package \texttt{b2bii}. It is part of the Belle~II Analysis Software Framework. This allows the validation of the analysis software and the improvement of analyses based on the recorded Belle dataset using newly developed analysis tools.

hep-ex

A Quantile Variant of the EM Algorithm and Its Applications to Parameter Estimation with Interval Data

The expectation-maximization (EM) algorithm is a powerful computational technique for finding the maximum likelihood estimates for parametric models when the data are not fully observed. The EM is best suited for situations where the expectation in each E-step and the maximization in each M-step are straightforward. A difficulty with the implementation of the EM algorithm is that each E-step requires the integration of the log-likelihood function in closed form. The explicit integration can be avoided by using what is known as the Monte Carlo EM (MCEM) algorithm. The MCEM uses a random sample to estimate the integral at each E-step. However, the problem with the MCEM is that it often converges to the integral quite slowly and the convergence behavior can also be unstable, which causes a computational burden. In this paper, we propose what we refer to as the quantile variant of the EM (QEM) algorithm. We prove that the proposed QEM method has an accuracy of $O(1/K^2)$ while the MCEM method has an accuracy of $O_p(1/\sqrt{K})$. Thus, the proposed QEM method possesses faster and more stable convergence properties when compared with the MCEM algorithm. The improved performance is illustrated through the numerical studies. Several practical examples illustrating its use in interval-censored data problems are also provided.

stat.CO

Determination of the Joint Confidence Region of Optimal Operating Conditions in Robust Design by Bootstrap Technique

Robust design has been widely recognized as a leading method in reducing variability and improving quality. Most of the engineering statistics literature mainly focuses on finding "point estimates" of the optimum operating conditions for robust design. Various procedures for calculating point estimates of the optimum operating conditions are considered. Although this point estimation procedure is important for continuous quality improvement, the immediate question is "how accurate are these optimum operating conditions?" The answer for this is to consider interval estimation for a single variable or joint confidence regions for multiple variables. In this paper, with the help of the bootstrap technique, we develop procedures for obtaining joint "confidence regions" for the optimum operating conditions. Two different procedures using Bonferroni and multivariate normal approximation are introduced. The proposed methods are illustrated and substantiated using a numerical example.

stat.ME

Parameter Estimation from Censored Samples using the Expectation-Maximization Algorithm

This paper deals with parameter estimation when the data are randomly right censored. The maximum likelihood estimates from censored samples are obtained by using the expectation-maximization (EM) and Monte Carlo EM (MCEM) algorithms. We introduce the concept of the EM and MCEM algorithms and develop parameter estimation methods for a variety of distributions such as normal, Laplace and Rayleigh distributions. These proposed methods are illustrated with three examples.

stat.CO

A Note On the Use of Fiducial Limits for Control Charts

Many of the early works in the quality control literature construct control limits through the use of graphs and tables as described in Wortham and Ringer (1972). However, the methods used in this literature are restricted to using only the values that the graphs and tables can provide and to the case where the parameters of the underlying distribution are known. In this note, we briefly describe a technique which can be used to calculate exact control limits without the use of graphs or tables. We also describe what are commonly referred to in the literature as fiducial limits. Fiducial limits are often used as the limits in control charting when the parameters of the underlying distribution are unknown.

stat.AP

Note on the closed-form MLEs of k-component load-sharing systems

Recently Kim and Kvam (2004) and Singh, Sharma, Kumar (2008) proposed different load-sharing models and developed parametric inference for the these models. However, their parametric estimates are calculated using iterative numerical methods. In this note, we provide the general closed-form MLEs for the two load-sharing models provided by them.

stat.OT