SearcharxivSearch

arXiv subjects

Ehsan Zamanzade

Publications and source records attributed to Ehsan Zamanzade.

10 recordsLinked to original sources

L-Estimation of Population Quantiles Using Ranked Set Sampling

Quantile estimation is central when interest lies in thresholds or tail behavior rather than the mean. When exact measurement is costly but units can be ranked cheaply, ranked set sampling (RSS) provides an attractive alternative to simple random sampling (SRS). We develop two families of RSS-based L-estimators for population quantiles that extend Stigler-type and Harrell--Davis estimators to the RSS framework. The first applies weighted-order-statistic estimation directly to the pooled ordered RSS sample and serves primarily as an exact conceptual benchmark, since its computational burden increases rapidly with the set size. The second exploits a decomposition induced by the RSS design that constructs $k$ pooled transformed-scale component estimators indexed by rank stratum and leads to a computationally scalable procedure. We derive large-sample results for these component estimators under regularity conditions; these results provide a principled first-order motivation for the combined estimators employed in practice. Simulation results across several distributions, quantile levels, and ranking qualities show consistent efficiency gains over empirical quantile estimators under both SRS and RSS, with the RSS Harrell--Davis version performing especially well for moderate and upper quantiles. Beyond the simulation study, we demonstrate the practical relevance of the proposed estimators through an application to NHANES transient elastography data, highlighting their usefulness for estimating clinically meaningful quantiles in a biomedical setting

stat.ME

Statistical Inference on the Cumulative Distribution Function using Judgment Post Stratification

In this work, we discuss a general class of the estimators for the cumulative distribution function (CDF) based on judgment post stratification (JPS) sampling scheme which includes both empirical and kernel distribution functions. Specifically, we obtain the expectation of the estimators in this class and show that they are asymptotically more efficient than their competitors in simple random sampling (SRS), as long as the rankings are better than random guessing. We find a mild condition that is necessary and sufficient for them to be asymptotically unbiased. We also prove that given the same condition, the estimators in this class are strongly uniformly consistent estimators of the true CDF, and converge in distribution to a normal distribution when the sample size goes to infinity. We then focus on the kernel distribution function (KDF) in the JPS design and obtain the optimal bandwidth. We next carry out a comprehensive Monte Carlo simulation to compare the performance of the KDF in the JPS design for different choices of sample size, set size, ranking quality, parent distribution, kernel function as well as both perfect and imperfect rankings set-ups with its counterpart in SRS design. It is found that the JPS estimator dramatically improves the efficiency of the KDF as compared to its SRS competitor for a wide range of the settings. Finally, we apply the described procedure on a real dataset from medical context to show their usefulness and applicability in practice.

stat.ME

Statistical Inference from Partially Nominated Sets: An Application to Estimating the Prevalence of Osteoporosis

This paper focuses on drawing statistical inference based on a novel variant of maxima or minima nomination sampling (NS) designs. These sampling designs are useful for obtaining more representative sample units from the tails of the population distribution using the available auxiliary ranking information. However, one common difficulty in performing NS in practice is that the researcher cannot obtain a nominated sample unless he/she uniquely determines the sample unit with the highest or the lowest rank in each set. To overcome this problem, a variant of NS which is called partial nomination sampling is proposed in which the researcher is allowed to declare that two or more units are tied in the ranks whenever he/she cannot find with high confidence the sample unit with the highest or the lowest rank with high confidence. Based on this sampling design, two asymptotically unbiased estimators are developed for the cumulative distribution function, which are obtained using maximum likelihood and moment-based approaches, and their asymptotic normality is proved. Several numerical studies have shown that the developed estimators have higher relative efficiencies than their counterpart in simple random sampling in analyzing either the upper or the lower tail of the parent distribution. The procedures that we are developed are then implemented on a real dataset from the Third National Health and Nutrition Examination Survey (NHANES III) to estimate the prevalence of osteoporosis among adult women aged 50 and over. It is shown that in some certain circumstances, the techniques that we have developed require only one-third of the sample size needed in SRS to achieve the desired precision. This results in a considerable reduction of time and cost compared to the standard SRS method.

stat.ME

A ranked-based estimator of the mean past lifetime with its application

The mean past lifetime (MPL) is an important tool in reliability and survival analysis for measuring the average time elapsed since the occurrence of an event, under the condition that the event has occurred before a specific time $t>0$. This article develops a nonparametric estimator for MPL based on observations collected according to ranked set sampling (RSS) design. It is shown that the estimator that we have developed is a strongly uniform consistent. It is also proved that the introduced estimator tends to a Gaussian process under some mild conditions. A Monte Carlo simulation study is employed to evaluate the performance of the proposed estimator with its competitor in simple random sampling (SRS). Our findings show the introduced estimator is more efficient than its counterpart estimator in SRS as long as the quality of ranking is better than random. Finally, an illustrative example is provided to describe the potential application of the developed estimator in assessing the average time between the infection and diagnosis in HIV patients.

stat.ME

Using a rank-based design in estimating prevalence of breast cancer

It is highly important for governments and health organizations to monitor the prevalence of breast cancer as a leading source of cancer-related death among women. However, the accurate diagnosis of this disease is expensive, especially in developing countries. This article concerns a cost-efficient method for estimating prevalence of breast cancer, when diagnosis is based on a comprehensive biopsy procedure. Multistage ranked set sampling (MSRSS) is utilized to develop a proportion estimator. This design employs some visually assessed cytological covariates, which are pertinent to determination of breast cancer, so as to provide the experimenter with a more informative sample. Theoretical properties of the proposed estimator are explored. Evidence from numerical studies is reported. The developed procedure can be substantially more efficient than its competitor in simple random sampling (SRS). In some situations, the proportion estimation in MSRSS needs around 76% fewer observations than that in SRS, given a precision level. Thus, using MSRSS may lead to a considerable reduction in cost with respect to SRS. In many medical studies, e.g. diagnosing breast cancer based on a full biopsy procedure, exact quantification is difficult (costly and/or time-consuming), but the potential sample units can be ranked fairly accurately without actual measurements. In this setup, multistage ranked set sampling is an appropriate design for developing cost-efficient statistical methods.

stat.ME

Inference on a Distribution Function from Ranked Set Samples

Consider independent observations $(X_1,R_1)$, $(X_2,R_2)$, \ldots, $(X_n,R_n)$ with random or fixed ranks $R_i \in \{1,2,\ldots,k\}$, while conditional on $R_i = r$, the random variable $X_i$ has the same distribution as the $r$-th order statistic within a random sample of size $k$ from an unknown continuous distribution function $F$. Such observation schemes are utilized in situations in which ranking observations is much easier than obtaining their precise values. Two well-known special cases are ranked set sampling (McIntyre 1952) and judgement post-stratification (MacEachern et al. 2004). Within a general setting including unbalanced ranked set sampling we derive and compare the asymptotic distributions of three different estimators of the distribution function $F$ as $n \to \infty$ with fixed $k$: The stratified estimator of Stokes and Sager (1988), the nonparametric maximum-likelihood estimator of Kvam and Samaniego (1994) and a moment-based estimator of Chen (2001). Our functional central limit theorems generalize and refine previous asymptotic analyses. In addition we discuss briefly pointwise and simultaneous confidence intervals for the distribution function $F$ with guaranteed coverage probability for finite sample sizes. The methods are illustrated with a real data example, and the potential impact of imperfect rankings is investigated in a small simulation experiment. All in all, the moment-based estimator seems to offer a good compromise between efficiency and robustness versus imperfect ranking, in addition to computational efficiency.

stat.ME

New ranked set sampling for estimating the population parameters

In this paper, a new modification of ranked set sampling (RSS) is suggested, namely; unified ranked set sampling (URSS) for estimating the population mean and variance. The performance of the empirical mean and variance estimators based on URSS are compared with their counterparts in ranked set sampling and simple random sampling (SRS) via Monte Carlo simulation. Simulation results indicate that the URSS estimators perform better than their counterparts using RSS and SRS designs when the ranking is perfect. When the ranking is imperfect, the URSS estimators still are superior than their counterparts in ranked set sampling and simple random sampling methods. Finally, an illustrative example is provided to show the efficiency of the new method in practice.

stat.ME

Variance Estimation in Ranked Set Sampling Using a Concomitant Variable

We propose a nonparametric variance estimator when ranked set sampling (RSS) and judgment post stratification (JPS) are applied by measuring a concomitant variable. Our proposed estimator is obtained by conditioning on observed concomitant values and using nonparametric kernel regression.

stat.ME

Permutation-Based Tests of Perfect Ranking

We improve three tests of perfect ranking in ranked set sampling proposed by Li and Balakrishnan (2008) using a permutation approach. This simple way of extending all three concepts to comparisons across different cycles increases the power. Two of the proposed tests are equivalent to tests from the literature, which were derived differently and are therefore generalized by the permutation-based tests.

stat.ME