SearcharxivSearch

arXiv subjects

Edsel A. Pena

Publications and source records attributed to Edsel A. Pena.

9 recordsLinked to original sources

"Game, Set, Match": Double Delight Watching a Grand Slam Tennis Match

Probabilistic properties of tennis scoring systems are examined and compared with best-of-K systems. A model, where each player has his/her own probability of winning his/her service point and which remains invariant for the duration of the match, and where outcomes of points played are independent of each other, is assumed. Probabilities of winning a game tie-breaker, a game, a set tie-breaker, a set, and the match are obtained. Since tennis scoring systems are unique, probability calculations require decomposing big and complicated problems into smaller and simpler constituent problems, solving these sub-problems, then combining to obtain the solution to the big problem. The problems that arise from tennis scoring systems offer excellent pedagogical venues for teaching probability, in particular, the use of the Theorem of Total Probability and the Iterated Rules for Mean, Variance, and Covariance. There are also many interesting questions in tennis, foremost of which is whether a tennis match under this assumption will actually end with probability one; or whether when two players of `equal abilities' play a match, the first server possesses an advantage. These questions are addressed in this work. Tennis scoring systems are technically statistical decision systems to determine the better player. Since such a decision system is based on a finite number of points played, erroneous decisions could arise, such as the inferior player winning the match. We compare different systems in terms of the probability of the better player winning, as well as the duration of the match in terms of the number of points played.

math.PR

Maximum Agreement Linear Predictors

This paper studies predictor functions motivated by maximizing a measure of agreement with the predictand. Specifically, it examines distributional properties and predictive performance of the estimated maximum agreement linear predictor (MALP), the linear predictor maximizing Lin's concordance correlation coefficient (CCC) between the predictor and the predictand. It is compared and contrasted, theoretically and through computer experiments, with the estimated least-squares linear predictor (LSLP), with respect to some performance measures. Finite-sample and asymptotic properties are obtained, and confidence intervals and prediction intervals are also presented. Predictors are illustrated using two real data sets: an eye data set and a body fat data set. Results indicate that the estimated MALP is a viable alternative to the estimated LSLP if one desires a predictor whose predicted values possesses higher agreement with the predictand values, as measured by the CCC.

stat.ME

Improved Multiple Confidence Intervals via Thresholding Informed by Prior Information

Consider a statistical problem where a set of parameters are of interest to a researcher. Then multiple confidence intervals can be constructed to infer the set of parameters simultaneously. The constructed multiple confidence intervals are the realization of a multiple interval estimator (MIE), the main focus of this study. In particular, a thresholding approach is introduced to improve the performance of the MIE. The developed thresholds require additional information, so a prior distribution is assumed for this purpose. The MIE procedure is then evaluated by two performance measures: a global coverage probability and a global expected content, which are averages with respect to the prior distribution. The procedure defined by the performance measures will be called a Bayes MIE with thresholding (BMIE Thres). In this study, a normal-normal model is utilized to build up the BMIE Thres for a set of location parameters. Then, the behaviors of BMIE Thres are investigated in terms of the performance measures, which approach those of the corresponding z-based MIE as the thresholding parameter, C, goes to infinity. In addition, an optimization procedure is introduced to achieve the best thresholding parameter C. For illustrations, in-season baseball batting average data and leukemia gene expression data are used to demonstrate the procedure for the known and unknown standard deviations situations, respectively. In the ensuing simulations, the target parameters are generated from different true generating distributions to consider the misspecified prior situation. The simulation also involves Bayes credible MIEs, and the effectiveness among the different MIEs are compared with respect to the performance measures. In general, the thresholding procedure helps to achieve a meaningful reduction in the global expected content while maintaining a nominal level of the global coverage probability.

stat.ME

The Search for Truth through Data: NP Decision Processes, ROC Functions, $P$-Functionals, Knowledge Updating and Sequential Learning

This paper re-visits the problem of deciding between two simple hypotheses, the setting considered by Neyman and Pearson in developing their fundamental lemma. It studies the decision process induced by the most powerful test and the receiver operating characteristic function associated with this decision process. It addresses the question of how to report the decision arising from the decision function. It also examines the P-functional (the P-value statistic) and its role in the decision-making process. The impetus of this work is the continuing criticisms of statistical decision-making procedures that use the P-functional and a level of significance (LoS) of 0.05. A point made is that if one is going to use the value of the P-functional, then it should be used in an equivalent manner as the most powerful decision function, but if one wants to obtain from its value the degree of support for either hypotheses, then the value of its density under the alternative is the proper quantity to use. Replicability of results are discussed. Knowledge updating through Bayes theorem when given the decision or the value of the P-functional is also discussed, and it is argued that sequential learning is a coherent way of finding the truth. But the impact of publication bias is also demonstrated to be quite serious in the search for truth. It is argued that decision-makers are free to choose their own LoS, since the additional summary measures will automatically take their LoS choices into consideration. Three approaches for choosing an optimal LoS are discussed and a procedure for sample size determination is described. Ideas are illustrated by concrete problems and by the lady tea-tasting experiment of Fisher which ushered null hypothesis significance testing. It is hoped that by considering this fundamental setting of simple hypotheses, a better understanding of more complex settings will ensue.

math.ST

Median Confidence Regions in a Nonparametric Model

The problem of constructing confidence regions for the median in the nonparametric measurement error model (NMEM) is considered. This problem arises in many settings, including inference about the median lifetime of a complex system arising in engineering, reliability, biomedical, and public health settings. Current methods of constructing CRs are discussed, including the T-statistic based CR and the Wilcoxon signed-rank statistic based CR, arguably the two default methods in applied work when a confidence interval about the center of a distribution is desired. Optimal equivariant CRs are developed with focus on subclasses of of the class of all distributions. Applications to a real car mileage efficiency data set and Proschan's air-conditioning data set are demonstrated. Simulation studies to compare the performances of the different CR methods were undertaken. Results of these studies indicate that the sign-statistic based CR and the optimal CR focused on symmetric distributions satisfy the confidence level requirement, though they tended to have higher contents; while two of the bootstrap-based CR procedures and one of the developed adaptive CR tended to be a tad more liberal but with smaller contents. A critical recommendation is that, under the NMEM, both the T-statistic based and Wilcoxon signed-rank statistic based confidence regions should not be used since they have degraded confidence levels and/or inflated contents.

math.ST

Model Selection and Estimation with Quantal-Response Data in Benchmark Risk Assessment

This paper describes several approaches for estimating the benchmark dose (BMD) in a risk assessment study with quantal dose-response data and when there are competing model classes for the dose-response function. Strategies involving a two-step approach, a model-averaging approach, a focused-inference approach, and a nonparametric approach based on a PAVA-based estimator of the dose-response function are described and compared. Attention is raised to the perils involved in data "double-dipping" and the need to adjust for the model-selection stage in the estimation procedure. Simulation results are presented comparing the performance of five model selectors and eight BMD estimators. An illustration using a real quantal-response data set from a carcinogenecity study is provided.

math.ST

Asymptotics for a Class of Dynamic Recurrent Event Models

Asymptotic properties, both consistency and weak convergence, of estimators arising in a general class of dynamic recurrent event models are presented. The class of models take into account the impact of interventions after each event occurrence, the impact of accumulating event occurrences, the induced informative and dependent right-censoring mechanism due to the data-accrual scheme, and the effect of covariate processes on the recurrent event occurrences. The class of models subsumes as special cases many of the recurrent event models that have been considered in biostatistics, reliability, and in the social sciences. The asymptotic properties presented have the potential of being useful in developing goodness-of-fit and model validation procedures, confidence intervals and confidence bands constructions, and hypothesis testing procedures for the finite- and infinite-dimensional parameters of a general class of dynamic recurrent event models, albeit the models without frailties.

math.ST

Compound p-Value Statistics for Multiple Testing Procedures

Many multiple testing procedures make use of the p-values from the individual pairs of hypothesis tests, and are valid if the p-value statistics are independent and uniformly distributed under the null hypotheses. However, it has recently been shown that these types of multiple testing procedures are inefficient since such p-values do not depend upon all of the available data. This paper provides tools for constructing compound p-value statistics, which are those that depend upon all of the available data, but still satisfy the conditions of independence and uniformity under the null hypotheses. As an example, a class of compound p-value statistics for testing for location shifts is developed. It is demonstrated, both analytically and through simulations, that multiple testing procedures tend to reject more false null hypotheses when applied to these compound p-values rather than the usual p-values, and at the same time still guarantee the desired type I error rate control. The compound p-values, in conjunction with two different multiple testing methods, are used to analyze a real microarray data set. Applying either multiple testing method to the compound p-values, instead of the usual p-values, enhances their powers.

stat.ME

Classes of Multiple Decision Functions Strongly Controlling FWER and FDR

This paper provides two general classes of multiple decision functions where each member of the first class strongly controls the family-wise error rate (FWER), while each member of the second class strongly controls the false discovery rate (FDR). These classes offer the possibility that an optimal multiple decision function with respect to a pre-specified criterion, such as the missed discovery rate (MDR), could be found within these classes. Such multiple decision functions can be utilized in multiple testing, specifically, but not limited to, the analysis of high-dimensional microarray data sets.

math.ST