Searcharxiv⌕ Search

arXiv subjects

Wolfgang Rolke

Publications and source records attributed to Wolfgang Rolke.

10 recordsLinked to original sources

Power Studies For Two-Sample and Goodness-of-Fit Methods For Multivariate Data

We present the results of a large number of simulation studies regarding the power of various goodness-of-fit as well as non-parametric two-sample tests for multivariate data. In two dimensions this includes both continuous and discrete data, in higher dimensions continuous data only. In general no single method can be relied upon to provide good power, any one method may be quite good for some combination of null hypothesis and alternative and may fail badly for another. Based on the results of these studies we propose a fairly small number of methods chosen such that for any of the case studies included here at least one of the methods has good power. The studies were carried out using the R packages MD2sample and MDgof, available from CRAN.

stat.ME↗

Power Studies For Two-sample Methods For Multivariate Data

We present the results of a large number of simulation studies regarding the power of various non-parametric two-sample tests for multivariate data. This includes both continuous and discrete data. In general no single method can be relied upon to provide good power, any one method may be quite good for some combination of null hypothesis and alternative and may fail badly for another. Based on the results of these studies we propose a fairly small number of methods chosen such that for any of the case studies included here at least one of the methods has good power. The studies were carried out using the R package MD2sample, available from CRAN.

stat.ME↗

Estimating Counts Through an Average Rounded to the Nearest Non-negative Integer and its Theoretical & Practical Effects

In practice, the use of rounding is ubiquitous. Although researchers have looked at the implications of rounding continuous random variables, rounding may also be applied to functions of discrete random variables. For example, to infer the number of excess deaths due to falls after a national emergency, authorities may only provide a rounded average of deaths before and after the emergency started. Deaths from falling tend to be relatively low in most places, and such rounding may seriously affect inference on the change in the rate of deaths. In this paper, we study drawing inference on a parameter fromthe probability mass function of a non-negative discrete random variableY , when for rounding coarsening width h we get U = h[Y /h] as a proxy forY . We show that the probability generating function of U, E(U), and Var(U) capture the effect of the coarsening of the support of Y . Theoretical properties are explored further under some probability distributions. Moreover, we introduce two relative risks of rounding metrics to aid the numerical assessment of how sensitive the results may be to rounding. Under certain conditions, rounding has little impact. However, we also find scenarios where rounding can significantly affect statistical inference. The methods are applied to infer the probability of success of a binomial distribution and estimate the excess deaths due to Hurricane Maria. The simple methods we propose can partially counter rounding error effects.

math.ST↗

Simulation Studies For Goodness-of-Fit and Two-Sample Methods For Univariate Data

We present the results of a large number of simulation studies regarding the power of various goodness-of-fit as well as nonparametric two-sample tests for univariate data. This includes both continuous and discrete data. In general no single method can be relied upon to provide good power, any one method may be quite good for some combination of null hypothesis and alternative and may fail badly for another. Based on the results of these studies we propose a fairly small number of methods chosen such that for any of the case studies included here at least one of the methods has good power. The studies were carried out using the R packages R2sample and Rgof, available from CRAN.

stat.ME↗

R Package moodlequizR: Fully Randomized Moodle Tests

This article describes the R package moodlequizR, which allows the user to easily create fully randomized quizzes and exams for Moodle, or indeed any online assessment platform that uses XML files for importing questions. In such a quiz the students are presented with the essentially same problem, but with various parts sufficiently different to make cheating very difficult. For example, the problem might require the students to find the sample mean but each student is presented with a different data set. Moodle does include some facilities for randomization, but these are rudimentary and wholly insufficient for a course in Statistics. The package is available on CRAN.

stat.AP↗

Supplemental Studies for Simultaneous Goodness-of-Fit Testing

Testing to see whether a given data set comes from some specified distribution is among the oldest types of problems in Statistics. Many such tests have been developed and their performance studied. The general result has been that while a certain test might perform well, aka have good power, in one situation it will fail badly in others. This is not a surprise given the great many ways in which a distribution can differ from the one specified in the null hypothesis. It is therefore very difficult to decide a priori which test to use. The obvious solution is not to rely on any one test but to run several of them. This however leads to the problem of simultaneous inference, that is, if several tests are done even if the null hypothesis were true, one of them is likely to reject it anyway just by random chance. In this paper we present a method that yields a p value that is uniform under the null hypothesis no matter how many tests are run. This is achieved by adjusting the p value via simulation. While this adjustment method is not new, it has not previously been used in the context of goodness-of-fit testing. We present a number of simulation studies that show the uniformity of the p value and others that show that this test is superior to any one test if the power is averaged over a large number of cases.

stat.AP↗

A Chi-square Goodness-of-Fit Test for Continuous Distributions against a known Alternative

The chi square goodness-of-fit test is among the oldest known statistical tests, first proposed by Pearson in 1900 for the multinomial distribution. It has been in use in many fields ever since. However, various studies have shown that when applied to data from a continuous distribution it is generally inferior to other methods such as the Kolmogorov-Smirnov or Anderson-Darling tests. However, the performance, that is the power, of the chi square test depends crucially on the way the data is binned. In this paper we describe a method that automatically finds a binning that is very good against a specific alternative. We show that then the chi square test is generally competitive and sometimes even superior to other standard tests.

stat.ME↗

Modeling Excess Deaths After a Natural Disaster with Application to Hurricane Maria

Estimation of excess deaths due to a natural disaster is an important public health problem. The CDC provides guidelines to fill death certificates to help determine the death toll of such events. But, even when followed by medical examiners, the guidelines can not guarantee a precise calculation of excess deaths.%particularly due to the ambiguity of indirect deaths. We propose two models to estimate excess deaths due to an emergency. The first model is simple, permitting excess death estimation with little data through a profile likelihood method. The second model is more flexible, incorporating: temporal variation, covariates, and possible population displacement; while allowing inference on how the emergency's effect changes with time. The models are implemented to build confidence intervals estimating Hurricane Maria's death toll.

stat.AP↗

How to Claim a Discovery

We describe a statistical hypothesis test for the presence of a signal. The test allows the researcher to fix the signal location and/or width a priori, or perform a search to find the signal region that maximizes the signal. The background rate and/or distribution can be known or might be estimated from the data. Cuts can be used to bring out the signal.

physics.data-an↗