SearcharxivSearch

arXiv subjects

Pamela Shaw

Publications and source records attributed to Pamela Shaw.

4 recordsLinked to original sources

De-meaning Simulation Studies

In simulation studies evaluating asymptotic approximations it is common practice to report averages and standard deviations over repeated simulations. We argue that quantile-based summaries are more appropriate from both a theoretical and practical point of view. Theoretically, convergence of moments -- or even existence of moments -- is not guaranteed by convergence in distribution, so sample moments are not ideal for assessing the accuracy of a distributional approximation. In practice, means and variances are not good summaries of approximately-Normal distributions that may have occasional outliers. We suggest the median and median absolute deviation, and empirical confidence interval coverage, as better general summaries, and argue that moments should be reserved for simulation settings where they are of substantive interest.

stat.ME

Missing at Random or Not: A Semiparametric Testing Approach

Practical problems with missing data are common, and statistical methods have been developed concerning the validity and/or efficiency of statistical procedures. On a central focus, there have been longstanding interests on the mechanism governing data missingness, and correctly deciding the appropriate mechanism is crucially relevant for conducting proper practical investigations. The conventional notions include the three common potential classes -- missing completely at random, missing at random, and missing not at random. In this paper, we present a new hypothesis testing approach for deciding between missing at random and missing not at random. Since the potential alternatives of missing at random are broad, we focus our investigation on a general class of models with instrumental variables for data missing not at random. Our setting is broadly applicable, thanks to that the model concerning the missing data is nonparametric, requiring no explicit model specification for the data missingness. The foundational idea is to develop appropriate discrepancy measures between estimators whose properties significantly differ only when missing at random does not hold. We show that our new hypothesis testing approach achieves an objective data oriented choice between missing at random or not. We demonstrate the feasibility, validity, and efficacy of the new test by theoretical analysis, simulation studies, and a real data analysis.

stat.ME

Novel Non-Negative Variance Estimator for (Modified) Within-Cluster Resampling

This article proposes a novel variance estimator for within-cluster resampling (WCR) and modified within-cluster resampling (MWCR) - two existing methods for analyzing longitudinal data. WCR is a simple but computationally intensive method, in which a single observation is randomly sampled from each cluster to form a new dataset. This process is repeated numerous times, and in each resampled dataset (or outputation), we calculate beta using a generalized linear model. The final resulting estimator is an average across estimates from all outputations. MWCR is an extension of WCR that can account for the within-cluster correlation of the dataset; consequently, there are two noteworthy differences: 1) in MWCR, each resampled dataset is formed by randomly sampling multiple observations without replacement from each cluster and 2) generalized estimating equations (GEEs) are used to estimate the parameter of interest. While WCR and MWCR are relatively simple to implement, a key challenge is that the proposed moment-based estimator is often times negative in practice. Our modified variance estimator is not only strictly positive, but simulations show that it preserves the type I error and allows statistical power gains associated with MWCR to be realized.

stat.ME

Regression calibration to correct correlated errors in outcome and exposure

Measurement error arises through a variety of mechanisms. A rich literature exists on the bias introduced by covariate measurement error and on methods of analysis to address this bias. By comparison, less attention has been given to errors in outcome assessment and non-classical covariate measurement error. We consider an extension of the regression calibration method to settings with errors in a continuous outcome, where the errors may be correlated with prognostic covariates or with covariate measurement error. This method adjusts for the measurement error in the data and can be applied with either a validation subset, on which the true data are also observed (e.g., a study audit), or a reliability subset, where a second observation of error prone measurements are available. For each case, we provide conditions under which the proposed method is identifiable and leads to unbiased estimates of the regression parameter. When the second measurement on the reliability subset has no error or classical unbiased measurement error, the proposed method is unbiased even when the primary outcome and exposures of interest are subject to both systematic and random error. We examine the performance of the method with simulations for a variety of measurement error scenarios and sizes of the reliability subset. We illustrate the method's application using data from the Women's Health Initiative Dietary Modification Trial.

stat.ME