SearcharxivSearch

arXiv subjects

Mark Rubin

Publications and source records attributed to Mark Rubin.

8 recordsLinked to original sources

p-Hacking Inflates Type I Error Rates in the Error Statistical Approach but not in the Formal Inference Approach

p-hacking occurs when researchers conduct multiple significance tests (e.g., p1;H0,1 and p2;H0,2) and then selectively report tests that yield desirable (usually significant) results (e.g., p2 < 0.05;H0,2) without correcting for multiple testing (e.g., 0.05/2 = 0.025). In the present article, I consider p-hacking in the context of two philosophies of significance testing - the error statistical approach and the formal inference approach. I argue that although p-hacking inflates Type I error rates in the error statistical approach, it does not inflate them in the formal inference approach. Specifically, in the error statistical approach, the "actual" familywise error rate (e.g., 1 - [1 - 0.05]2 = 0.098 for two independent tests) is relevant because it covers both the reported and unreported tests in the "actual" test procedure (i.e., p1;H0,1 and p2;H0,2). In this approach, Type I error rate inflation occurs because the "actual" error rate (0.098) is higher than the nominal error rate (0.05). In contrast, in the formal inference approach, the "actual" familywise error rate is irrelevant because (a) the researcher does not report a statistical inference about the corresponding intersection null hypothesis (i.e., H0,1 & H0,2), and (b) the "actual" familywise error rate does not license inferences about the reported individual hypotheses (i.e., H0,2). Instead, in the formal inference approach, only the nominal error rate is relevant, and a comparison with the "actual" error rate is inappropriate. Implications for conceptualizing, demonstrating, and reducing p-hacking are discussed.

stat.OT

Preregistration does not improve the transparent evaluation of severity in Popper's philosophy of science or when deviations are allowed

One justification for preregistering research hypotheses, methods, and analyses is that it improves the transparent evaluation of the severity of hypothesis tests. In this article, I consider two cases in which preregistration does not improve this evaluation. First, I argue that, although preregistration may facilitate the transparent evaluation of severity in Mayo's error statistical philosophy of science, it does not facilitate this evaluation in Popper's theory-centric approach. To illustrate, I show that associated concerns about Type I error rate inflation are only relevant in the error statistical approach and not in a theory-centric approach. Second, I argue that a test procedure that is preregistered but that also allows deviations in its implementation (i.e., "a plan, not a prison") does not provide a more transparent evaluation of Mayoian severity than a non-preregistered procedure. In particular, I argue that sample-based validity-enhancing deviations cause an unknown inflation of the test procedure's Type I error rate and, consequently, an unknown reduction in its capability to license inferences severely. I conclude that preregistration does not improve the transparent evaluation of severity (a) in Popper's philosophy of science or (b) in Mayo's approach when deviations are allowed.

stat.ME

Inconistent multiple testing corrections: The fallacy of using family-based error rates to make inferences about individual hypotheses

During multiple testing, researchers often adjust their alpha level to control the familywise error rate for a statistical inference about a joint union alternative hypothesis (e.g., "H1,1 or H1,2"). However, in some cases, they do not make this inference. Instead, they make separate inferences about each of the individual hypotheses that comprise the joint hypothesis (e.g., H1,1 and H1,2). For example, a researcher might use a Bonferroni correction to adjust their alpha level from the conventional level of 0.050 to 0.025 when testing H1,1 and H1,2, find a significant result for H1,1 (p < 0.025) and not for H1,2 (p > .0.025), and so claim support for H1,1 and not for H1,2. However, these separate individual inferences do not require an alpha adjustment. Only a statistical inference about the union alternative hypothesis "H1,1 or H1,2" requires an alpha adjustment because it is based on "at least one" significant result among the two tests, and so it refers to the familywise error rate. Hence, an inconsistent correction occurs when a researcher corrects their alpha level during multiple testing but does not make an inference about a union alternative hypothesis. In the present article, I discuss this inconsistent correction problem, including its reduction in statistical power for tests of individual hypotheses and its potential causes vis-a-vis error rate confusions and the alpha adjustment ritual. I also provide three illustrations of inconsistent corrections from recent psychology studies. I conclude that inconsistent corrections represent a symptom of statisticism, and I call for a more nuanced inference-based approach to multiple testing corrections.

stat.ME

Type I Error Rates are Not Usually Inflated

The inflation of Type I error rates is thought to be one of the causes of the replication crisis. Questionable research practices such as p-hacking are thought to inflate Type I error rates above their nominal level, leading to unexpectedly high levels of false positives in the literature and, consequently, unexpectedly low replication rates. In this article, I offer an alternative view. I argue that questionable and other research practices do not usually inflate relevant Type I error rates. I begin by introducing the concept of Type I error rates and distinguishing between statistical errors and theoretical errors. I then illustrate my argument with respect to model misspecification, multiple testing, selective inference, forking paths, exploratory analyses, p-hacking, optional stopping, double dipping, and HARKing. In each case, I demonstrate that relevant Type I error rates are not usually inflated above their nominal level, and in the rare cases that they are, the inflation is easily identified and resolved. I conclude that the replication crisis may be explained, at least in part, by researchers' misinterpretation of statistical errors and their underestimation of theoretical errors.

stat.ME

When to adjust alpha during multiple testing: A consideration of disjunction, conjunction, and individual testing

Scientists often adjust their significance threshold (alpha level) during null hypothesis significance testing in order to take into account multiple testing and multiple comparisons. This alpha adjustment has become particularly relevant in the context of the replication crisis in science. The present article considers the conditions in which this alpha adjustment is appropriate and the conditions in which it is inappropriate. A distinction is drawn between three types of multiple testing: disjunction testing, conjunction testing, and individual testing. It is argued that alpha adjustment is only appropriate in the case of disjunction testing, in which at least one test result must be significant in order to reject the associated joint null hypothesis. Alpha adjustment is inappropriate in the case of conjunction testing, in which all relevant results must be significant in order to reject the joint null hypothesis. Alpha adjustment is also inappropriate in the case of individual testing, in which each individual result must be significant in order to reject each associated individual null hypothesis. The conditions under which each of these three types of multiple testing is warranted are examined. It is concluded that researchers should not automatically (mindlessly) assume that alpha adjustment is necessary during multiple testing. Illustrations are provided in relation to joint studywise hypotheses and joint multiway ANOVAwise hypotheses.

stat.ME

Does preregistration improve the credibility of research findings?

Preregistration entails researchers registering their planned research hypotheses, methods, and analyses in a time-stamped document before they undertake their data collection and analyses. This document is then made available with the published research report to allow readers to identify discrepancies between what the researchers originally planned to do and what they actually ended up doing. This historical transparency is supposed to facilitate judgments about the credibility of the research findings. The present article provides a critical review of 17 of the reasons behind this argument. The article covers issues such as HARKing, multiple testing, p-hacking, forking paths, optional stopping, researchers' biases, selective reporting, test severity, publication bias, and replication rates. It is concluded that preregistration's historical transparency does not facilitate judgments about the credibility of research findings when researchers provide contemporary transparency in the form of (a) clear rationales for current hypotheses and analytical approaches, (b) public access to research data, materials, and code, and (c) demonstrations of the robustness of research conclusions to alternative interpretations and analytical approaches.

stat.OT

Spitzer Observations of Bok Globule B335: Isolated Star Formation Efficiency and Cloud Structure

We present infrared and millimeter observations of Barnard 335, the prototypical isolated Bok globule with an embedded protostar. Using Spitzer data we measure the source luminosity accurately; we also constrain the density profile of the innermost globule material near the protostar using the observation of an 8.0 um shadow. HHT observations of 12CO 2 --> 1 confirm the detection of a flattened molecular core with diameter ~10000 AU and the same orientation as the circumstellar disk (~100 to 200 AU in diameter). This structure is probably the same as that generating the 8.0 um shadow and is expected from theoretical simulations of collapsing embedded protostars. We estimate the mass of the protostar to be only ~5% of the mass of the parent globule.

astro-ph

Spitzer observations of a 24 micron shadow: Bok Globule CB190

We present Spitzer observations of the dark globule CB190 (L771). We observe a roughly circular 24 micron shadow with a 70 arcsec radius. The extinction profile of this shadow matches the profile derived from 2MASS photometry at the outer edges of the globule and reaches a maximum of ~32 visual magnitudes at the center. The corresponding mass of CB190 is ~10 Msun. Our 12CO and 13CO J = 2-1 data over a 10 arcmin X 10 arcmin region centered on the shadow show a temperature ~10 K. The thermal continuum indicates a similar temperature for the dust. The molecular data also show evidence of freezeout onto dust grains. We estimate a distance to CB190 of 400 pc using the spectroscopic parallax of a star associated with the globule. Bonnor-Ebert fits to the density profile, in conjunction with this distance, yield xi_max = 7.2, indicating that CB190 may be unstable. The high temperature (56 K) of the best fit Bonnor-Ebert model is in contradiction with the CO and thermal continuum data, leading to the conclusion that the thermal pressure is not enough to prevent free-fall collapse. We also find that the turbulence in the cloud is inadequate to support it. However, the cloud may be supported by the magnetic field, if this field is at the average level for dark globules. Since the magnetic field will eventually leak out through ambipolar diffusion, it is likely that CB190 is collapsing or in a late pre-collapse stage.

astro-ph