SearcharxivSearch

arXiv subjects

Zad Rafi

Publications and source records attributed to Zad Rafi.

4 recordsLinked to original sources

Technical Issues in the Interpretation of S-values and Their Relation to Other Information Measures

An extended technical discussion of $S$-values and unconditional information can be found in Greenland, 2019. Here we briefly cover several technical topics mentioned in our main paper, Rafi & Greenland, 2020: Different units for (scaling of) the $S$-value besides base-2 logs (bits); the importance of uniformity (validity) of the $P$-value for interpretation of the $S$-value; and the relation of the $S$-value to other measures of statistical information about a test hypothesis or model.

stat.ME

To Aid Scientific Inference, Emphasize Unconditional Compatibility Descriptions of Statistics

All scientific interpretations of statistical outputs depend on background (auxiliary) assumptions that are rarely delineated or explicitly interrogated. These include not only the usual modeling assumptions, but also deeper assumptions about the data-generating mechanism that are implicit in conventional statistical interpretations yet are unrealistic in most health, medical and social research. We provide arguments and methods for reinterpreting statistics such as P-values and interval estimates in unconditional terms, which describe compatibility of observations with an entire set of underlying assumptions, rather than with a narrow target hypothesis conditional on the assumptions. Emphasizing unconditional interpretations helps avoid overconfident and misleading inferences in light of uncertainties about the assumptions used to arrive at the statistical results. These include not only mathematical assumptions, but also those about absence of systematic errors, protocol violations, and data corruption. Unconditional descriptions introduce assumption uncertainty directly into the primary statistical interpretations of results, rather than leaving it for the discussion of limitations after presentation of conditional interpretations. The unconditional approach does not entail different methods or calculations, only different interpretation of the usual results. We view use of unconditional description as a vital component of effective statistical training and presentation. By interpreting statistical outputs in unconditional terms, researchers can avoid making overconfident statements based on statistical outputs. Instead, reports should emphasize the compatibility of results with a range of plausible explanations, including assumption violations.

stat.ME

Semantic and Cognitive Tools to Aid Statistical Science: Replace Confidence and Significance by Compatibility and Surprise

Researchers often misinterpret and misrepresent statistical outputs. This abuse has led to a large literature on modification or replacement of testing thresholds and $P$-values with confidence intervals, Bayes factors, and other devices. Because the core problems appear cognitive rather than statistical, we review simple aids to statistical interpretations. These aids emphasize logical and information concepts over probability, and thus may be more robust to common misinterpretations than are traditional descriptions. We use the Shannon transform of the $P$-value $p$, also known as the binary surprisal or $S$-value $s=-\log_{2}(p)$, to measure the information supplied by the testing procedure, and to help calibrate intuitions against simple physical experiments like coin tossing. We also use tables or graphs of test statistics for alternative hypotheses, and interval estimates for different percentile levels, to thwart fallacies arising from arbitrary dichotomies. Finally, we reinterpret $P$-values and interval estimates in unconditional terms, which describe compatibility of data with the entire set of analysis assumptions. We illustrate these methods with a reanalysis of data from an existing record-based cohort study. In line with other recent recommendations, we advise that teaching materials and research reports discuss $P$-values as measures of compatibility rather than significance, compute $P$-values for alternative hypotheses whenever they are computed for null hypotheses, and interpret interval estimates as showing values of high compatibility with data, rather than regions of confidence. Our recommendations emphasize cognitive devices for displaying the compatibility of the observed data with various hypotheses of interest, rather than focusing on single hypothesis tests or interval estimates. We believe these simple reforms are well worth the minor effort they require.

stat.ME

Misplaced Confidence in Observed Power

A recently published randomized controlled trial in JAMA investigated the impact of the selective serotonin reuptake inhibitor, escitalopram, on the risk of major adverse events (MACE). The authors estimated a hazard ratio (HR) of 0.69 (95% CI: 0.49, 0.96; $p$ = 0.03) and then attempted to calculate how much statistical power their study (test) had attained, and used this measure to assess how reliable their results were. Here, we discuss why this approach, along with other post-hoc power analyses, are highly misleading.

stat.AP