SearcharxivSearch

arXiv subjects

S. Stanley Young

Publications and source records attributed to S. Stanley Young.

At least 19 recordsLinked to original sources

The reliability of the gender Implicit Association Test (gIAT) for high-ability careers

Males outnumber females in many high-ability careers in the fields of science, technology, engineering, and mathematics, STEM, and academic medicine, to name a few. These differences are often attributed to subconscious bias as measured by the gender Implicit Association Test, gIAT. We compute p-value plots for results from two meta-analyses, one examines the predictive power of gIAT, and the other examines the predictive power of vocational interests, i.e. personal interests, and behaviors, for explaining gender differences in high-ability careers. The results are clear, the gender Implicit Association Test provides little or no information on male versus female differences, whereas vocational interests are strongly predictive. Researchers of implicit bias should expand their modeling to include additional relevant covariates. In short, these meta-analyses provide no support for the gender Implicit Association Test influencing choice and gender differences of high-ability careers.

stat.AP

Reproducibility of Implicit Association Test (IAT) -- Case study of meta-analysis of racial bias research claims

The Implicit Association Test, IAT, is widely used to measure hidden (subconscious) human biases, implicit bias, of many topics: race, gender, age, ethnicity, religion stereotypes. There is a need to understand the reliability of these measures as they are being used in many decisions in society today. A case study was undertaken to independently test the reliability of (ability to reproduce) racial bias research claims of Black White relations based on IAT (implicit bias) and explicit bias measurements using statistical p value plots. These claims were for IAT, real world behavior correlations and explicit bias, real world behavior correlations of Black White relations. The p value plots were constructed using data sets from published literature and the plots exhibited considerable randomness for all correlations examined. This randomness supports a lack of correlation between IAT, implicit bias, and explicit bias measurements with real world behaviors of Whites towards Blacks. These findings were for microbehaviors (measures of nonverbal and subtle verbal behavior) and person perception judgments (explicit judgments about others). Findings of the p value plots were consistent with the case study research claim that the IAT provides little insight into who will discriminate against whom. It was also observed that the amount of real world variance explained by the IAT and explicit bias measurements was small, less than 5 percent. Others have noted that the poor performance of both the IAT and explicit bias measurements are mostly consistent with a (flawed instruments explanation) problems in theories that motivated development and use of these instruments.

stat.AP

Statistical reliability of meta_analysis research claims for gas stove cooking_childhood respiratory health associations

Odds ratios or p_values from individual observational studies can be combined to examine a common cause_effect research question in meta_analysis. However, reliability of individual studies used in meta_analysis should not be taken for granted as claimed cause_effect associations may not reproduce. An evaluation was undertaken on meta_analysis of base papers examining gas stove cooking, including nitrogen dioxide, NO2, and childhood asthma and wheeze associations. Numbers of hypotheses tested in 14 of 27 base papers, 52 percent, used in meta_analysis of asthma and wheeze were counted. Test statistics used in the meta_analysis, 40 odds ratios with 95 percent confidence limits, were converted to p_values and presented in p_value plots. The median and interquartile range of possible numbers of hypotheses tested in the 14 base papers was 15,360, 6,336_49,152. None of the 14 base papers made mention of correcting for multiple testing, nor was any explanation offered if no multiple testing procedure was used. Given large numbers of hypotheses available, statistics drawn from base papers and used for meta-analysis are likely biased. Even so, p-value plots for gas stove_current asthma and gas stove_current wheeze associations show randomness consistent with unproven gas stove harms. The meta-analysis fails to provide reliable evidence for public health policy making on gas stove harms to children in North America. NO2 is not established as a biologically plausible explanation of a causal link with childhood asthma. Biases_multiple testing and p-hacking_cannot be ruled out as explanations for a gas stove_current asthma association claim. Selective reporting is another bias in published literature of gas stove_childhood respiratory health studies. Keywords gas stove, asthma, meta-analysis, p-value plot, multiple testing, p_hacking

stat.AP

Mortality Rates of US Counties: Are they Reliable and Predictable?

We examine US County-level observational data on Lung Cancer mortality rates in 2012 and overall Circulatory Respiratory mortality rates in 2016 as well as their "Top Ten" potential causes from Federal or State sources. We find that these two mortality rates for 2,812 US Counties have remarkably little in common. Thus, for predictive modeling, we use a single "compromise" measure of mortality that has several advantages. The vast majority of our new findings have simple implications that we illustrate graphically.

stat.AP

Statistical reproducibility of meta-analysis research claims for medical mask use in community settings to prevent COVID infection

The coronavirus pandemic (COVID) has been an exceptional test of current scientific evidence that inform and shape policy. Many US states, cities, and counties implemented public orders for mask use on the notion that this intervention would delay and flatten the epidemic peak and largely benefit public health outcomes. P-value plotting was used to evaluate statistical reproducibility of meta-analysis research claims of a benefit for medical (surgical) mask use in community settings to prevent COVID infection. Eight studies (seven meta-analyses, one systematic review) published between 1 January 2020 and 7 December 2022 were evaluated. Base studies were randomized control trials with outcomes of medical diagnosis or laboratory-confirmed diagnosis of viral (Influenza or COVID) illness. Self-reported viral illness outcomes were excluded because of awareness bias. No evidence was observed for a medical mask use benefit to prevent viral infections in six p-value plots (five meta-analyses and one systematic review). Research claims of no benefit in three meta-analyses and the systematic review were reproduced in p-value plots. Research claims of a benefit in two meta-analyses were not reproduced in p-value plots. Insufficient data were available to construct p-value plots for two meta-analyses because of overreliance on self-reported outcomes. These findings suggest a benefit for medical mask use in community settings to prevent viral, including COVID infection, is unproven.

q-bio.QM

Reproducibility of health claims in meta-analysis studies of COVID quarantine (stay-at-home) orders

The coronavirus pandemic (COVID) has been an extraordinary test of modern government scientific procedures that inform and shape policy. Many governments implemented COVID quarantine (stay-at-home) orders on the notion that this nonpharmaceutical intervention would delay and flatten the epidemic peak and largely benefit public health outcomes. The overall research capacity response to COVID since late 2019 has been massive. Given lack of research transparency, only a small fraction of published research has been judged by others to be reproducible before COVID. Independent evaluation of published meta-analysis on a common research question can be used to assess the reproducibility of a claim coming from that field of research. We used a p-value plotting statistical method to independently evaluate reproducibility of specific research claims made in four meta-analysis studies related to benefits/risks of COVID quarantine orders. Outcomes we investigated included: mortality, mental health symptoms, incidence of domestic violence, and suicidal ideation (thoughts of killing yourself). Three of the four meta-analyses that we evaluated (mortality, mental health symptoms, incidence of domestic violence) raise further questions about benefits/risks of this form of intervention. The fourth meta-analysis study (suicidal ideation) is unreliable. Given lack of research transparency and irreproducibility of published research, independent evaluation of meta-analysis studies using p-value plotting is offered as a way to strengthen or refute (falsify) claims made in COVID research.

q-bio.OT

EPA Particulate Matter Data -- Analyses using Local Control Strategy

Statistical Learning methodology for analysis of large collections of cross-sectional observational data can be most effective when the approach used is both Nonparametric and Unsupervised. We illustrate use of our NU Learning approach on 2016 US environmental epidemiology data that we have made freely available. We encourage other researchers to download these data, apply whatever methodology they wish, and contribute to development of a broad-based ``consensus view'' of potential effects of Secondary Organic Aerosols (volatile organic compounds of predominantly biogenic or anthropogenic origin) within PM2.5 particulate matter on circulatory and/or respiratory mortality. Our analyses here focus on the question: ``Are regions with relatively high air-borne biogenic particulate matter also expected to have relatively high circulatory and/or respiratory mortality?''

cs.CY

Case Study: Evaluation of a meta-analysis of the association between soy protein and cardiovascular disease

It is well-known that claims coming from observational studies most often fail to replicate. Experimental (randomized) trials, where conditions are under researcher control, have a high reputation and meta-analysis of experimental trials are considered the best possible evidence. Given the irreproducibility crisis, experiments lately are starting to be questioned. There is a need to know the reliability of claims coming from randomized trials. A case study is presented here independently examining a published meta-analysis of randomized trials claiming that soy protein intake improves cardiovascular health. Counting and p-value plotting techniques (standard p-value plot, p-value expectation plot, and volcano plot) are used. Counting (search space) analysis indicates that reported p-values from the meta-analysis could be biased low due to multiple testing and multiple modeling. Plotting techniques used to visualize the behavior of the data set used for meta-analysis suggest that statistics drawn from the base papers do not satisfy key assumptions of a random-effects meta-analysis. These assumptions include using unbiased statistics all drawn from the same population. Also, publication bias is unaddressed in the meta-analysis. The claim that soy protein intake should improve cardiovascular health is not supported by our analysis.

stat.AP

Evaluation of a meta-analysis of the association between red and processed meat and selected human health effects

Background: Risk ratios or p-values from multiple, independent studies, observational or randomized, can be computationally combined to provide an overall assessment of a research question in meta-analysis. However, an irreproducibility crisis currently afflicts a wide range of scientific disciplines, including nutritional epidemiology. An evaluation was undertaken to assess the reliability of a meta-analysis examining the association between red and processed meat and selected human health effects (all-cause mortality, cardiovascular mortality, overall cancer mortality, breast cancer incidence, colorectal cancer incidence, type 2 diabetes incidence). Methods: The number of statistical tests and models were counted in 15 randomly selected base papers (14%) from 105 used in the meta-analysis. Relative risk with 95% confidence limits for 125 risk results were converted to p-values and p-value plots were constructed to evaluate the effect heterogeneity of the p-values. Results: The number of statistical tests possible in the 15 randomly selected base papers was large, median = 20,736 (interquartile range = 1,728 to 331,776). Each p-value plot for the six selected health effects showed either a random pattern (p-values > 0.05), or a two-component mixture with small p-values < 0.001 while other p-values appeared random. Given potentially large numbers of statistical tests conducted in the 15 selected base papers, questionable research practices cannot be ruled out as explanations for small p-values. Conclusions: This independent analysis, which complements the findings of the original meta-analysis, finds that the base papers used in the red and resulting processed meat meta-analysis do not provide evidence for the claimed health effects.

stat.AP

Standard meta-analysis methods are not robust

P values or risk ratios from multiple, independent studies, observational or randomized, can be computationally combined to provide an overall assessment of a research question in meta-analysis. There is a need to examine the reliability of these methods of combination. It is typical in observational studies to statistically test many questions and not correct the analysis results for multiple testing or multiple modeling, MTMM. The same problem can happen for randomized, experimental trials. There is the additional problem that some of the base studies may be using fabricated or fraudulent data. If there is no attention to MTMM or fraud in the base studies, there is no guarantee that the results to be combined are unbiased, the key requirement for the valid combining of results. We note that methods of combination are not robust; even one extreme base study value can overwhelm standard methods of combination. It is possible that multiple, extreme (MTMM or fraudulent) results can feed from the base studies to bias the combined result. A meta-analysis of observational (or even randomized studies) may not be reliable. Examples are given along with some methods to evaluate existing base studies and meta-analysis studies.

stat.ME

Particulate Matter Exposure and Lung Cancer: A Review of two Meta-Analysis Studies

The current regulatory paradigm is that PM2.5, over time causes lung cancer. This claim is based on cohort studies and meta-analysis that use cohort studies as their base studies. There is a need to evaluate the reliability of this causal claim. Our idea is to examine the base studies with respect to multiple testing and multiple modeling and to look closer at the meta-analysis using p-value plots. For two meta-analysis we investigated, some extremely small p-values were observed in some of the base studies, which we think are due to a combination of bias and small standard errors. The p-value plot for one meta-analysis indicates no effect. For the other meta-analysis, we note the p-value plot is consistent with a two-component mixture. Small p-values might be real or due to some combination of p-hacking, publication bias, covariate problems, etc. The large p-values could indicate no real effect, or be wrong due to low power, missing covariates, etc. We conclude that the results are ambiguous at best. These meta-analyses do not establish that PM2.5 is causal of lung tumors.

stat.AP

PM2.5 and all-cause mortality

The US EPA and the WHO claim that PM2.5 is causal of all-cause deaths. Both support and fund research on air quality and health effects. WHO funded a massive systematic review and meta-analyses of air quality and health-effect papers. 1,632 literature papers were reviewed and 196 were selected for meta-analyses. The standard air components, particulate matter, PM10 and PM2.5, nitrogen dioxide, NO2, and ozone, were selected as causes and all-cause and cause-specific mortalities were selected as outcomes. A claim was made for PM2.5 and all-cause deaths, risk ratio of 1.0065, with confidence limits of 1.0044 to 1.0086. There is a need to evaluate the reliability of this causal claim. Based on a p-value plot and discussion of several forms of bias, we conclude that the association is not causal.

stat.AP

Reliability of meta-analysis of an association between ambient air quality and development of asthma later in life

Claims from observational studies often fail to replicate. A study was undertaken to assess the reliability of cohort studies used in a highly cited meta-analysis of the association between ambient nitrogen dioxide, NO2, and fine particulate matter, PM2.5, concentrations early in life and development of asthma later in life. The numbers of statistical tests possible were estimated for 19 base papers considered for the meta-analysis. A p-value plot for NO2 and PM2.5 was constructed to evaluate effect heterogeneity of p-values used from the base papers. The numbers of statistical tests possible in the base papers were large - median 13,824, interquartile range 1,536-221,184; range 96-42M, in comparison to statistical test results presented. Statistical test results drawn from the base papers are unlikely to provide unbiased measures for meta-analysis. The p-value plot indicated that heterogeneity of the NO2 results across the base papers is consistent with a two-component mixture. First, it makes no sense to average across a mixture in meta-analysis. Second, the shape of the p-value plot for NO2 appears consistent with the possibility of analysis manipulation to obtain small p-values in several of the cohort studies. As for PM2.5, all corresponding p-values fall on a 45-degree line indicating complete randomness rather than a true association. Our interpretation of the meta-analysis is that the random p-values indicating no cause-effect associations are more plausible and that their meta-analysis will not likely replicate in the absence of bias. We conclude that claims made in the base papers used for meta-analysis are unreliable due to bias induced by multiple testing and multiple modelling, MTMM. We also show there is evidence that the heterogeneity across the base papers used for meta-analysis is more complex than simple sampling from a normal process.

stat.AP

Evaluation of a meta-analysis of ambient air quality as a risk factor for asthma exacerbation

False-positive results and bias may be common features of the biomedical literature today, including risk factor-chronic disease research. A study was undertaken to assess the reliability of base studies used in a meta-analysis examining whether carbon monoxide, particulate matter 10 and 2.5 micro molar, sulfur dioxide, nitrogen dioxide and ozone are risk factors for asthma exacerbation (hospital admission and emergency room visits for asthma attack). The number of statistical tests and models were counted in 17 randomly selected base papers from 87 used in the meta-analysis. P-value plots for each air component were constructed to evaluate the effect heterogeneity of p-values used from all 87 base papers The number of statistical tests possible in the 17 selected base papers was large, median=15,360 (interquartile range=1,536 to 40,960), in comparison to results presented. Each p-value plot showed a two-component mixture with small p-values less than .001 while other p-values appeared random (p-values greater than .05). Given potentially large numbers of statistical tests conducted in the 17 selected base papers, p-hacking cannot be ruled out as explanations for small p-values. Our interpretation of the meta-analysis is that the random p-values indicating null associations are more plausible and that the meta-analysis will not likely replicate in the absence of bias. We conclude the meta-analysis and base papers used are unreliable and do not offer evidence of value to inform public health practitioners about air quality as a risk factor for asthma exacerbation. The following areas are crucial for enabling improvements in risk factor chronic disease observational studies at the funding agency and journal level: preregistration, changes in funding agency and journal editor (and reviewer) practices, open sharing of data and facilitation of reproducibility research.

stat.AP

Evaluation of a meta-analysis of air quality and heart attacks, a case study

It is generally acknowledged that claims from observational studies often fail to replicate. An exploratory study was undertaken to assess the reliability of base studies used in meta-analysis of short-term air quality-myocardial infarction risk and to judge the reliability of statistical evidence from meta-analysis that uses data from observational studies. A highly cited meta-analysis paper examining whether short-term air quality exposure triggers myocardial infarction was evaluated as a case study. The paper considered six air quality components - carbon monoxide, nitrogen dioxide, sulfur dioxide, particulate matter 10 and 2.5 micrometers in diameter (PM10 and PM2.5), and ozone. The number of possible questions and statistical models at issue in each of 34 base papers used were estimated and p-value plots for each of the air components were constructed to evaluate the effect heterogeneity of p-values used from the base papers. Analysis search spaces (number of statistical tests possible) in the base papers were large, median of 12,288, interquartile range: 2,496 to 58,368, in comparison to actual statistical test results presented. Statistical test results taken from the base papers may not provide unbiased measures of effect for meta-analysis. Shapes of p-value plots for the six air components were consistent with the possibility of analysis manipulation to obtain small p-values in several base papers. Results suggest the appearance of heterogeneous, researcher-generated p-values used in the meta-analysis rather than unbiased evidence of real effects for air quality. We conclude that this meta-analysis does not provide reliable evidence for an association of air quality components with myocardial risk.

stat.AP

The reliability of an environmental epidemiology meta-analysis, a case study

Summary Background Claims made in science papers are coming under increased scrutiny with many claims failing to replicate. Meta-analysis studies that use unreliable observational studies should be in question. We examine the reliability of the base studies used in an air quality/heart attack meta-analysis and the resulting meta-analysis. Methods A meta-analysis study that includes 14 observational air quality/heart attack studies is examined for its statistical reliability. We use simple counting to evaluate the reliability of the base papers and a p-value plot of the p-values from the base studies to examine study heterogeneity. Findings We find that the based papers have massive multiple testing and multiple modeling with no statistical adjustments. Statistics coming from the base papers are not guaranteed to be unbiased, a requirement for a valid meta-analysis. There is study heterogeneity for the base papers with strong evidence for so called p-hacking. Interpretation We make two observations: there are many claims at issue in each of the 14 base studies so uncorrected multiple testing is a serious issue. We find the base papers and the resulting meta-analysis are unreliable.

stat.AP

Combined background information for meta-analysis evaluation

Massive numbers of meta-analysis studies are being published. A Google Scholar search of "systematic review and meta-analysis" returns about 452k hits since 2014. The search was done on Jan 14, 2019. There is a need to have some way to judge the reliability of a positive claim made in a meta-analysis that uses observational studies. Our idea is to examine the quality of the observational studies used in the meta-analysis and to examine the heterogeneity of those studies. We provide background information and examples: a listing of negative studies, a simulation of p-value plots, and multiple examples of p-value plots.

stat.AP

The reliability of a nutritional meta-analysis study

Background: Many researchers have studied the relationship between diet and health. There are papers showing an association between the consumption of sugar-sweetened beverages and Type 2 diabetes. Many meta-analyses use individual studies that do not adjust for multiple testing or multiple modeling and thus provide biased estimates of effect. Hence the claims reported in a meta-analysis paper may be unreliable if the primary papers do not ensure unbiased estimates of effect. Objective: Determine the statistical reliability of 10 papers and indirectly the reliability of the meta-analysis study. Method: Ten primary papers used in a meta-analysis paper and counted the numbers of outcomes, predictors, and covariates. We estimated the size of the potential analysis search space available to the authors of these papers; i.e. the number of comparisons and models available. Since we noticed that there were differences between predictors and covariates cited in the abstract and in the text, we applied this formula to information found in the abstracts, Space A, as well as the text, Space T, of each primary paper. Results: The median and range of the number of comparisons possible across the primary papers are 6.5 and (2-12,288) for abstracts, and 196,608 and (3,072-117,117,952) the texts. Note that the median of 6.5 for Space A is misleading as each primary study has 60-165 foods not mentioned in the abstract. Conclusion: Given that testing is at the 0.05 level and the number of comparisons is very large, nominal statistical significance is very weak support for a claim. The claims in these papers are not statistically supported and hence are unreliable. Thus, the claims of the meta-analysis paper lack evidentiary confirmation.

stat.AP