SearcharxivSearch

arXiv subjects

Taojun Hu

Publications and source records attributed to Taojun Hu.

6 recordsLinked to original sources

Copas-Jackson-type bounds for publication bias over a general class of selection models

Publication bias (PB) is one of the most vital threats to the accuracy of meta-analysis. Adjustment or sensitivity analysis based on selection models, which describe the probability of a study being published, provide a more objective evaluation of PB than widely-used simple graphical methods such as the trim-and-fill method. Most existing methods rely on parametric selection models. The Copas-Jackson bound (C-J bound) provides a worst-case bound of an analytical form over a nonparametric class of selection models, which would provide more robust conclusions than parametric sensitivity analysis. The nonparametric class of the selection models in the C-J bound is restrictive and only covers parametric selection models monotonic to the standard errors of outcomes. The novelty of this paper is to develop a method that constructs worst-case bounds over a general class of selection models weakening the assumption in the C-J bound. We propose an efficient numerical method to obtain an approximate worst-case bound via tractable nonlinear programming with linear constraints. We substantiate the effectiveness of the proposed bound with extensive simulation studies and show its applicability with two real-world meta-analyses.

stat.ME

A likelihood-based sensitivity analysis for addressing publication bias in meta-analysis of diagnostic studies using exact likelihood

Publication bias (PB) poses a significant threat to meta-analysis, as studies yielding notable results are more likely to be published in scientific journals. Sensitivity analysis provides a flexible method to address PB and to examine the impact of unpublished studies. A selection model based on t-statistics to sensitivity analysis is proposed by Copas. This t-statistics selection model is interpretable and enables the modeling of biased publication sampling across studies, as indicated by the asymmetry in the funnel-plot. In meta-analysis of diagnostic studies, the summary receiver operating characteristic curve is an essential tool for synthesizing the bivariate outcomes of sensitivity and specificity reported by individual studies. Previous studies address PB upon the bivariate normal model but these methods rely on the normal approximation for the empirical logit-transformed sensitivity and specificity, which is not suitable for sparse data scenarios. Compared to the bivariate normal model, the bivariate binomial model which replaces the normal approximation in the within-study model with the exact within-study model has better finite sample properties. In this study, we applied the Copas t-statistics selection model to the meta-analysis of diagnostic studies using the bivariate binomial model. To our knowledge, this is the first study to apply the Copas t-statistics selection model to the bivariate binomial model. We have evaluated our proposed method through several real-world meta-analyses of diagnostic studies and simulation studies.

stat.AP

Copas-Heckman-type sensitivity analysis for publication bias in rare-event meta-analysis under generalized linear mixed models

In systematic reviews and meta-analyses, publication bias (PB) is one of the serious concerns and mainly induced by selective publication of academic literatures. Although many methods have been proposed to deal with PB, almost all the methods are based on the normal-normal (NN) random-effects model assuming that data are normally distributed in both the within-study and the between-study levels. For rare-event meta-analysis where data contain rare occurrences of events, the standard NN random-effects model may perform poorly. Instead, some generalized linear mixed models (GLMMs) which employ the exact distribution for the number of events in within-study level provide alternatives and have been widely used in practice. However, limited methods can be applied to deal with PB in the GLMMs. To address this limitation, we propose a framework of sensitivity analysis for evaluating the impact of PB in various GLMMs. The proposed framework is developed based on the famous Copas-Heckman-type sensitivity analysis methods and can be easily implemented with the standard software with small computational cost. In this paper, we conduct simulation studies to assess the performance of proposed methods in adjusting PB and compare the results with related existing methods. Several real-world examples are also analyzed to show the broad applicability of our proposal in evaluating the potential impact of PB in meta-analysis of odds ratios and proportions with rare-event outcomes.

stat.ME

Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions

Natural Language Processing (NLP) is witnessing a remarkable breakthrough driven by the success of Large Language Models (LLMs). LLMs have gained significant attention across academia and industry for their versatile applications in text generation, question answering, and text summarization. As the landscape of NLP evolves with an increasing number of domain-specific LLMs employing diverse techniques and trained on various corpus, evaluating performance of these models becomes paramount. To quantify the performance, it's crucial to have a comprehensive grasp of existing metrics. Among the evaluation, metrics which quantifying the performance of LLMs play a pivotal role. This paper offers a comprehensive exploration of LLM evaluation from a metrics perspective, providing insights into the selection and interpretation of metrics currently in use. Our main goal is to elucidate their mathematical formulations and statistical interpretations. We shed light on the application of these metrics using recent Biomedical LLMs. Additionally, we offer a succinct comparison of these metrics, aiding researchers in selecting appropriate metrics for diverse tasks. The overarching goal is to furnish researchers with a pragmatic guide for effective LLM evaluation and metric selection, thereby advancing the understanding and application of these large language models.

cs.CL

Sensitivity analysis for publication bias in meta-analysis of sparse data based on exact likelihood

Meta-analysis is a powerful tool to synthesize findings from multiple studies. The normal-normal random-effects model is widely used to account for between-study heterogeneity. However, meta-analysis of sparse data, which may arise when the event rate is low for binary or count outcomes, poses a challenge to the normal-normal random-effects model in the accuracy and stability in inference since the normal approximation in the within-study model may not be good. To reduce bias arising from data sparsity, the generalized linear mixed model can be used by replacing the approximate normal within-study model with an exact model. Publication bias is one of the most serious threats in meta-analysis. Several quantitative sensitivity analysis methods for evaluating the potential impacts of selective publication are available for the normal-normal random-effects model. We propose a sensitivity analysis method by extending the likelihood-based sensitivity analysis with the t-statistic selection function of Copas to several generalized linear mixed-effects models. Through applications of our proposed method to several real-world meta-analysis and simulation studies, the proposed method was proven to outperform the likelihood-based sensitivity analysis based on the normal-normal model. The proposed method would give useful guidance to address publication bias in meta-analysis of sparse data.

stat.ME

ADRNet: A Generalized Collaborative Filtering Framework Combining Clinical and Non-Clinical Data for Adverse Drug Reaction Prediction

Adverse drug reaction (ADR) prediction plays a crucial role in both health care and drug discovery for reducing patient mortality and enhancing drug safety. Recently, many studies have been devoted to effectively predict the drug-ADRs incidence rates. However, these methods either did not effectively utilize non-clinical data, i.e., physical, chemical, and biological information about the drug, or did little to establish a link between content-based and pure collaborative filtering during the training phase. In this paper, we first formulate the prediction of multi-label ADRs as a drug-ADR collaborative filtering problem, and to the best of our knowledge, this is the first work to provide extensive benchmark results of previous collaborative filtering methods on two large publicly available clinical datasets. Then, by exploiting the easy accessible drug characteristics from non-clinical data, we propose ADRNet, a generalized collaborative filtering framework combining clinical and non-clinical data for drug-ADR prediction. Specifically, ADRNet has a shallow collaborative filtering module and a deep drug representation module, which can exploit the high-dimensional drug descriptors to further guide the learning of low-dimensional ADR latent embeddings, which incorporates both the benefits of collaborative filtering and representation learning. Extensive experiments are conducted on two publicly available real-world drug-ADR clinical datasets and two non-clinical datasets to demonstrate the accuracy and efficiency of the proposed ADRNet. The code is available at https://github.com/haoxuanli-pku/ADRnet.

cs.IR