SearcharxivSearch

arXiv subjects

Harlan Campbell

Publications and source records attributed to Harlan Campbell.

At least 19 recordsLinked to original sources

Don't Disregard the Data for Lack of a Likelihood: Bayesian Synthetic Likelihood for Enhanced Multilevel Network Meta-Regression

Multilevel network meta-regression (ML-NMR) enables population-adjusted indirect treatment comparisons by combining individual patient data (IPD) with aggregate data. When individual-level covariates are unavailable, ML-NMR marginalizes over the covariate distribution, but this strategy cannot exploit subgroup-level summary results that are often available and potentially highly informative. We propose using Bayesian Synthetic Likelihood (BSL) to leverage this ancillary summary information and present an implementation strategy for Hamiltonian Monte Carlo (HMC), a gradient-based Markov chain Monte Carlo (MCMC) algorithm. At each MCMC iteration, the BSL method imputes missing covariates by sampling from the model-implied conditional distribution, computes synthetic subgroup summaries from the imputed data, and matches these synthetic summaries to observed summaries via a multivariate normal synthetic likelihood. Fitting this model with HMC presents multiple challenges: first, gradients cannot be computed exactly but must be estimated stochastically; and second, the model's likelihood may be non-differentiable at certain points, a pathology that can deeply frustrate the performance of HMC. We address these challenges with pre-drawn random numbers, continuous relaxation of the likelihood, and Pareto-smoothed importance sampling. This work (1) introduces a novel application of BSL to missing data problems where summary statistics from the complete dataset are available despite substantial missingness in the individual-level data, (2) demonstrates how BSL strategies can be implemented within Stan's HMC framework, and (3) shows, using a network of plaque psoriasis trials, that BSL-enhanced ML-NMR can substantially improve upon standard ML-NMR by leveraging informative ancillary information.

stat.ME

Hidden in Plain Sight: How Non-Collapsibility Biases Treatment Effects in (Network) Meta-Analysis

Network meta-analysis (NMA) is widely used to compare multiple interventions simultaneously by synthesizing direct and indirect evidence. The general fixed or random effects contrast-based NMA model can be applied to different outcomes and data structures by opting for either an arm-based or contrast-based likelihood depending on the data available. Depending on the outcome and link-function, we estimate either collapsible or non-collapsible effect measures. Using an illustrative example involving binary outcomes and the non-collapsible odds ratio, we demonstrate that the standard NMA model produces estimates for non-collapsible effect measures that are biased toward the null when studies in the evidence base enroll heterogeneous populations (mixtures of distinct risk groups) that vary across studies. Importantly, this also holds when there are no differences in effect-modifiers across studies; the standard assumption of a common treatment effect when there are no differences in the distribution of effect-modifiers across studies is not appropriate when studies have different baseline risks. As a potential solution, we propose a ``bookend'' approach that explicitly models mixed-population studies as weighted combinations of two homogeneous subpopulations identified from studies with extreme baseline risks and provide guidance for practitioners to determine if bias due to non-collapsibility may be a concern.

stat.ME

Doubly robust augmented weighting estimators for the analysis of externally controlled single-arm trials and unanchored indirect treatment comparisons

Externally controlled single-arm trials are critical to assess treatment efficacy across therapeutic indications for which randomized controlled trials are not feasible. A closely-related research design, the unanchored indirect treatment comparison, is often required for disconnected treatment networks in health technology assessment. We present a unified causal inference framework for both research designs. We develop a novel estimator that augments a popular weighting approach based on entropy balancing -- matching-adjusted indirect comparison (MAIC) -- by fitting a model for the conditional outcome expectation. The predictions of the outcome model are combined with the entropy balancing MAIC weights. While the standard MAIC estimator is singly robust where the outcome model is non-linear, our augmented MAIC approach is doubly robust, providing increased robustness against model misspecification. This is demonstrated in a simulation study with binary outcomes and a logistic outcome model, where the augmented estimator demonstrates its doubly robust property, while exhibiting higher precision than all non-augmented weighting estimators and near-identical precision to G-computation. We describe the extension of our estimator to the setting with unavailable individual participant data for the external control, illustrating it through an applied example. Our findings reinforce the understanding that entropy balancing-based approaches have desirable properties compared to standard ``modeling'' approaches to weighting, but should be augmented to improve protection against bias and guarantee double robustness.

stat.ME

A fully Bayesian approach for the imputation and analysis of derived outcome variables with missingness

Derived variables are variables that are constructed from one or more source variables through established mathematical operations or algorithms. For example, body mass index (BMI) is a derived variable constructed from two source variables: weight and height. When using a derived variable as the outcome in a statistical model, complications arise when some of the source variables have missing values. In this paper, we propose how one can define a single fully Bayesian model to simultaneously impute missing values and sample from the posterior. We compare our proposed method with alternative approaches that rely on multiple imputation with examples including an analysis to estimate the risk of microcephaly (a derived variable based on sex, gestational age and head circumference at birth) in newborns exposed to the ZIKA virus.

stat.ME

Augmented two-stage estimation for treatment crossover in oncology trials: Leveraging external data for improved precision

Randomized controlled trials (RCTs) in oncology often allow control group participants to crossover to experimental treatments, a practice that, while often ethically necessary, complicates the accurate estimation of long-term treatment effects. When crossover rates are high or sample sizes are limited, commonly used methods for crossover adjustment (such as the rank-preserving structural failure time model, inverse probability of censoring weights, and two-stage estimation (TSE)) may produce imprecise estimates. Real-world data (RWD) can be used to develop an external control arm for the RCT, although this approach ignores evidence from trial subjects who did not crossover and ignores evidence from the data obtained prior to crossover for those subjects who did. This paper introduces ''augmented two-stage estimation'' (ATSE), a method that combines data from non-switching participants in a RCT with an external dataset, forming a ''hybrid non-switching arm''. With a simulation study, we evaluate the ATSE method's performance compared to TSE crossover adjustment and an external control arm approach. Results indicate that performance is dependent on scenario characteristics, but when unconfounded external data are available, ATSE may result in less bias and improved precision compared to TSE and external control arm approaches. When external data are affected by unmeasured confounding, ATSE becomes prone to bias, but to a lesser extent compared to an external control arm approach.

stat.ME

Integrating representative and non-representative survey data for efficient inference

Non-representative surveys are commonly used and widely available but suffer from selection bias that generally cannot be entirely eliminated using weighting techniques. Instead, we propose a Bayesian method to synthesize longitudinal representative unbiased surveys with non-representative biased surveys by estimating the degree of selection bias over time. We show using a simulation study that synthesizing biased and unbiased surveys together out-performs using the unbiased surveys alone, even if the selection bias may evolve in a complex manner over time. Using COVID-19 vaccination data, we are able to synthesize two large sample biased surveys with an unbiased survey to reduce uncertainty in now-casting and inference estimates while simultaneously retaining the empirical credible interval coverage. Ultimately, we are able to conceptually obtain the properties of a large sample unbiased survey if the assumed unbiased survey, used to anchor the estimates, is unbiased for all time-points.

stat.ME

Estimating marginal treatment effects from observational studies and indirect treatment comparisons: When are standardization-based methods preferable to those based on propensity score weighting?

In light of newly developed standardization methods, we evaluate, via simulation study, how propensity score weighting and standardization -based approaches compare for obtaining estimates of the marginal odds ratio and the marginal hazard ratio. Specifically, we consider how the two approaches compare in two different scenarios: (1) in a single observational study, and (2) in an anchored indirect treatment comparison (ITC) of randomized controlled trials. We present the material in such a way so that the matching-adjusted indirect comparison (MAIC) and the (novel) simulated treatment comparison (STC) methods in the ITC setting may be viewed as analogous to the propensity score weighting and standardization methods in the single observational study setting. Our results suggest that current recommendations for conducting ITCs can be improved and underscore the importance of adjusting for purely prognostic factors.

stat.ME

Equivalence testing for linear regression

We introduce equivalence testing procedures for linear regression analyses. Such tests can be very useful for confirming the lack of a meaningful association between a continuous outcome and a continuous or binary predictor. Specifically, we propose an equivalence test for unstandardized regression coefficients and an equivalence test for semipartial correlation coefficients. We review how to define valid hypotheses, how to calculate p-values, and how these tests compare to an alternative Bayesian approach with applications to examples in the literature.

stat.ME

Defining a credible interval is not always possible with "point-null'' priors: A lesser-known correlate of the Jeffreys-Lindley paradox

In many common situations, a Bayesian credible interval will be, given the same data, very similar to a frequentist confidence interval, and researchers will interpret these intervals in a similar fashion. However, no predictable similarity exists when credible intervals are based on model-averaged posteriors whenever one of the two nested models under consideration is a so called ''point-null''. Not only can this model-averaged credible interval be quite different than the frequentist confidence interval, in some cases it may be undefined. This is a lesser-known correlate of the Jeffreys-Lindley paradox and is of particular interest given the popularity of the Bayes factor for testing point-null hypotheses.

math.ST

Bayes factors and posterior estimation: Two sides of the very same coin

Recently, several researchers have claimed that conclusions obtained from a Bayes factor (or the posterior odds) may contradict those obtained from Bayesian posterior estimation. In this short paper, we wish to point out that no such "incompatibility" exists if one is willing to consistently define one's priors and posteriors. The key for compatibility is that the (implied) prior model odds used for testing are the same as those used for estimation. Our recommendation is simple: If one reports a Bayes factor comparing two models, then one should also report posterior estimates which appropriately acknowledge the uncertainty with regards to which of the two models is correct.

math.ST

re:Linde et al. (2021): The Bayes factor, HDI-ROPE and frequentist equivalence tests can all be reverse engineered -- almost exactly -- from one another

Following an extensive simulation study comparing the operating characteristics of three different procedures used for establishing equivalence (the frequentist "TOST", the Bayesian "HDI-ROPE", and the Bayes factor interval null procedure), Linde et al. (2021) conclude with the recommendation that "researchers rely more on the Bayes factor interval null approach for quantifying evidence for equivalence." We redo the simulation study of Linde et al. (2021) in its entirety but with the different procedures calibrated to have the same predetermined maximum type 1 error rate. Our results suggest that, when calibrated in this way, the Bayes Factor, HDI-ROPE, and frequentist equivalence tests all have similar -- almost exactly -- type 2 error rates. In general any advocating for frequentist testing as better or worse than Bayesian testing in terms of empirical findings seems dubious at best. If one decides on which underlying principle to subscribe to in tackling a given problem, then the method follows naturally. Bearing in mind that each procedure can be reverse-engineered from the others (at least approximately), trying to use empirical performance to argue for one approach over another seems like tilting at windmills.

stat.ME

Adjusting for misclassification of an exposure in an individual participant data meta-analysis

A common problem in the analysis of multiple data sources, including individual participant data meta-analysis (IPD-MA), is the misclassification of binary variables. Misclassification may lead to biased estimates of model parameters, even when the misclassification is entirely random. We aimed to develop statistical methods that facilitate unbiased estimation of adjusted and unadjusted exposure-outcome associations and between-study heterogeneity in IPD-MA, where the extent and nature of exposure misclassification may vary across studies. We present Bayesian methods that allow misclassification of binary exposure variables to depend on study- and participant-level characteristics. In an example of the differential diagnosis of dengue using two variables, where the gold standard measurement for the exposure variable was unavailable for some studies which only measured a surrogate prone to misclassification, our methods yielded more accurate estimates than analyses naive with regard to misclassification or based on gold standard measurements alone. In a simulation study, the evaluated misclassification model yielded valid estimates of the exposure-outcome association, and was more accurate than analyses restricted to gold standard measurements. Our proposed framework can appropriately account for the presence of binary exposure misclassification in IPD-MA. It requires that some studies supply IPD for the surrogate and gold standard exposure and misclassification is exchangeable across studies conditional on observed covariates (and outcome). The proposed methods are most beneficial when few large studies that measured the gold standard are available, and when misclassification is frequent.

stat.ME

Measurement Error in Meta-Analysis (MEMA) -- a Bayesian framework for continuous outcome data

Ideally, a meta-analysis will summarize data from several unbiased studies. Here we consider the less than ideal situation in which contributing studies may be compromised by measurement error. Measurement error affects every study design, from randomized controlled trials to retrospective observational studies. We outline a flexible Bayesian framework for continuous outcome data which allows one to obtain appropriate point and interval estimates with varying degrees of prior knowledge about the magnitude of the measurement error. We also demonstrate how, if individual-participant data (IPD) are available, the Bayesian meta-analysis model can adjust for multiple participant-level covariates, measured with or without measurement error.

stat.ME

Bayesian adjustment for preferential testing in estimating the COVID-19 infection fatality rate

A key challenge in estimating the infection fatality rate (IFR) -- and its relation with various factors of interest -- is determining the total number of cases. The total number of cases is not known because not everyone is tested, but also, more importantly, because tested individuals are not representative of the population at large. We refer to the phenomenon whereby infected individuals are more likely to be tested than non-infected individuals, as "preferential testing." An open question is whether or not it is possible to reliably estimate the IFR without any specific knowledge about the degree to which the data are biased by preferential testing. In this paper we take a partial identifiability approach, formulating clearly where deliberate prior assumptions can be made and presenting a Bayesian model which pools information from different samples. When the model is fit to European data obtained from seroprevalence studies and national official COVID-19 statistics, we estimate the overall COVID-19 IFR for Europe to be 0.53%, 95% C.I. = [0.39%, 0.69%].

stat.ME

What to make of non-inferiority and equivalence testing with a post-specified margin?

In order to determine whether or not an effect is absent based on a statistical test, the recommended frequentist tool is the equivalence test. Typically, it is expected that an appropriate equivalence margin has been specified before any data are observed. Unfortunately, this can be a difficult task. If the margin is too small, then the test's power will be substantially reduced. If the margin is too large, any claims of equivalence will be meaningless. Moreover, it remains unclear how defining the margin afterwards will bias one's results. In this short article, we consider a series of hypothetical scenarios in which the margin is defined post-hoc or is otherwise considered controversial. We also review a number of relevant, potentially problematic actual studies from clinical trials research, with the aim of motivating a critical discussion as to what is acceptable and desirable in the reporting and interpretation of equivalence tests.

stat.ME

The consequences of checking for zero-inflation and overdispersion in the analysis of count data

Count data are ubiquitous in ecology and the Poisson generalized linear model (GLM) is commonly used to model the association between counts and explanatory variables of interest. When fitting this model to the data, one typically proceeds by first confirming that the data is not overdispersed and that there is no excess of zeros. If the data appear to be overdispersed or if there is any zero-inflation, key assumptions of the Poison GLM may be violated and researchers will then typically consider alternatives to the Poison GLM. An important question is whether the potential model selection bias introduced by this data-driven multi-stage procedure merits concern. In this paper, we conduct a large-scale simulation study to investigate the potential consequences of model selection bias that can arise in the simple scenario of analyzing a sample of potentially overdispersed, potentially zero-heavy, count data.

stat.ME

A non-inferiority test for R-squared with random regressors

Determining the lack of association between an outcome variable and a number of different explanatory variables is frequently necessary in order to disregard a proposed model. This paper proposes a non-inferiority test for the coefficient of determination (or squared multiple correlation coefficient), R-squared, in a linear regression analysis with random predictors. The test is derived from inverting a one-sided confidence interval based on a scaled central F distribution.

stat.ME

Can we disregard the whole model? Omnibus non-inferiority testing for $R^{2}$ in multivariable linear regression and $\hatη^{2}$ in ANOVA

Determining a lack of association between an outcome variable and a number of different explanatory variables is frequently necessary in order to disregard a proposed model (i.e., to confirm the lack of an association between an outcome and predictors). Despite this, the literature rarely offers information about, or technical recommendations concerning, the appropriate statistical methodology to be used to accomplish this task. This paper introduces non-inferiority tests for ANOVA and linear regression analyses, that correspond to the standard widely used $F$-test for $\hatη^2$ and $R^{2}$, respectively. A simulation study is conducted to examine the type I error rates and statistical power of the tests, and a comparison is made with an alternative Bayesian testing approach. The results indicate that the proposed non-inferiority test is a potentially useful tool for 'testing the null.'

stat.ME