SearcharxivSearch

arXiv subjects

Stephen Burgess

Publications and source records attributed to Stephen Burgess.

At least 19 recordsLinked to original sources

Multivariable Mendelian randomization with weak instruments: a comparison of Bayesian and frequentist methods

Weak instruments are a well known limitation for valid causal inference in Mendelian randomization studies. In the single exposure setting, weak instrument bias can be mitigated by selecting genetic instruments which are strongly associated with the exposure according to p-value and/or F-statistic thresholds. However, in the multi-exposure setting, genetic instruments may be strongly associated with an exposure but weakly associated with it conditional on all other exposures in the analysis. It is therefore more difficult to guarantee conditionally strong instruments in multivariable Mendelian randomization. Weak instrument bias can be mitigated using modelling approaches, however there are fewer methods for doing this in the multivariable case compared with the single exposure case. In this paper, we consider a method for mitigating weak instrument bias in multivariable Mendelian randomization using a Bayesian framework: MVMR-Pony. We compare this method with existing frequentist methods. We show using simulation studies that the MVMR-Pony method outperforms the frequentist approaches with respect to bias, coverage, type I error rates, and power, across settings where weak instrument bias arises due to correlated genetic effects, measurement error, and mediation.

stat.ME

Efficient semiparametric estimation of marginal treatment effects with genetic instrumental variables

Alcohol misuse is a key target of public health strategies aimed at reducing cardiovascular risk. The effect of excessive alcohol consumption on blood pressure may vary systematically with individuals' unobserved propensity to engage in heavy drinking, complicating causal inference with observational data. The marginal treatment effects framework uses an instrumental variable for treatment choice (excessive alcohol consumption) to study how selection into treatment is linked with the treatment effect. We explore the use of a genetic instrument within this framework, which is challenging because genetic compliers (individuals for whom a change in the instrument changes their treatment choice) are likely to be a small proportion of the overall sample. This can lead to greater sampling uncertainty in the tails of the propensity score distribution, i.e., the conditional probability of choosing treatment, and in turn poor estimation of causal estimands that measure heterogeneous treatment effects. We show that the use of efficient influence functions of target estimands improves estimation in terms of robustness to sampling uncertainty in nonparametrically estimated propensity scores. We find evidence of reverse selection on gains: individuals most prone to excessive alcohol consumption experience larger adverse effects on blood pressure.

stat.ME

Mendelian randomization in a multi-ancestry world: reflections and practical advice

Many Mendelian randomization (MR) papers have been conducted only in people of European ancestry, limiting transportability of results to the global population. Expanding MR to diverse ancestry groups is essential to ensure equitable biomedical insights, yet presents analytical and conceptual challenges. This review examines the practical challenges of MR analyses beyond the European only context, including use of data from multi-ancestry, mismatched ancestry, and admixed populations. We explain how apparent heterogeneity in MR estimates between populations can arise from differences in genetic variant frequencies and correlation patterns, as well as from differences in the distribution of phenotypic variables, complicating the detection of true differences in the causal pathway. We summarize published strategies for selecting genetic instruments and performing analyses when working with limited ancestry-specific data, discussing the assumptions needed in each case for incorporating external data from different ancestry populations. We conclude that differences in MR estimates by ancestry group should be interpreted cautiously, with consideration of how the identified differences may arise due to social and cultural factors. Corroborating evidence of a biological mechanism altering the causal pathway is needed to support a conclusion of differing causal pathways between ancestry groups.

q-bio.PE

Context-stratified Mendelian randomization: exploiting regional exposure variation to explore causal effect heterogeneity and non-linearity

Mendelian randomization (MR) uses genetic variants as instrumental variables to make causal claims. Standard MR approaches typically report a single population-averaged estimate, limiting their ability to explore effect heterogeneity or non-linear dose-response relationships. Existing stratification methods, such as residual-based and doubly-ranked stratified MR, attempt to overcome this but rely on strong and unverifiable assumptions. We propose an alternative, context-stratified Mendelian randomization, which exploits exogenous variation in the exposure across subgroups -- such as recruitment centres, geographic regions, or time periods -- to investigate effect heterogeneity and non-linearity. Separate MR analyses are performed within each context, and heterogeneity in the resulting estimates is assessed using Cochran's Q statistic and meta-regression. We demonstrate through simulations that the approach detects heterogeneity when present while maintaining nominal false positive rates under homogeneity when appropriate methods are used. In an applied example using UK Biobank data, we assess the effect of vitamin D levels on coronary artery disease risk across 20 recruitment centres. Despite some regional variation in vitamin D distributions, there is no evidence for a causal effect or heterogeneity in estimates. Compared to stratification methods requiring model-based assumptions, the context-stratified approach is simple to implement and robust to collider bias, provided the context variable is exogenous. However, the method's power and interpretability depend critically on meaningful exogenous variation in exposure distributions between contexts. In the example of vitamin D, subgroups from other stratification methods explored a much wider range of the exposure distribution.

stat.ME

Stratification-based Instrumental Variable Analysis Framework for Nonlinear Effect Analysis

Nonlinear causal effects are prevalent in many research scenarios involving continuous exposures, and instrumental variables (IVs) can be employed to investigate such effects, particularly in the presence of unmeasured confounders. However, common IV methods for nonlinear effect analysis, such as IV regression or the control-function method, have inherent limitations, leading to either low statistical power or potentially misleading conclusions. In this work, we propose an alternative IV framework for nonlinear effect analysis, which has recently emerged in genetic epidemiology and addresses many of the drawbacks of existing IV methods. This framework enables study of the effect function while avoiding unnecessary model assumptions. In particular, it facilitates the identification of change points or threshold values in causal effects. Through a wide variety of simulations, we demonstrate that our framework outperforms other representative nonlinear IV methods in predicting the effect shape when the instrument is weak and can accurately estimate the effect function as well as identify the change point and predict its value under various structural model and effect shape scenarios. We further apply our framework to assess the nonlinear effect of alcohol consumption on systolic blood pressure using a genetic instrument (i.e. Mendelian randomization) with UK Biobank data. Our analysis detects a threshold beyond which alcohol intake exhibits a clear causal effect on the outcome. Our results are consistent with published medical guidelines.

stat.ME

Outlier Detection in Mendelian Randomisation

Mendelian Randomisation (MR) uses genetic variants as instrumental variables to infer causal effects of exposures on an outcome. One key assumption of MR is that the genetic variants used as instrumental variables are independent of the outcome conditional on the risk factor and unobserved confounders. Violations of this assumption, i.e. the effect of the instrumental variables on the outcome through a path other than the risk factor included in the model (which can be caused by pleiotropy), are common phenomena in human genetics. Genetic variants, which deviate from this assumption, appear as outliers to the MR model fit and can be detected by the general heterogeneity statistics proposed in the literature, which are known to suffer from overdispersion, i.e. too many genetic variants are declared as false outliers. We propose a method that corrects for overdispersion of the heterogeneity statistics in uni- and multivariable MR analysis by making use of the estimated inflation factor to correctly remove outlying instruments and therefore account for pleiotropic effects. Our method is applicable to summary-level data.

stat.ME

Weak instruments in multivariable Mendelian randomization: methods and practice

The method of multivariable Mendelian randomization uses genetic variants to instrument multiple exposures, to estimate the effect that a given exposure has on an outcome conditional on all other exposures included in a linear model. Unfortunately, the inclusion of every additional exposure makes a weak instruments problem more likely, because we require conditionally strong genetic predictors of each exposure. This issue is well appreciated in practice, with different versions of F-statistics routinely reported as measures of instument strength. Less transparently, however, these F-statistics are sometimes used to guide instrument selection, and even to decide whether to report empirical results. Rather than discarding findings with low F-statistics, weak instrument-robust methods can provide valid inference under weak instruments. For multivariable Mendelian randomization with two-sample summary data, we encourage use of the inference strategy of Andrews (2018) that reports both robust and non-robust confidence sets, along with a statistic that measures how reliable the non-robust confidence set is in terms of coverage. We also propose a novel adjusted-Kleibergen statistic that corrects for overdispersion heterogeneity in genetic associations with the outcome.

stat.ME

Estimating time-varying exposure effects through continuous-time modelling in Mendelian randomization

Mendelian randomization is an instrumental variable method that utilizes genetic information to investigate the causal effect of a modifiable exposure on an outcome. In most cases, the exposure changes over time. Understanding the time-varying causal effect of the exposure can yield detailed insights into mechanistic effects and the potential impact of public health interventions. Recently, a growing number of Mendelian randomization studies have attempted to explore time-varying causal effects. However, the proposed approaches oversimplify temporal information and rely on overly restrictive structural assumptions, limiting their reliability in addressing time-varying causal problems. This paper considers a novel approach to estimate time-varying effects through continuous-time modelling by combining functional principal component analysis and weak-instrument-robust techniques. Our method effectively utilizes available data without making strong structural assumptions and can be applied in general settings where the exposure measurements occur at different timepoints for different individuals. We demonstrate through simulations that our proposed method performs well in estimating time-varying effects and provides reliable inference results when the time-varying effect form is correctly specified. The method could theoretically be used to estimate arbitrarily complex time-varying effects. However, there is a trade-off between model complexity and instrument strength. Estimating complex time-varying effects requires instruments that are unrealistically strong. We illustrate the application of this method in a case study examining the time-varying effects of systolic blood pressure on urea levels.

stat.ME

A frequentist test of proportional colocalization after selecting relevant genetic variants

Colocalization analyses assess whether two traits are affected by the same or distinct causal genetic variants in a single gene region. A class of Bayesian colocalization tests are now routinely used in practice; for example, for genetic analyses in drug development pipelines. In this work, we consider an alternative frequentist approach to colocalization testing that examines the proportionality of genetic associations with each trait. The proportional colocalization approach uses markedly different assumptions to Bayesian colocalization tests, and therefore can provide valuable complementary evidence in cases where Bayesian colocalization results are inconclusive or sensitive to priors. We propose a novel conditional test of proportional colocalization, prop-coloc-cond, that aims to account for the uncertainty in variant selection, in order to recover accurate type I error control. The test can be implemented straightforwardly, requiring only summary data on genetic associations. Simulation evidence and an empirical investigation into GLP1R gene expression demonstrates how tests of proportional colocalization can offer important insights in conjunction with Bayesian colocalization tests.

stat.ME

Illustrating the structures of bias from immortal time using directed acyclic graphs

Background: Immortal time is a period of follow-up during which death or the study outcome cannot occur by design. Bias from immortal time has been increasingly recognized in epidemiologic studies. However, the fundamental causes and structures of bias from immortal time have not been explained systematically using a structural approach. Methods: We use an example "Do Nobel Prize winners live longer than less recognized scientists?" for illustration. We illustrate how immortal time arises and present the structures of bias from immortal time using time-varying directed acyclic graphs (DAGs). We further explore the structures of bias with the exclusion of immortal time and with the presence of competing risks. We discuss how these structures are shared by different study designs in pharmacoepidemiology and provide solutions, where possible, to address the bias. Results: We illustrate that immortal time arises from using postbaseline information to define exposure or eligibility. We use time-varying DAGs to explain the structures of bias from immortal time are confounding by survival until exposure allocation or selection bias from selecting on survival until eligibility. We explain that excluding immortal time from the follow-up does not fully address this confounding or selection bias, and that the presence of competing risks can worsen the bias. Bias from immortal time may be avoided by aligning time zero, exposure allocation and eligibility, and by excluding individuals with prior exposure. Conclusions: Understanding bias from immortal time in terms of confounding or selection bias helps researchers identify and thereby avoid or ameliorate this bias.

stat.ME

Selecting invalid instruments to improve Mendelian randomization with two-sample summary data

Mendelian randomization (MR) is a widely-used method to estimate the causal relationship between a risk factor and disease. A fundamental part of any MR analysis is to choose appropriate genetic variants as instrumental variables. Genome-wide association studies often reveal that hundreds of genetic variants may be robustly associated with a risk factor, but in some situations investigators may have greater confidence in the instrument validity of only a smaller subset of variants. Nevertheless, the use of additional instruments may be optimal from the perspective of mean squared error even if they are slightly invalid; a small bias in estimation may be a price worth paying for a larger reduction in variance. For this purpose, we consider a method for "focused" instrument selection whereby genetic variants are selected to minimise the estimated asymptotic mean squared error of causal effect estimates. In a setting of many weak and locally invalid instruments, we propose a novel strategy to construct confidence intervals for post-selection focused estimators that guards against the worst case loss in asymptotic coverage. In empirical applications to: (i) validate lipid drug targets; and (ii) investigate vitamin D effects on a wide range of outcomes, our findings suggest that the optimal selection of instruments does not involve only a small number of biologically-justified instruments, but also many potentially invalid instruments.

stat.ME

Statistical Methods for cis-Mendelian Randomization with Two-sample Summary-level Data

Mendelian randomization is the use of genetic variants to assess the existence of a causal relationship between a risk factor and an outcome of interest. Here, we focus on two-sample summary-data Mendelian randomization analyses with many correlated variants from a single gene region, and particularly on cis-Mendelian randomization studies which use protein expression as a risk factor. Such studies must rely on a small, curated set of variants from the studied region; using all variants in the region requires inverting an ill-conditioned genetic correlation matrix and results in numerically unstable causal effect estimates. We review methods for variable selection and estimation in cis-Mendelian randomization with summary-level data, ranging from stepwise pruning and conditional analysis to principal components analysis, factor analysis and Bayesian variable selection. In a simulation study, we show that the various methods have a comparable performance in analyses with large sample sizes and strong genetic instruments. However, when weak instrument bias is suspected, factor analysis and Bayesian variable selection produce more reliable inferences than simple pruning approaches, which are often used in practice. We conclude by examining two case studies, assessing the effects of LDL-cholesterol and serum testosterone on coronary heart disease risk using variants in the HMGCR and SHBG gene regions respectively.

q-bio.QM

Bias in multivariable Mendelian randomization studies due to measurement error on exposures

Multivariable Mendelian randomization estimates the causal effect of multiple exposures on an outcome, typically using summary statistics of genetic variant associations. However, exposures of interest in Mendelian randomization applications will often be measured with error. The summary statistics will therefore not be of the genetic associations with the exposure, but with the exposure measured with error. Classical measurement error will not bias genetic association estimates but will increase their standard errors. With a single exposure, this will result in bias toward the null in a two sample framework. However, this will not necessarily be the case with multiple correlated exposures. In this paper, we examine how the direction and size of bias, as well as coverage, power and type I error rates in multivariable Mendelian randomization studies are affected by measurement error on exposures. We show how measurement error can be accounted for in a maximum likelihood framework. We consider two applied examples. In the first, we show that measurement error leads to the effect of body mass index on coronary heart disease risk to be overestimated, and that of waist-to-hip ratio to be underestimated. In the second, we show that the proportion of the effect of education on coronary heart disease risk which is mediated by body mass index, smoking and blood pressure may be underestimated if measurement error is not taken into account.

stat.ME

Conditional inference in cis-Mendelian randomization using weak genetic factors

Mendelian randomization is a widely-used method to estimate the unconfounded effect of an exposure on an outcome by using genetic variants as instrumental variables. Mendelian randomization analyses which use variants from a single genetic region (cis-MR) have gained popularity for being an economical way to provide supporting evidence for drug target validation. This paper proposes methods for cis-MR inference which use the explanatory power of many correlated variants to make valid inferences even in situations where those variants only have weak effects on the exposure. In particular, we exploit the highly structured nature of genetic correlations in single gene regions to reduce the dimension of genetic variants using factor analysis. These genetic factors are then used as instrumental variables to construct tests for the causal effect of interest. Since these factors may often be weakly associated with the exposure, size distortions of standard t-tests can be severe. Therefore, we consider two approaches based on conditional testing. First, we extend results of commonly-used identification-robust tests to account for the use of estimated factors as instruments. Secondly, we propose a test which appropriately adjusts for first-stage screening of genetic factors based on their relevance. Our empirical results provide genetic evidence to validate cholesterol-lowering drug targets aimed at preventing coronary heart disease.

stat.ME

Disentangling the effects of traits with shared clustered genetic predictors using multivariable Mendelian randomization

When genetic variants in a gene cluster are associated with a disease outcome, the causal pathway from the variants to the outcome can be difficult to disentangle. For example, the chemokine receptor gene cluster contains genetic variants associated with various cytokines. Associations between variants in this cluster and stroke risk may be driven by any of these cytokines. Multivariable Mendelian randomization is an extension of standard univariable Mendelian randomization to estimate the direct effects of related exposures with shared genetic predictors. However, when genetic variants are clustered, a Goldilocks dilemma arises: including too many highly-correlated variants in the analysis can lead to ill-conditioning, but pruning variants too aggressively can lead to imprecise estimates or even lack of identification. We propose multivariable methods that use principal component analysis to reduce many correlated genetic variants into a smaller number of orthogonal components that are used as instrumental variables. We show in simulations that these methods result in more precise estimates that are less sensitive to numerical instability due to both strong correlations and small changes in the input data. We apply the methods to demonstrate the most likely causal risk factor for stroke at the chemokine gene cluster is monocyte chemoattractant protein-1.

stat.ME

Pleiotropy robust methods for multivariable Mendelian randomization

Mendelian randomization is a powerful tool for inferring the presence, or otherwise, of causal effects from observational data. However, the nature of genetic variants is such that pleiotropy remains a barrier to valid causal effect estimation. There are many options in the literature for pleiotropy robust methods when studying the effects of a single risk factor on an outcome. However, there are few pleiotropy robust methods in the multivariable setting, that is, when there are multiple risk factors of interest. In this paper we introduce three methods which build on common approaches in the univariable setting: MVMR-Robust; MVMR-Median; and MVMR-Lasso. We discuss the properties of each of these methods and examine their performance in comparison to existing approaches in a simulation study. MVMR-Robust is shown to outperform existing outlier robust approaches when there are low levels of pleiotropy. MVMR-Lasso provides the best estimation in terms of mean squared error for moderate to high levels of pleiotropy, and can provide valid inference in a three sample setting. MVMR-Median performs well in terms of estimation across all scenarios considered, and provides valid inference up to a moderate level of pleiotropy. We demonstrate the methods in an applied example looking at the effects of intelligence, education and household income on the risk of Alzheimer's disease.

stat.ME

An efficient and robust approach to Mendelian randomization with measured pleiotropic effects in a high-dimensional setting

Valid estimation of a causal effect using instrumental variables requires that all of the instruments are independent of the outcome conditional on the risk factor of interest and any confounders. In Mendelian randomization studies with large numbers of genetic variants used as instruments, it is unlikely that this condition will be met. Any given genetic variant could be associated with a large number of traits, all of which represent potential pathways to the outcome which bypass the risk factor of interest. Such pleiotropy can be accounted for using standard multivariable Mendelian randomization with all possible pleiotropic traits included as covariates. However, the estimator obtained in this way will be inefficient if some of the covariates do not truly sit on pleiotropic pathways to the outcome. We present a method which uses regularization to identify which out of a set of potential covariates need to be accounted for in a Mendelian randomization analysis in order to produce an efficient and robust estimator of a causal effect. The method can be used in the case where individual-level data are not available and the analysis must rely on summary-level data only. It can also be used in the case where there are more covariates under consideration than instruments, which is not possible using standard multivariable Mendelian randomization. We show the results of simulation studies which demonstrate the performance of the proposed regularization method in realistic settings. We also illustrate the method in an applied example which looks at the causal effect of urate plasma concentration on coronary heart disease.

stat.ME

Robust instrumental variable methods using multiple candidate instruments with application to Mendelian randomization

Mendelian randomization is the use of genetic variants to make causal inferences from observational data. The field is currently undergoing a revolution fuelled by increasing numbers of genetic variants demonstrated to be associated with exposures in genome-wide association studies, and the public availability of summarized data on genetic associations with exposures and outcomes from large consortia. A Mendelian randomization analysis with many genetic variants can be performed relatively simply using summarized data. However, a causal interpretation is only assured if each genetic variant satisfies the assumptions of an instrumental variable. To provide some protection against failure of these assumptions, robust methods for instrumental variable analysis have been proposed. Here, we develop three extensions to instrumental variable methods using: i) robust regression, ii) the penalization of weights from candidate instruments with heterogeneous causal estimates, and iii) L1 penalization. Results from a wide variety of robust methods, including the recently-proposed MR-Egger and median-based methods, are compared in an extensive simulation study. We demonstrate that two methods, robust regression in an inverse-variance weighted method and a simple median of the causal estimates from the individual variants, have considerably improved Type 1 error rates compared with conventional methods in a wide variety of scenarios when up to 30% of the genetic variants are invalid instruments. While the MR-Egger method gives unbiased estimates when its assumptions are satisfied, these estimates are less efficient than those from other methods and are highly sensitive to violations of the assumptions. Methods that make different assumptions should be used routinely to assess the robustness of findings from applied Mendelian randomization investigations with multiple genetic variants.

stat.ME