SearcharxivSearch

arXiv subjects

Saonli Basu

Publications and source records attributed to Saonli Basu.

5 recordsLinked to original sources

A Random Effects Model-based Method of Moments Estimation of Causal Effect in Mendelian Randomization Studies

Recent advances in genotyping technology have delivered a wealth of genetic data, which is rapidly advancing our understanding of the underlying genetic architecture of complex diseases. Mendelian Randomization (MR) leverages such genetic data to estimate the causal effect of an exposure factor on an outcome from observational studies. In this paper, we utilize genetic correlations to summarize information on a large set of genetic variants associated with the exposure factor. Our proposed approach is a generalization of the MR-inverse variance weighting (IVW) approach where we can accommodate many weak and pleiotropic effects. Our approach quantifies the variation explained by all valid instrumental variables (IVs) instead of estimating the individual effects and thus could accommodate weak IVs. This is particularly useful for performing MR estimation in small studies, or minority populations where the selection of valid IVs is unreliable and thus has a large influence on the MR estimation. Through simulation and real data analysis, we demonstrate that our approach provides a robust alternative to the existing MR methods. We illustrate the robustness of our proposed approach under the violation of MR assumptions and compare the performance with several existing approaches.

stat.ME

Causal Effects in Twin Studies: the Role of Interference

The use of twins designs to address causal questions is becoming increasingly popular. A standard assumption is that there is no interference between twins---that is, no twin's exposure has a causal impact on their co-twin's outcome. However, there may be settings in which this assumption would not hold, and this would (1) impact the causal interpretation of parameters obtained by commonly used existing methods; (2) change which effects are of greatest interest; and (3) impact the conditions under which we may estimate these effects. We explore these issues, and we derive semi-parametric efficient estimators for causal effects in the presence of interference between twins. Using data from the Minnesota Twin Family Study, we apply our estimators to assess whether twins' consumption of alcohol in early adolescence may have a causal impact on their co-twins' substance use later in life.

stat.ME

A Robust and Unified Framework for Estimating Heritability in Twin Studies using Generalized Estimating Equations

The development of a complex disease is an intricate interplay of genetic and environmental factors. "Heritability" is defined as the proportion of total trait variance due to genetic factors within a given population. Studies with monozygotic (MZ) and dizygotic (DZ) twins allow us to estimate heritability by fitting an "ACE" model which estimates the proportion of trait variance explained by additive genetic (A), common shared environment (C), and unique non-shared environmental (E) latent effects, thus helping us better understand disease risk and etiology. In this paper, we develop a flexible generalized estimating equations framework ("GEE2") for fitting twin ACE models that requires minimal distributional assumptions, rather only the first two moments need to be correctly specified. We prove that two commonly used methods for estimating heritability, the normal ACE model ("NACE") and Falconer's method, can both be fit within this unified GEE2 framework, which additionally provides robust standard errors. Although the traditional Falconer's method cannot directly adjust for covariates, we show that the corresponding GEE2 version ("GEE2-Falconer") can incorporate covariate effects for both mean and variance-level parameters (e.g. let heritability vary by sex or age). Given non-normal data, we show that the GEE2 models attain significantly better coverage of the true heritability compared to the traditional NACE and Falconer's methods. Finally, we demonstrate an important scenario where the NACE model produces biased estimates of heritability while Falconer's method remains unbiased. Overall, we recommend using the robust and flexible GEE2-Falconer model for estimating heritability in twin studies.

stat.ME

Simultaneous Selection of Multiple Important Single Nucleotide Polymorphisms in Familial Genome Wide Association Studies Data

We propose a resampling-based fast variable selection technique for detecting relevant single nucleotide polymorphisms (SNP) in a multi-marker mixed effect model. Due to computational complexity, current practice primarily involves testing the effect of one SNP at a time, commonly termed as `single SNP association analysis'. Joint modeling of genetic variants within a gene or pathway may have better power to detect associated genetic variants, especially the ones with weak effects. In this paper, we propose a computationally efficient model selection approach -- based on the e-values framework -- for single SNP detection in families while utilizing information on multiple SNPs simultaneously. To overcome computational bottleneck of traditional model selection methods, our method trains one single model, and utilizes a fast and scalable bootstrap procedure. We illustrate through numerical studies that our proposed method is more effective in detecting SNPs associated with a trait than either single-marker analysis using family data or model selection methods that ignore the familial dependency structure. Further, we perform gene-level analysis in Minnesota Center for Twin and Family Research (MCTFR) dataset using our method to detect several SNPs using this that have been implicated to be associated with alcohol consumption.

stat.AP

USAT: A Unified Score-based Association Test for Multiple Phenotype-Genotype Analysis

Genome-wide Association Studies (GWASs) for complex diseases often collect data on multiple correlated endo-phenotypes. Multivariate analysis of these correlated phenotypes can improve the power to detect genetic variants. Multivariate analysis of variance (MANOVA) can perform such association analysis at a GWAS level, but the behavior of MANOVA under different trait models has not been carefully investigated. In this paper, we show that MANOVA is generally very powerful for detecting association but there are situations, such as when a genetic variant is associated with all the traits, where MANOVA may not have any detection power. We investigate the behavior of MANOVA, both theoretically and using simulations, and derive the conditions where MANOVA loses power. Based on our findings, we propose a unified score-based test statistic USAT that can perform better than MANOVA in such situations and nearly as well as MANOVA elsewhere. Our proposed test reports an approximate asymptotic p-value for association and is computationally very efficient to implement at a GWAS level. We have studied through extensive simulation the performance of USAT, MANOVA and other existing approaches and demonstrated the advantage of using the USAT approach to detect association between a genetic variant and multivariate phenotypes. We applied USAT to data from three correlated traits collected on 5,816 Caucasian individuals from the Atherosclerosis Risk in Communities (ARIC) Study and detected some interesting associations.

stat.ME