SearcharxivSearch

arXiv subjects

Hon-Cheong So

Publications and source records attributed to Hon-Cheong So.

7 recordsLinked to original sources

SumVg: Total heritability explained by all variants in genome-wide association studies based on summary statistics with standard error estimates

Genome-wide association studies (GWAS) are commonly employed to study the genetic basis of complex traits and diseases, and a key question is how much heritability could be explained by all variants in GWAS. One widely used approach that relies on summary statistics only is LD score regression (LDSC), however the approach requires certain assumptions on the SNP effects (all SNPs contribute to heritability and each SNP contributes equal variance). More flexible modeling methods may be useful. We previously developed an approach recovering the true z-statistics from a set of observed z-statistics with an empirical Bayes approach, using only summary statistics. However, methods for standard error (SE) estimation are not available yet, limiting the interpretation of results and applicability of the approach. In this study we developed several resampling-based approaches to estimate the SE of SNP-based heritability, including two jackknife and three parametric bootstrap methods. Simulations showed that delete-d-jackknife and parametric bootstrap approaches provide good estimates of the SE. Particularly, the parametric bootstrap approaches yield the lowest root-mean-squared-error (RMSE) of the true SE. In addition, we applied our method to estimate SNP-based heritability of 12 immune-related traits (levels of cytokines and growth factors) to shed light on their genetic architecture. We also implemented the methods to compute the sum of heritability explained and the corresponding SE in an R package SumVg, available at https://github.com/lab-hcso/Estimating-SE-of-total-heritability/ . In conclusion, SumVg may provide a useful alternative tool for SNP heritability and SE estimates, which does not rely on distributional assumptions of SNP effects.

q-bio.GN

Analysis of genetic differences between psychiatric disorders: Exploring pathways and cell-types/tissues involved and ability to differentiate the disorders by polygenic scores

Although displaying genetic correlations, psychiatric disorders are clinically defined as categorical entities as they each have distinguishing clinical features and may involve different treatments. Identifying differential genetic variations between these disorders may reveal how the disorders differ biologically and help to guide more personalized treatment. Here we presented a comprehensive analysis to identify genetic markers differentially associated with various psychiatric disorders/traits based on GWAS summary statistics, covering 18 psychiatric traits/disorders and 26 comparisons. We also conducted comprehensive analysis to unravel the genes, pathways and SNP functional categories involved, and the cell types and tissues implicated. We also assessed how well one could distinguish between psychiatric disorders by polygenic risk scores (PRS). SNP-based heritabilities (h2SNP) were significantly larger than zero for most comparisons. Based on current GWAS data, PRS have mostly modest power to distinguish between psychiatric disorders. For example, we estimated that AUC for distinguishing schizophrenia from major depressive disorder (MDD), bipolar disorder (BPD) from MDD and schizophrenia from BPD were 0.694, 0.602 and 0.618 respectively, while the maximum AUC (based on h2SNP) were 0.763, 0.749 and 0.726 respectively. We also uncovered differences in each pair of studied traits in terms of their differences in genetic correlation with comorbid traits. For example, clinically-defined MDD appeared to more strongly genetically correlated with other psychiatric disorders and heart disease, when compared to non-clinically-defined depression in UK Biobank. Our findings highlight genetic differences between psychiatric disorders and the mechanisms involved. PRS may aid differential diagnosis of selected psychiatric disorders in the future with larger GWAS samples.

q-bio.GN

A framework to decipher the genetic architecture of combinations of complex diseases: applications in cardiovascular medicine

Genome-wide association studies(GWAS) have proven to be highly useful in revealing the genetic basis of complex diseases. At present, most GWAS are studies of a particular single disease diagnosis against controls. However, in practice, an individual is often affected by more than one condition/disorder. For example, patients with coronary artery disease(CAD) are often comorbid with diabetes mellitus(DM). Along a similar line, it is often clinically meaningful to study patients with one disease but without a comorbidity. For example, obese DM may have different pathophysiology from non-obese DM. Here we developed a statistical framework to uncover susceptibility variants for comorbid disorders (or a disorder without comorbidity), using GWAS summary statistics only. In essence, we mimicked a case-control GWAS in which the cases are affected with comorbidities or a disease without a relevant comorbid condition (in either case, we may consider the cases as those affected by a specific subtype of disease, as characterized by the presence or absence of comorbid conditions). We extended our methodology to deal with continuous traits with clinically meaningful categories (e.g. lipids). In addition, we illustrated how the analytic framework may be extended to more than two traits. We verified the feasibility and validity of our method by applying it to simulated scenarios and four cardiometabolic (CM) traits. We also analyzed the genes, pathways, cell-types/tissues involved in CM disease subtypes. LD-score regression analysis revealed some subtypes may indeed be biologically distinct with low genetic correlations. Further Mendelian randomization analysis found differential causal effects of different subtypes to relevant complications. We believe the findings are of both scientific and clinical value, and the proposed method may open a new avenue to analyzing GWAS data.

q-bio.GN

Turning genome-wide association study findings into opportunities for drug repositioning

Drug development is a very costly and lengthy process, while repositioned or repurposed drugs could be brought into clinical practice within a shorter time-frame and at a much reduced cost. The past decade has observed a massive growth in the amount of data from genome-wide association studies (GWAS). The rich information contained in GWAS data has great potential to guide drug discovery or repositioning. Here we provide an overview of different computational approaches which employ GWAS data to guide drug repositioning. These methods include selection of top candidate genes from GWAS as drug targets, deducing drug candidates based on drug-drug and disease-disease similarity, searching for reversed expression profiles between drugs and diseases, pathway-based methods as well as repositioning based on analysis of biological networks. Each method is illustrated with examples, and their respective strengths and limitations are discussed. Finally we discussed several areas for future research.

q-bio.GN

A machine learning approach to drug repositioning based on drug expression profiles: Applications to schizophrenia and depression/anxiety disorders

Development of new medications is a very lengthy and costly process. Finding novel indications for existing drugs, or drug repositioning, can serve as a useful strategy to shorten the development cycle. In this study, we present an approach to drug discovery or repositioning by predicting indication for a particular disease based on expression profiles of drugs, with a focus on applications in psychiatry. Drugs that are not originally indicated for the disease but with high predicted probabilities serve as good candidates for repurposing. This framework is widely applicable to any chemicals or drugs with expression profiles measured, even if the drug targets are unknown. It is also highly flexible as virtually any supervised learning algorithms can be used. We applied this approach to identify repositioning opportunities for schizophrenia as well as depression and anxiety disorders. We applied various state-of-the-art machine learning (ML) approaches for prediction, including deep neural networks, support vector machines (SVM), elastic net, random forest and gradient boosted machines. The performance of the five approaches did not differ substantially, with SVM slightly outperformed the others. However, methods with lower predictive accuracy can still reveal literature-supported candidates that are of different mechanisms of actions. As a further validation, we showed that the repositioning hits are enriched for psychiatric medications considered in clinical trials. Notably, many top repositioning hits are supported by previous preclinical or clinical studies. Finally, we propose that ML approaches may provide a new avenue to explore drug mechanisms via examining the variable importance of gene features.

q-bio.GN

Exploring shared genetic bases and causal relationships of schizophrenia and bipolar disorder with 28 cardiovascular and metabolic traits

Cardiovascular diseases (CVD) represent a major health issue in patients with schizophrneia (SCZ) and bipolar disorder (BD), but the exact nature of cardiometabolic (CM) abnormalities involved and the underlying mechanisms remain unclear. Using polygenic risk scores (PRS) and LD score regression, we investigated the shared genetic bases of SCZ and BD with a panel of 28 cardiometabolic traits. We performed Mendelian randomization (MR) to elucidate casual relationships between the two groups of disorders. The analysis was based on large-scale meta-analyses of genome-wide association studies (GWAS). We also identified the potential shared genetic variants by a statistical approach based on local true discovery rates, and inferred the pathways involved. We found polygenic associations of SCZ with glucose metabolism abnormalities, adverse adipokine profiles, increased wait-hip ratio and raised visceral adiposity. However, BMI showed inverse genetic correlation and polygenic link with SCZ. On the other hand, we observed polygenic associations with an overall favorable CM profile in BD. MR analysis showed that SCZ may be causally linked to raised triglyceride and that lower fasting glucose may be linked to BD; otherwise MR did not reveal other significant causal relationships in general. We also identified numerous SNPs and pathways shared between SCZ/BD with cardiometabolic traits, some of which are related to inflammation or the immune system. In conclusion, SCZ patients may be genetically associated with several CM abnormalities independent of medication side-effects, and proper surveillance and management of CV risk factors may be required from the onset of the disease. On the other hand, CM abnormalities in BD are more likely to be secondary.

q-bio.GN

Epigenome-wide association study and integrative analysis with the transcriptome based on GWAS summary statistics

The past decade has seen a rapid growth in omics technologies. Genome-wide association studies (GWAS) have uncovered susceptibility variants for a variety of complex traits. However, the functional significance of most discovered variants are still not fully understood. On the other hand, there is increasing interest in exploring the role of epigenetic variations such as DNA methylation in disease pathogenesis. In this work, we present a general framework for epigenome-wide association study and integrative analysis with the transcriptome based on GWAS summary statistics and data from methylation and expression quantitative trait loci (QTL) studies. The framework is based on Mendelian randomization, which is much less vulnerable to confounding and reverse causation compared to conventional studies. The framework was applied to five complex diseases. We first identified loci that are differentially methylated due to genetic variations, and then developed several approaches for joint testing with the GWAS-imputed transcriptome. We discovered a number of novel candidate genes that are not implicated in the original GWAS studies. We also observed strong evidence (lowest p = 2.01e-184) for differential expression among the top genes mapped to methylation loci. The framework proposed here opens a new way of analyzing GWAS summary data and will be useful for gaining deeper insight into disease mechanisms.

q-bio.GN