Searcharxiv⌕ Search

arXiv subjects

Young Ju Suh

Publications and source records attributed to Young Ju Suh.

4 recordsLinked to original sources

Using Volcano Plots and Regularized-Chi Statistics in Genetic Association Studies

Labor intensive experiments are typically required to identify the causal disease variants from a list of disease associated variants in the genome. For designing such experiments, candidate variants are ranked by their strength of genetic association with the disease. However, the two commonly used measures of genetic association, the odds-ratio (OR) and p-value, may rank variants in different order. To integrate these two measures into a single analysis, here we transfer the volcano plot methodology from gene expression analysis to genetic association studies. In its original setting, volcano plots are scatter plots of fold-change and t-test statistic (or -log of the p-value), with the latter being more sensitive to sample size. In genetic association studies, the OR and Pearson's chi-square statistic (or equivalently its square root, chi; or the standardized log(OR)) can be analogously used in a volcano plot, allowing for their visual inspection. Moreover, the geometric interpretation of these plots leads to an intuitive method for filtering results by a combination of both OR and chi-square statistic, which we term "regularized-chi". This method selects associated markers by a smooth curve in the volcano plot instead of the right-angled lines which corresponds to independent cutoffs for OR and chi-square statistic. The regularized-chi incorporates relatively more signals from variants with lower minor-allele-frequencies than chi-square test statistic. As rare variants tend to have stronger functional effects, regularized-chi is better suited to the task of prioritization of candidate genes.

q-bio.QM↗

Exploring Case-Control Genetic Association Tests Using Phase Diagrams

Background: By a new concept called "phase diagram", we compare two commonly used genotype-based tests for case-control genetic analysis, one is a Cochran-Armitage trend test (CAT test at $x=0.5$, or CAT0.5) and another (called MAX2) is the maximization of two chi-square test results: one from the two-by-two genotype count table that combines the baseline homozygotes and heterozygotes, and another from the table that combines heterozygotes with risk homozygotes. CAT0.5 is more suitable for multiplicative disease models and MAX2 is better for dominant/recessive models. Methods: We define the CAT0.5-MAX2 phase diagram on the disease model space such that regions where MAX2 is more powerful than CAT0.5 are separated from regions where the CAT0.5 is more powerful, and the task is to choose the appropriate parameterization to make the separation possible. Results: We find that using the difference of allele frequencies ($δ_p$) and the difference of Hardy-Weinberg disequilibrium coefficients ($δ_ε$) can separate the two phases well, and the phase boundaries are determined by the angle $tan^{-1}(δ_p/δ_ε)$, which is an improvement over the disease model selection using $δ_ε$ only. Conclusions: We argue that phase diagrams similar to the one for CAT0.5-MAX2 have graphical appeals in understanding power performance of various tests, clarifying simulation schemes, summarizing case-control datasets, and guessing the possible mode of inheritance.

q-bio.PE↗

Genotype-based Case-Control Analysis, Violation of Hardy-Weinberg Equilibrium, and Phase Diagrams

We study in detail a particular statistical method in genetic case-control analysis, labeled "genotype-based association", in which the two test results from assuming dominant and recessive model are combined in one optimal output. This method differs both from the allele-based association which artificially doubles the sample size, and the direct chi-square test on 3-by-2 contingency table which may overestimate the degree of freedom. We conclude that the comparative advantage (or disadvantage) of the genotype-based test over the allele-based test mainly depends on two parameters, the allele frequency difference delta and the Hardy-Weinberg disequilibrium coefficient difference delta_epsilon. Six different situations, called "phases", characterized by the two X^2 test statistics in allele-based and genotype-based test, are well separated in the phase diagram parameterized by delta and delta_epsilon. For two major groups of phases, a single parameter theta = tan^-1 (delta/delta_epsilon) is able to achieves an almost perfect phase separation. We also applied the analytic result to several types of disease models. It is shown that for dominant and additive models, genotype-based tests are favored over allele-based tests.

q-bio.QM↗

Does Logarithm Transformation of Microarray Data Affect Ranking Order of Differentially Expressed Genes?

A common practice in microarray analysis is to transform the microarray raw data (light intensity) by a logarithmic transformation, and the justification for this transformation is to make the distribution more symmetric and Gaussian-like. Since this transformation is not universally practiced in all microarray analysis, we examined whether the discrepancy of this treatment of raw data affect the "high level" analysis result. In particular, whether the differentially expressed genes as obtained by $t$-test, regularized t-test, or logistic regression have altered rank orders due to presence or absence of the transformation. We show that as much as 20%--40% of significant genes are "discordant" (significant only in one form of the data and not in both), depending on the test being used and the threshold value for claiming significance. The t-test is more likely to be affected by logarithmic transformation than logistic regression, and regularized $t$-test more affected than t-test. On the other hand, the very top ranking genes (e.g. up to top 20--50 genes, depending on the test) are not affected by the logarithmic transformation.

q-bio.QM↗