SearcharxivSearch

arXiv subjects

Ruth M. Pfeiffer

Publications and source records attributed to Ruth M. Pfeiffer.

3 recordsLinked to original sources

Incorporating survival data into case-control studies with incident and prevalent cases

Typically, case-control studies to estimate odds-ratios associating risk factors with disease incidence from logistic regression only include cases with newly diagnosed disease. Recently proposed methods allow incorporating information on prevalent cases, individuals who survived from disease diagnosis to sampling, into cross-sectionally sampled case-control studies under parametric assumptions for the survival time after diagnosis. Here we propose and study methods to additionally use prospectively observed survival times from prevalent and incident cases to adjust logistic models for the time between disease diagnosis and sampling, the backward time, for prevalent cases. This adjustment yields unbiased odds-ratio estimates from case-control studies that include prevalent cases. We propose a computationally simple two-step generalized method-of-moments estimation procedure. First, we estimate the survival distribution based on a semi-parametric Cox model using an expectation-maximization algorithm that yields fully efficient estimates and accommodates left truncation for the prevalent cases and right censoring. Then, we use the estimated survival distribution in an extension of the logistic model to three groups (controls, incident and prevalent cases), to accommodate the survival bias in prevalent cases. In simulations, when the amount of censoring was modest, odds-ratios from the two-step procedure were equally efficient as those estimated by jointly optimizing the logistic and survival data likelihoods under parametric assumptions. Even with 90% censoring they were as efficient as estimates obtained using only cross-sectionally available information under parametric assumptions. This indicates that utilizing prospective survival data from the cases lessens model dependency and improves precision of association estimates for case-control studies with prevalent cases.

stat.ME

Subset Testing and Analysis of Multiple Phenotypes (STAMP)

Meta-analysis of multiple genome-wide association studies (GWAS) is effective for detecting single or multi marker associations with complex traits. We develop a flexible procedure ("STAMP") based on mixture models to perform region based meta-analysis of different phenotypes using data from different GWAS and identify subsets of associated phenotypes. Our model framework helps distinguish true associations from between-study heterogeneity. As a measure of association we compute for each phenotype the posterior probability that the genetic region under investigation is truly associated. Extensive simulations show that STAMP is more powerful than standard approaches for meta analyses when the proportion of truly associated outcomes is $\leq$ 50\%. For other settings, the power of STAMP is similar to that of existing methods. We illustrate our method on two examples, the association of a region on chromosome 9p21 with risk of fourteen cancers, and the associations of expression of quantitative traits loci (eQTLs) from two genetic regions with their cis-SNPs measured in seventeen tissue types using data from The Cancer Genome Atlas (TCGA).

stat.ME

On Combining Data From Genome-Wide Association Studies to Discover Disease-Associated SNPs

Combining data from several case-control genome-wide association (GWA) studies can yield greater efficiency for detecting associations of disease with single nucleotide polymorphisms (SNPs) than separate analyses of the component studies. We compared several procedures to combine GWA study data both in terms of the power to detect a disease-associated SNP while controlling the genome-wide significance level, and in terms of the detection probability ($\mathit{DP}$). The $\mathit{DP}$ is the probability that a particular disease-associated SNP will be among the $T$ most promising SNPs selected on the basis of low $p$-values. We studied both fixed effects and random effects models in which associations varied across studies. In settings of practical relevance, meta-analytic approaches that focus on a single degree of freedom had higher power and $\mathit{DP}$ than global tests such as summing chi-square test-statistics across studies, Fisher's combination of $p$-values, and forming a combined list of the best SNPs from within each study.

stat.ME