SearcharxivSearch

arXiv subjects

Jinko Graham

Publications and source records attributed to Jinko Graham.

4 recordsLinked to original sources

Log-F-penalized Conditional Logistic Regression for Sparse Data

We investigate penalized likelihood methods for estimation and inference in conditional logistic regression. The standard conditional maximum likelihood estimator is known to be biased away from zero in small or sparse matched case-control studies. A widely used remedy is Firth's penalized likelihood approach, which has good frequentist operating characteristics but provides limited control over the degree of shrinkage applied to individual regression coefficients. We develop point and interval estimators by penalizing the conditional likelihood with independent log-$F$ distributions. The log-\(F\)-penalized approach allows analysts to calibrate shrinkage using interpretable prior assumptions about plausible effect sizes. We also provide practical guidance for calibrating the amount of shrinkage and show that the method can be implemented through data augmentation using standard conditional logistic regression software. We illustrate the methods using data from (i) a study of maternal exposure to diethylstilbestrol and the risk of vaginal cancer in daughters, and (ii) a genetic association study of type 2 diabetes. We then compare the log-$F$-penalized approach with Firth's penalized likelihood method in a simulation study. In simulations, the log-$F$-penalized estimators had confidence-interval coverage comparable to that of Firth's method and lower mean squared error, with similar type~1 error rates and power. These results support the use of log-$F$-penalized conditional logistic regression for inference in sparse matched and stratified studies.

stat.ME

The Contribution Plot: Decomposition and Graphical Display of the RV Coefficient, with Application to Genetic and Brain Imaging Biomarkers of Alzheimer's Disease

Alzheimer's disease (AD) is a chronic neurodegenerative disease that causes memory loss and decline in cognitive abilities. AD is the sixth leading cause of death in the United States, affecting an estimated 5 million Americans. To assess the association between multiple genetic variants and multiple measurements of structural changes in the brain a recent study of AD used a multivariate measure of linear dependence, the RV coefficient. The authors decomposed the RV coefficient into contributions from individual variants and displayed these contributions graphically. We investigate the properties of such a `contribution plot' in terms of an underlying linear model, and discuss estimation of the components of the plot when the correlation signal may be sparse. The contribution plot is applied to simulated data and to genomic and brain imaging data from the Alzheimer's Disease Neuroimaging Initiative.

stat.OT

Simple Measures of Individual Cluster-Membership Certainty for Hard Partitional Clustering

We propose two probability-like measures of individual cluster-membership certainty which can be applied to a hard partition of the sample such as that obtained from the Partitioning Around Medoids (PAM) algorithm, hierarchical clustering or k-means clustering. One measure extends the individual silhouette widths and the other is obtained directly from the pairwise dissimilarities in the sample. Unlike the classic silhouette, however, the measures behave like probabilities and can be used to investigate an individual's tendency to belong to a cluster. We also suggest two possible ways to evaluate the hard partition. We evaluate the performance of both measures in individuals with ambiguous cluster membership, using simulated binary datasets that have been partitioned by the PAM algorithm or continuous datasets that have been partitioned by hierarchical clustering and k-means clustering. For comparison, we also present results from soft clustering algorithms such as soft analysis clustering (FANNY) and two model-based clustering methods. Our proposed measures perform comparably to the posterior-probability estimators from either FANNY or the model-based clustering methods. We also illustrate the proposed measures by applying them to Fisher's classic iris data set.

stat.AP

A Bayesian Group Sparse Multi-Task Regression Model for Imaging Genetics

Motivation: Recent advances in technology for brain imaging and high-throughput genotyping have motivated studies examining the influence of genetic variation on brain structure. Wang et al. (Bioinformatics, 2012) have developed an approach for the analysis of imaging genomic studies using penalized multi-task regression with regularization based on a novel group $l_{2,1}$-norm penalty which encourages structured sparsity at both the gene level and SNP level. While incorporating a number of useful features, the proposed method only furnishes a point estimate of the regression coefficients; techniques for conducting statistical inference are not provided. A new Bayesian method is proposed here to overcome this limitation. Results: We develop a Bayesian hierarchical modeling formulation where the posterior mode corresponds to the estimator proposed by Wang et al. (Bioinformatics, 2012), and an approach that allows for full posterior inference including the construction of interval estimates for the regression parameters. We show that the proposed hierarchical model can be expressed as a three-level Gaussian scale mixture and this representation facilitates the use of a Gibbs sampling algorithm for posterior simulation. Simulation studies demonstrate that the interval estimates obtained using our approach achieve adequate coverage probabilities that outperform those obtained from the nonparametric bootstrap. Our proposed methodology is applied to the analysis of neuroimaging and genetic data collected as part of the Alzheimer's Disease Neuroimaging Initiative (ADNI), and this analysis of the ADNI cohort demonstrates clearly the value added of incorporating interval estimation beyond only point estimation when relating SNPs to brain imaging endophenotypes.

stat.ME