SearcharxivSearch

arXiv subjects

Patrick J. Breheny

Publications and source records attributed to Patrick J. Breheny.

3 recordsLinked to original sources

A Novel Approach to Instrumental Variable Estimation: TEAM-IV

Instrumental-variable (IV) analyses can be undermined when some instruments violate the exclusion restriction through direct effects on the outcome. Existing robust IV methods, including sisVIVE and CIIV, rely on majority- or plurality-validity conditions. We propose TEAM-IV, which targets joint validity by identifying sets of instruments that appear valid together and aggregating them, allowing reliable estimation even when only a small number of candidate instruments are valid (potentially as few as two). TEAM-IV fits MCP-penalized models of the outcome on projected exposure over overlapping instrument subsets. Instruments whose estimated direct effects are frequently shrunk to zero together (suggesting concordant validity status) form candidate ``teams'' (instrument sets). Teams are ranked by their out-of-sample predictive performance; near-best teams are unioned and used as the instrument set for the IV estimate. Across simulations varying the number, magnitude, and signs of invalid direct effects, instrument correlation, and exposure endogeneity, TEAM-IV generally attains lower absolute estimation error than sisVIVE and CIIV. Relative gains are largest when invalid instruments are prevalent. TEAM-IV is demonstrated in an empirical Mendelian randomization application using the Multi-Ethnic Study of Atherosclerosis (MESA) concerning the effect of LDL cholesterol on carotid intima--media thickness. Two additional MESA-anchored validation studies demonstrate differences between TEAM-IV and sisVIVE under challenging invalid-instrument configurations.

stat.ME

Cross-Validation in Penalized Linear Mixed Models: Addressing Common Implementation Pitfalls

In this paper, we develop an implementation of cross-validation for penalized linear mixed models. While these models have been proposed for correlated high-dimensional data, the current literature implicitly assumes that tuning parameter selection procedures developed for independent data will also work well in this context. We argue that such naive assumptions make analysis prone to pitfalls, several of which we will describe. Here we present a correct implementation of cross-validation for penalized linear mixed models, addressing these common pitfalls. We support our methods with mathematical proof, simulation study, and real data analysis.

stat.ME

plmmr: an R package to fit penalized linear mixed models for genome-wide association data with complex correlation structure

Correlation among the observations in high-dimensional regression modeling can be a major source of confounding. We present a new open-source package, plmmr, to implement penalized linear mixed models in R. This R package estimates correlation among observations in high-dimensional data and uses those estimates to improve prediction with the best linear unbiased predictor. The package uses memory-mapping so that genome-scale data can be analyzed on ordinary machines even if the size of data exceeds RAM. We present here the methods, workflow, and file-backing approach upon which plmmr is built, and we demonstrate its computational capabilities with two examples from real GWAS data.

stat.CO