SearcharxivSearch

arXiv subjects

Wanyi Ling

Publications and source records attributed to Wanyi Ling.

3 recordsLinked to original sources

How does limma-trend work? An empirical partially Bayes perspective

In high-throughput biology, it is common to fit thousands of linear regressions -- one per gene, protein, or other unit -- with very few samples per unit. Limma-trend, one of the most widely used methods in this setting, improves power by shrinking variance estimates parametrically toward a fitted curve (the trend) relating variance to a unit-level summary (e.g., average intensity, peptide count), before computing p-values and applying the Benjamini-Hochberg procedure to control the false discovery rate (FDR). We study limma-trend through the lens of empirical partially Bayes inference, a paradigm in which a prior is posited and estimated for the nuisance parameters while parameters of interest remain fixed. From this perspective, limma-trend computes approximate partially Bayes p-values that condition on the residual sample variance and the unit-level summary. The same framework explains why MAnorm2, a popular variant for ChIP-seq, can sometimes fail to control FDR. We then derive a nonparametric generalization of limma-trend that estimates the residual variance prior using nonparametric maximum likelihood. Under dense signals, this procedure asymptotically controls the FDR -- even when the trend is misspecified or inconsistently estimated. To allow the full shape of the conditional variance distribution to depend on the unit-level summary, we develop a second procedure that learns it directly.

stat.ME

Empirical Bayes Rebiasing

We study methods for simultaneous analysis of many noisy and biased estimates, each paired with an even noisier estimate of its own bias. The analyst's goal is to construct short calibrated intervals for each parameter. The standard debiasing approach, which subtracts the bias estimate from each biased estimate, inflates variance and yields long intervals. In this paper, we propose an empirical Bayes rebiasing strategy that starts from the fully debiased estimates and learns from data how much bias to reintroduce by estimating the unknown bias distribution. We provide convergence rates for the coverage of our intervals when the bias distribution is estimated using nonparametric maximum likelihood. Furthermore, we demonstrate substantial precision gains in prediction-powered inference, including pairwise LLM win-rate evaluations, as well as for inference of direct genetic effects in family-based GWAS.

stat.ME

Empirical partially Bayes two sample testing

A common task in high-throughput biology is to test for differences in means between two samples across thousands of features (e.g., genes or proteins), often with only a handful of replicates per sample. Moderated t-tests handle this problem by assuming normality and equal variances, and by applying the empirical partially Bayes principle: a prior is posited and estimated for the nuisance parameters (variances) but not for the primary parameters (means). This approach has been highly successful in genomics, yet the equal variance assumption is often violated in practice. Meanwhile, Welch's unequal variance t-test with few replicates suffers from inflated type-I error and low power. Taking inspiration from moderated t-tests, we extend the empirical partially Bayes paradigm to two-sample testing with unequal variances. We develop two procedures: one that models the ratio of the two sample-specific variances and another that models the two variances jointly, with prior distributions estimated by nonparametric maximum likelihood. Our empirical partially Bayes methods yield p-values that are asymptotically uniform as the number of features grows while the number of replicates remains fixed, ensuring asymptotic type-I error control. Simulations and applications to genomic data demonstrate substantial gains in power.

stat.ME