SearcharxivSearch

arXiv subjects

Amy D Willis

Publications and source records attributed to Amy D Willis.

6 recordsLinked to original sources

Consensus Tree Estimation with False Discovery Control via Partially Ordered Sets

Trees are data objects that hierarchically organize categories. Collections of trees arise in a diverse variety of fields, including evolutionary biology, machine learning, social sciences and anatomy. Summarizing a collection of trees by a single representative is challenging, in part due to the dimensions of both the sample space and the parameter space. We frame consensus tree estimation as a structured feature-selection problem, where leaves and edges are the features. We introduce a partial order on trees, use it to define false discoveries for a candidate summary tree, and develop novel estimation algorithms that control the false discovery rate at a nominal level for a broad class of generative models. We also use the partial order to assess the stability of features in a selected tree. Importantly, our method accommodates unequal leaf sets and non-binary trees, which commonly arise in modern datasets. Our feature-selection perspective yields finite-sample and model-free guarantees and provides a foundation for integrating multiple testing tools into tree estimation. We apply the method to study the origins of complex life. Our estimated tree recapitulates well-known divisions but highlights that there is insufficient data to determine the most recent archaeal ancestor of eukaryotic life.

stat.ME

Nonparametric Identification and Estimation of Ratios of Multi-Category Means under Preferential Sampling

Multi-category data arise in diverse fields including marketing, chemistry, public policy, genomics, political science, and ecology. We consider the problem of estimating ratios of category-specific means in a fully nonparametric setting, allowing for both observational units and categories to be preferentially sampled. We consider covariate-adjusted and unadjusted estimands that are non-parametrically defined and straightforward to interpret. While identifiability for related models has been established through parametric distributions or restrictions on the conditional mean (e.g., log-linearity), we show that identifiability can be obtained through an independence assumption or a category constraint, such as a reference category or a centering function. We develop an efficient, doubly-robust targeted minimum loss based estimator with excellent finite-sample performance, including in the setting of a large number of infrequently observed categories. We contrast the performance of our method with related approaches via simulation, and apply it to identify bacteria that are differentially abundant in diarrheal cases compared to controls. Our work provides a general framework for studying parameter identifiability in compositional data settings without requiring parametric assumptions on the data distribution.

stat.ME

Geometry of the space of phylogenetic trees with non-identical leaves

Phylogenetic trees summarize evolutionary relationships. The Billera-Holmes-Vogtmann (BHV) space for comparing phylogenetic trees has many elegant mathematical properties, but it does not encompass trees with differing leaf sets. To overcome this, we introduce Towering space: a complete metric space that extends BHV space to trees with non-identical leaf sets. Towering space is a structured collection of BHV spaces connected via pruning and regrafting operations. We study the geometry of paths in Towering space and present an algorithm for computing metric distances. By addressing a major limitation of BHV space, Towering space facilitates the analysis of modern phylogenetic datasets such as multi-domain gene trees.

q-bio.PE

Modeling complex measurement error in microbiome experiments to estimate relative abundances and detection effects

Accurate estimates of microbial species abundances are needed to advance our understanding of the role that microbiomes play in human and environmental health. However, artificially constructed microbiomes demonstrate that intuitive estimators of microbial relative abundances are biased. To address this, we propose a semiparametric method to estimate relative abundances, species detection effects, and/or cross-sample contamination in microbiome experiments. We show that certain experimental designs result in identifiable model parameters, and we present consistent estimators and asymptotically valid inference procedures. Notably, our procedure can estimate relative abundances on the boundary of the simplex. We demonstrate the utility of the method for comparing experimental protocols, removing cross-sample contamination, and estimating species' detectability.

stat.ME

Estimating Fold Changes from Partially Observed Outcomes with Applications in Microbial Metagenomics

We consider the problem of estimating fold-changes in the expected value of a multivariate outcome observed with unknown sample-specific and category-specific perturbations. This challenge arises in high-throughput sequencing studies of the abundance of microbial taxa because microbes are systematically over- and under-detected relative to their true abundances. Our model admits a partially identifiable estimand, and we establish full identifiability by imposing interpretable parameter constraints. To reduce bias and guarantee the existence of estimators in the presence of sparse observations, we apply an asymptotically negligible and constraint-invariant penalty to our estimating function. We develop a fast coordinate descent algorithm for estimation, and an augmented Lagrangian algorithm for estimation under null hypotheses. We construct a model-robust score test and demonstrate valid inference even for small sample sizes and violated distributional assumptions. The flexibility of the approach and comparisons to related methods are illustrated through a meta-analysis of microbial associations with colorectal cancer.

stat.ME

Distances between Extension Spaces of Phylogenetic Trees

Phylogenetic trees summarize evolutionary relationships between organisms, and tools to analyze collections of phylogenetic trees enable contrasts between different genes' ancestry. The BHV metric space has enabled the analysis of collections of trees that share a common set of leaves, but many genes are not shared, even between closely related species. BHV extension spaces represent trees with non-identical leaf sets in a common BHV space, but limited analytical tools exist for extension spaces. We define the distance between two phylogenetic trees with non-identical leaf sets as the shortest BHV distance between their extension spaces, and develop a reduced gradient algorithm to compute this distance. We study the scalability of our algorithm and apply it to analyze gene trees spanning multiple domains of life. Our distance and algorithm offer a fully general, interpretable approach to analyzing both ancient and recent evolutionary divergence.

q-bio.QM