SearcharxivSearch

arXiv subjects

Toby Kenney

Publications and source records attributed to Toby Kenney.

14 recordsLinked to original sources

Hypergraph Variable Selection with False Discovery Rate Control

Variable selection methods that control the false discovery rate often lose power when predictors exhibit complex dependence structures. We previously showed that selecting hierarchically clustered groups of predictors can mitigate this issue while maintaining false discovery rate control. When correlations are less structured, however, overlapping predictor sets may be more effective. We introduce a generalized false discovery rate for hypotheses defined on sets of predictors and propose a hypergraph-based selection method. This approach achieves higher power across diverse settings while preserving rigorous false discovery rate control.

stat.ME

Factor State Space Modelling of the Ornstein-Uhlenbeck Process with Measurement Error and its Application

Standard Ornstein-Uhlenbeck (OU) models often yield biased parameter estimates when measurement error is ignored. While the Ornstein-Uhlenbeck State Space Model (OUSSM) addresses this in univariate settings, multidimensional extensions remain limited. This paper introduces the factor OUSSM to model multi-dimensional, mean-reverting systems with observational noise. We resolve critical identifiability challenges in parameter estimation by establishing necessary constraints and validating the method through extensive simulations. We demonstrate the model's versatility by analyzing human gut microbiome dynamics and North Atlantic Sea Surface Temperature (SST) data. The results reveal distinct latent temporal structures in both biological and environmental systems, establishing the factor OUSSM as a robust framework for multivariate time series analysis.

stat.AP

Setwise Hierarchical Variable Selection and the Generalized Linear Step-Up Procedure for False Discovery Rate Control

Controlling the false discovery rate (FDR) in variable selection becomes challenging when predictors are correlated, as existing methods often exclude all members of correlated groups and consequently perform poorly for prediction. We introduce a new setwise variable-selection framework that identifies clusters of potential predictors rather than forcing selection of a single variable. By allowing any member of a selected set to serve as a surrogate predictor, our approach supports strong predictive performance while maintaining rigorous FDR control. We construct sets via hierarchical clustering of predictors based on correlation, then test whether each set contains any non-null effects. Similar clustering and setwise selection have been applied in the familywise error rate (FWER) control regime, but previous research has been unable to overcome the inherent challenges of extending this to the FDR control framework. To control the FDR, we develop substantial generalizations of linear step-up procedures, extending the Benjamini-Hochberg and Benjamini-Yekutieli methods to accommodate the logical dependencies among these composite hypotheses. We prove that these procedures control the FDR at the nominal level and highlight their broader applicability. Simulation studies and real-data analyses show that our methods achieve higher power than existing approaches while preserving FDR control, yielding more informative variable selections and improved predictive models.

stat.ME

Detection of evolutionary shifts in variance under an Ornsten-Uhlenbeck model

Sudden changes in environmental conditions can lead to evolutionary shifts not only in the optimal trait value, but also in the diffusion variance under the Ornstein-Uhlenbeck (OU) model. While several methods have been developed to detect shifts in optimal values, few explicitly account for concurrent shifts in both evolutionary variance and diffusion variance. We use a multi-optima and multi-variance OU model to describe trait evolution with shifts in both optimal value and diffusion variance and analyze how covariance between species is affected when shifts in variance occur along the phylogeny. We propose a new method that simultaneously detects shifts in both variance and optimal values by formulating the problem as a variable selection task using an L1-penalized loss function. Our method is implemented in the R package ShiVa (Detection of evolutionary shifts in variance). Through simulations, we compare ShiVa with existing methods that can automatically detect evolutionary shifts under the OU model (l1ou, PhylogeneticEM, and PCMFit). Our method demonstrates improved predictive ability and significantly reduces false positives in detecting optimal value shifts when variance shifts are present. When only shifts in optimal value occur, our method performs comparably to existing approaches. We apply ShiVa to empirical data on floral diameter in Euphorbiaceae and buccal morphology in Centrarchidae sunfishes.

q-bio.PE

Rank Selection for Non-negative Matrix Factorization

Non-Negative Matrix Factorization (NMF) is a widely used dimension reduction method that factorizes a non-negative data matrix into two lower dimensional non-negative matrices: One is the basis or feature matrix which consists of the variables and the other is the coefficients matrix which is the projections of data points to the new basis. The features can be interpreted as sub-structures of the data. The number of sub-structures in the feature matrix is also called the rank which is the only tuning parameter in NMF. An appropriate rank will extract the key latent features while minimizing the noise from the original data. In this paper, we develop a novel rank selection method based on hypothesis testing, using a deconvolved bootstrap distribution to assess the significance level accurately despite the large amount of optimization error. In the simulation section, we compare our method with a rank selection method based on hypothesis testing using bootstrap distribution without deconvolution, and with a cross-validated imputation method1. Through simulations, we demonstrate that our method is not only accurate at estimating the true ranks for NMF especially when the features are hard to distinguish but also efficient at computation. When applied to real microbiome data (e.g. OTU data and functional metagenomic data), our method also shows the ability to extract interpretable sub-communities in the data.

stat.AP

Evolutionary shift detection with ensemble variable selection

1. Abrupt environmental changes can lead to evolutionary shifts in trait evolution. Identifying these shifts is an important step in understanding the evolutionary history of phenotypes. 2. We propose an ensemble variable selection method (R package ELPASO) for the evolutionary shift detection task and compare it with existing methods (R packages l1ou and PhylogeneticEM) under several scenarios. 3. The performances of methods are highly dependent on the selection criterion. When the signal sizes are small, the methods using the Bayesian information criterion (BIC) have better performances. And when the signal sizes are large enough, the methods using the phylogenetic Bayesian information criterion (pBIC) (Khabbazian et al., 2016) have better performance. Moreover, the performance is heavily impacted by measurement error and tree reconstruction error. 4. Ensemble method + pBIC tends to perform less conservatively than l1ou + pBIC, and Ensemble method + BIC is more conservatively than l1ou + BIC. PhylogeneticEM is even more conservative with small signal sizes and falls between l1ou + pBIC and Ensemble method + BIC with large signal sizes. The results can differ between the methods, but none clearly outperforms the others. By applying multiple methods to a single dataset, we can access the robustness of each detected shift, based on the agreement among methods.

q-bio.PE

Stone Duality for Topological Convexity Spaces

A convexity space is a set X with a chosen family of subsets (called convex subsets) that is closed under arbitrary intersections and directed unions. There is a lot of interest in spaces that have both a convexity space and a topological space structure. In this paper, we study the category of topological convexity spaces and extend the Stone duality between coframes and topological spaces to an adjunction between topological convexity spaces and sup-lattices. We factor this adjunction through the category of preconvexity spaces (somtimes called closure spaces).

math.CT

Deconvolution density estimation with penalised MLE

Deconvolution is the important problem of estimating the distribution of a quantity of interest from a sample with additive measurement error. Nearly all methods in the literature are based on Fourier transformation because it is mathematically a very neat solution. However, in practice these methods are unstable, and produce bad estimates when signal-noise ratio or sample size are low. In this paper, we develop a new deconvolution method based on maximum likelihood with a smoothness penalty. We show that our new method has much better performance than existing methods, particularly for small sample size or signal-noise ratio.

stat.ME

Stochastic Generalized Lotka-Volterra Model with An Application to Learning Microbial Community Structures

Inferring microbial community structure based on temporal metagenomics data is an important goal in microbiome studies. The deterministic generalized Lotka-Volterra differential (GLV) equations have been used to model the dynamics of microbial data. However, these approaches fail to take random environmental fluctuations into account, which may negatively impact the estimates. We propose a new stochastic GLV (SGLV) differential equation model, where the random perturbations of Brownian motion in the model can naturally account for the external environmental effects on the microbial community. We establish new conditions and show various mathematical properties of the solutions including general existence and uniqueness, stationary distribution, and ergodicity. We further develop approximate maximum likelihood estimators based on discrete observations and systematically investigate the consistency and asymptotic normality of the proposed estimators. Our method is demonstrated through simulation studies and an application to the well-known "moving picture" temporal microbial dataset.

stat.ME

SuRF: a New Method for Sparse Variable Selection, with Application in Microbiome Data Analysis

In this paper, we present a new variable selection method for regression and classification purposes. Our method, called Subsampling Ranking Forward selection (SuRF), is based on LASSO penalised regression, subsampling and forward-selection methods. SuRF offers major advantages over existing variable selection methods in terms of both sparsity of selected models and model inference. We provide an R package that can implement our method for generalized linear models. We apply our method to classification problems from microbiome data, using a novel agglomeration approach to deal with the special tree-like correlation structure of the variables. Existing methods arbitrarily choose a taxonomic level a priori before performing the analysis, whereas by combining SuRF with these aggregated variables, we are able to identify the key biomarkers at the appropriate taxonomic level, as suggested by the data. We present simulations in multiple sparse settings to demonstrate that our approach performs better than several other popularly used existing approaches in recovering the true variables. We apply SuRF to two microbiome data sets: one about prediction of pouchitis and another for identifying samples from two healthy individuals. We find that SuRF can provide a better or comparable prediction with other methods while controlling the false positive rate of variable selection.

stat.ME

Consistency of Ranking Estimators

The ranking problem is to order a collection of units by some unobserved parameter, based on observations from the associated distribution. This problem arises naturally in a number of contexts, such as business, where we may want to rank potential projects by profitability; or science, where we may want to rank predictors potentially associated with some trait by the strength of the association. This approach provides a valuable alternative to the sparsity framework often used with big data. Most approaches to this problem are empirical Bayesian, where we use the data to estimate the hyperparameters of the prior distribution, then use that distribution to estimate the unobserved parameter values. There are a number of different approaches to this problem, based on different loss functions for mis-ranking units. Despite the number of papers developing methods for this problem, there is no work on the consistency of these methods. In this paper, we develop a general framework for consistency of empirical Bayesian ranking methods, which includes nearly all commonly used methods. We then determine conditions under which consistency holds. Given that little work has been done on selection of prior distribution, and that the loss functions developed are not strongly motivated, we consider the case where both of these are misspecified. We show that provided the loss function is reasonable; the prior distribution is not too light-tailed; and the error in measuring each unit converges to zero at a fast enough rate compared with the number of units (which is assumed to increase to infinity); all ranking methods are consistent.

math.ST

Poisson PCA: Poisson Measurement Error corrected PCA, with Application to Microbiome Data

In this paper, we study the problem of computing a Principal Component Analysis of data affected by Poisson noise. We assume samples are drawn from independent Poisson distributions. We want to estimate principle components of a fixed transformation of the latent Poisson means. Our motivating example is microbiome data, though the methods apply to many other situations. We develop a semiparametric approach to correct the bias of variance estimators, both for untransformed and transformed (with particular attention to log-transformation) Poisson means. Furthermore, we incorporate methods for correcting different exposure or sequencing depth in the data. In addition to identifying the principal components, we also address the non-trivial problem of computing the principal scores in this semiparametric framework. Most previous approaches tend to take a more parametric line. For example the Poisson-log-normal (PLN) model, approach. We compare our method with the PLN approach and find that our method is better at identifying the main principal components of the latent log-transformed Poisson means, and as a further major advantage, takes far less time to compute. Comparing methods on real data, we see that our method also appears to be more robust to outliers than the parametric method.

stat.ME

Prior Distributions for Ranking Problems

The ranking problem is to order a collection of units by some unobserved parameter, based on observations from the associated distribution. This problem arises naturally in a number of contexts, such as business, where we may want to rank potential projects by profitability; or science, where we may want to rank variables potentially associated with some trait by the strength of the association. Most approaches to this problem are empirical Bayesian, where we use the data to estimate the hyperparameters of the prior distribution, then use that distribution to estimate the unobserved parameter values. There are a number of different approaches to this problem, based on different loss functions for mis-ranking units. However, little has been done on the choice of prior distribution. Typical approaches involve choosing a conjugate prior for convenience, and estimating the hyperparameters by MLE from the whole dataset. In this paper, we look in more detail at the effect of choice of prior distribution on Bayesian ranking. We focus on the use of posterior mean for ranking, but many of our conclusions should apply to other ranking criteria, and it is not too difficult to adapt our methods to other choices of prior distributions.

stat.ME

The Adequate Bootstrap

There is a fundamental disconnect between what is tested in a model adequacy test, and what we would like to test. The usual approach is to test the null hypothesis "Model M is the true model." However, Model M is never the true model. A model might still be useful even if we have enough data to reject it. In this paper, we present a technique to assess the adequacy of a model from the philosophical standpoint that we know the model is not true, but we want to know if it is useful. Our solution to this problem is to measure the parameter uncertainty in our estimates caused by the model uncertainty. We use bootstrap inference on samples of a smaller size, for which the model cannot be rejected. We use a model adequacy test to choose a bootstrap size with limited probability of rejecting the model and perform inference for samples of this size based on a nonparametric bootstrap. Our idea is that if we base our inference on a sample size at which we do not reject the model, then we should be happy with this inference, because we would have been confident in it if our original dataset had been this size.

stat.ME