SearcharxivSearch

arXiv subjects

Mevin Hooten

Publications and source records attributed to Mevin Hooten.

3 recordsLinked to original sources

Network knockoffs: controlling false discovery in dyadic space

Phenomena such as epidemiological processes, hydrologic systems, social platforms, utility services, and supply chains can be represented as topological networks. A central question about these networks concerns connectivity and the permeability of edges. Dyadic regression and related approaches have been proposed to identify network features associated with pairwise node-level differences. In high-dimensional settings, it is important to control the number of spuriously selected features. However, controlling the false discovery rate for dyadic outcomes is challenging because dependence among dyads invalidates classic asymptotic procedures and complicates standard data splitting and knockoff approaches. We propose a novel knockoff variable selection procedure that simulates synthetic features directly on the topological network prior to constructing the augmented design matrix in dyadic space. Empirically, our method controls the false discovery rate for both node- and edge-level features. The Benjamini-Hochberg, Benjamini-Yekutieli, Storey Q-value, data-splitting, and standard knockoff procedures were all anticonservative. We applied our network knockoffs to assess the impassability of over 1000 stream barriers in North Carolina for Salvelinus fontinalis. Compared to data splitting and traditional knockoff approaches, our proposed approach selected a higher proportion of barriers previously assessed to impede fish movement.

stat.ME

Melding Wildlife Surveys to Improve Conservation Inference

Integrated models are a popular tool for analyzing species of conservation concern. Species of conservation concern are often monitored by multiple entities that generate several datasets. Individually, these datasets may be insufficient for guiding management due to low spatio-temporal resolution, biased sampling, or large observational uncertainty. Integrated models provide an approach for assimilating multiple datasets in a coherent framework that can compensate for these deficiencies. While conventional integrated models have been used to assimilate count data with surveys of survival, fecundity, and harvest, they can also assimilate ecological surveys that have differing spatio-temporal regions and observational uncertainties. Motivated by independent aerial and ground surveys of lesser prairie-chicken, we developed an integrated modeling approach that assimilates density estimates derived from surveys with distinct sources of observational error into a joint framework that provides shared inference on spatio-temporal trends. We model these data using a Bayesian Markov melding approach and apply several data augmentation strategies for efficient sampling. In a simulation study, we show that our integrated model improved predictive performance relative to models that analyzed the surveys independently. We use the integrated model to facilitate prediction of lesser prairie-chicken density at unsampled regions and perform a sensitivity analysis to quantify the inferential cost associated with reduced survey effort.

stat.ME

Model Selection using Multi-Objective Optimization

Choices in scientific research and management require balancing multiple, often competing objectives.Multiple-objective optimization (MOO) provides a unifying framework for solving multiple objective problems. Model selection is a critical component to scientific inference and prediction and concerns balancing the competing objectives of model fit and model complexity. The tradeoff between model fit and model complexity provides a basis for describing the model-selection problem within the MOO framework. We discuss MOO and two strategies for solving the MOO problem; modeling preferences pre-optimization and post-optimization. Most model selection methods are consistent with solving MOO problems via specification of preferences pre-optimization. We reconcile these methods within the MOO framework. We also consider model selection using post-optimization specification of preferences. That is, by first identifying Pareto optimal solutions, and then selecting among them. We demonstrate concepts with an ecological application of model selection using avian species richness data in the continental United States.

stat.AP