Searcharxiv⌕ Search

arXiv subjects

Jabed H Tomal

Publications and source records attributed to Jabed H Tomal.

2 recordsLinked to original sources

Ensembles of phalanxes across assessment metrics for robust ranking of homologous proteins

Two proteins are homologous if they have a common evolutionary origin, and the binary classification problem is to identify proteins in a candidate set that are homologous to a particular native protein. The feature (explanatory) variables available for classification are various measures of similarity of proteins. There are multiple classification problems of this type for different native proteins and their respective candidate sets. Homologous proteins are rare in a single candidate set, giving a highly unbalanced two-class problem. The goal is to rank proteins in a candidate set according to the probability of being homologous to the set's native protein. An ideal classifier will place all the homologous proteins at the head of such a list. Our approach uses an ensemble of models in a classifier and an ensemble of assessment metrics. For a given metric a classifier combines models, each based on a subset of the available feature variables which we call phalanxes. The proposed ensemble of phalanxes identifies strong and diverse subsets of feature variables. A second phase of ensembling aggregates classifiers based on diverse evaluation metrics. The overall result is called an ensemble of phalanxes and metrics. It provide robustness against both close and distant homologues.

stat.ML↗

Statistical methods for estimating ecological breakpoints and prediction intervals

The relationships among ecological variables are usually obtained by fitting statistical models that go through the conditional means of the dependent variables. For example, the nonparametric loess and the parametric piecewise linear regression models, which pass through the conditional mean of the response variable given the predictor, are used to analyze simple to complex relationships among variables. We used loess and bootstrapped confidence interval to subjectively identify the number and positions of potential ecological breakpoints in a bivariate relationship, and a piecewise linear regression model (PLRM) to quantitatively estimate the location of breakpoints and the associated precision. We also estimated breakpoint location and precision using a piecewise linear quantile regression model (PQRM), which is fitted to the quantiles of the conditional distribution of the response variable given the predictor and provides much richer information in terms of estimating relationships and breakpoints. We compared the precision of breakpoints estimated by PQRM relative to PLRM. We compared the precision of the methods using two examples from the ecological literature suspected to exhibit multiple breakpoints: relating a Fish Index of Biotic Integrity (an index of wetlands' fish community 'health') to the amount of human activity in wetlands' adjacent watersheds; and relating the biomass of cyanobacteria to the total phosphorus concentration in Canadian lakes. Statistically significant breakpoints were detected for both datasets, demarcating the boundaries of three line segments with markedly different slopes. We recommend the piecewise linear quantile regression as an effective means of characterizing bivariate environmental relationships where the scatter of points represents natural environmental variation rather than measurement error.

stat.AP↗