SearcharxivSearch

arXiv subjects

N. A. Cruz

Publications and source records attributed to N. A. Cruz.

11 recordsLinked to original sources

Identification of Separable OTUs for Multinomial Classification in Compositional Data Analysis

High-throughput sequencing has transformed microbiome research, but it also produces inherently compositional data that challenge standard statistical and machine learning methods. In this work, we propose a multinomial classification framework for compositional microbiome data based on penalized log-ratio regression and pairwise separability screening. The method quantifies the discriminative ability of each OTU through the area under the receiver operating characteristic curve ($AUC$) for all pairwise log-ratios and aggregates these values into a global separability index $S_k$, yielding interpretable rankings of taxa together with confidence intervals. We illustrate the approach by reanalyzing the Baxter colorectal adenoma dataset and comparing our results with Greenacre's ordination-based analysis using Correspondence Analysis and Canonical Correspondence Analysis. Our models consistently recover a core subset of taxa previously identified as discriminant, thereby corroborating Greenacre's main findings, while also revealing additional OTUs that become important once demographic covariates are taken into account. In particular, adjustment for age, gender, and diabetes medication improves the precision of the separation index and highlights new, potentially relevant taxa, suggesting that part of the original signal may have been influenced by confounding. Overall, the integration of log-ratio modeling, covariate adjustment, and uncertainty estimation provides a robust and interpretable framework for OTU selection in compositional microbiome data. The proposed method complements existing ordination-based approaches by adding a probabilistic and inferential perspective, strengthening the identification of biologically meaningful microbial signatures.

stat.AP

CrossCarry: An R package for the analysis of data from a crossover design with GEE

Crossover designs are widely applied in medicine, agriculture, and other biological sciences, yet their analysis remains challenging due to longitudinal observations within each unit and the presence of carry-over effects. Despite their prevalence, there is no comprehensive R package dedicated to the statistical modeling of crossover data. The CrossCarry package addresses this gap by providing a flexible and open-source framework for analyzing any crossover design with response variables from the exponential family, with or without washout periods. It extends the generalized estimating equations (GEE) methodology by incorporating correlation structures specifically tailored to crossover data, capturing both within- and between-period dependencies. Moreover, CrossCarry integrates a parametric component for treatment effects and a nonparametric spline-based component for time and carry-over effects. This combination allows users to model complex correlation patterns and temporal structures with minimal coding effort. By offering a domain-independent implementation of advanced statistical methodology, CrossCarry facilitates reproducible research and promotes the reuse of robust analytical tools across disciplines. Its potential applications span medical trials, agricultural field experiments, and other areas where crossover designs are essential, thus contributing to broader scientific discovery and cross-domain methodological standardization.

stat.CO

Penalized GEE for Complex Carry-Over in Repeated-Measures Crossover Designs

It has been argued for many years that models used to analyze data from crossover designs are not appropriate when simple carryover effects are assumed. Furthermore, a statistical model that could estimate complex carry-over effects in crossover designs had never been found. However, in this paper, the estimability conditions of the complex carryover effects and a theoretical result that supports them are found. In addition, a simulation example is developed in a non-linear dose-response test for a typical AB/BA crossover design with repeated measures. This simulation shows that a semiparametric model can detect complex carryover effects and that this estimation improves the precision of the estimators of the treatment effect. It is concluded that when there are at least five replicates in each observation period per individual, semiparametric statistical models provide a good estimator of the treatment effect and reduce bias with respect to models that assume the absence of carryover effects or simplex carryover effects. Furthermore, an application of the methodology is shown and the wealth of analysis gained by estimating complex carryover effects is evident.

stat.ME

SAR models with specific spatial coefficients and heteroskedastic innovations

This paper presents an innovative extension of spatial autoregressive (SAR) models, introducing spatial coefficients specific to each spatial region that evolve over time. The proposed estimation methodology covers both homoscedastic and heteroscedastic data, ensuring consistency and efficiency in the estimators of the parameters $\pmbρ$ and $\pmbβ$. The model is based on a robust theoretical framework, supported by the analysis of the asymptotic properties of the estimators, which reinforces its practical implementation. To facilitate its use, an algorithm has been developed in the R software, making it a standard tool for the analysis of complex spatial data. The proposed model proves to be more effective than other similar techniques, especially when modeling data with normal spatial structures and non-normal distributions, even when the residuals are not homoscedastic. Finally, the application of the model to homicide rates in the United States highlights its advantages in both statistical and social analysis, positioning it as a key tool for the analysis of spatial data in various disciplines.

stat.ME

Generalized spatial autoregressive model

This paper presents the generalized spatial autoregression (GSAR) model, a significant advance in spatial econometrics for non-normal response variables belonging to the exponential family. The GSAR model extends the logistic SAR, probit SAR, and Poisson SAR approaches by offering greater flexibility in modeling spatial dependencies while ensuring computational feasibility. Fundamentally, theoretical results are established on the convergence, efficiency, and consistency of the estimates obtained by the model. In addition, it improves the statistical properties of existing methods and extends them to new distributions. Simulation samples show the theoretical results and allow a visual comparison with existing methods. An empirical application is made to Republican voting patterns in the United States. The GSAR model outperforms standard spatial models by capturing nuanced spatial autocorrelation and accommodating regional heterogeneity, leading to more robust inferences. These findings underline the potential of the GSAR model as an analytical tool for researchers working with categorical or count data or skewed distributions with spatial dependence in diverse domains, such as political science, epidemiology, and market research. In addition, the R codes for estimating the model are provided, which allows its adaptability in these scenarios.

stat.ME

Analysis of longitudinal data with destructive sampling using linear mixed models

This paper proposes an analysis methodology for the case where there is longitudinal data with destructive sampling of observational units, which come from experimental units that are measured at all times of the analysis. A mixed linear model is proposed and compared with regression models with fixed and mixed effects, among which is a similar that is used for data called pseudo-panel, and one of multivariate analysis of variance, which are common in statistics. To compare the models, the mean square error was used, demonstrating the advantage of the proposed methodology. In addition, an application was made to real-life data that refers to the scores in the Saber 11 tests applied to students in Colombia to see the advantage of using this methodology in practical scenarios.

stat.ME

Spatial error models with heteroskedastic normal perturbations and joint modeling of mean and variance

This work presents the spatial error model with heteroskedasticity, which allows the joint modeling of the parameters associated with both the mean and the variance, within a traditional approach to spatial econometrics. The estimation algorithm is based on the log-likelihood function and incorporates the use of GAMLSS models in an iterative form. Two theoretical results show the advantages of the model to the usual models of spatial econometrics and allow obtaining the bias of weighted least squares estimators. The proposed methodology is tested through simulations, showing notable results in terms of the ability to recover all parameters and the consistency of its estimates. Finally, this model is applied to identify the factors associated with school desertion in Colombia.

stat.ME

Estimation and imputation of missing data in longitudinal models with Zero-Inflated Poisson response variable

This research deals with the estimation and imputation of missing data in longitudinal models with a Poisson response variable inflated with zeros. A methodology is proposed that is based on the use of maximum likelihood, assuming that data is missing at random and that there is a correlation between the response variables. In each of the times, the expectation maximization (EM) algorithm is used: in step E, a weighted regression is carried out, conditioned on the previous times that are taken as covariates. In step M, the estimation and imputation of the missing data are performed. The good performance of the methodology in different loss scenarios is demonstrated in a simulation study comparing the model only with complete data, and estimating missing data using the mode of the data of each individual. Furthermore, in a study related to the growth of corn, it is tested on real data to develop the algorithm in a practical scenario.

stat.ME

Joint spatial modeling of mean and non-homogeneous variance combining semiparametric SAR and GAMLSS models for hedonic prices

In the context of spatial econometrics, it is very useful to have methodologies that allow modeling the spatial dependence of the observed variables and obtaining more precise predictions of both the mean and the variability of the response variable, something very useful in territorial planning and public policies. This paper proposes a new methodology that jointly models the mean and the variance. Also, it allows to model the spatial dependence of the dependent variable as a function of covariates and to model the semiparametric effects in both models. The algorithms developed are based on generalized additive models that allow the inclusion of non-parametric terms in both the mean and the variance, maintaining the traditional theoretical framework of spatial regression. The theoretical developments of the estimation of this model are carried out, obtaining desirable statistical properties in the estimators. A simulation study is developed to verify that the proposed method has a remarkable predictive capacity in terms of the mean square error and shows a notable improvement in the estimation of the spatial autoregressive parameter, compared to other traditional methods and some recent developments. The model is also tested on data from the construction of a hedonic price model for the city of Bogota, highlighting as the main result the ability to model the variability of housing prices, and the wealth in the analysis obtained.

stat.ME

Semi-parametric generalized estimating equations for repeated measurements in cross-over designs

A model for cross-over designs with repeated measures within each period was developed. It is obtained using an extension of generalized estimating equations that includes a parametric component to model treatment effects and a non-parametric component to model time and carry-over effects; the estimation approach for the non-parametric component is based on splines. A simulation study was carried out to explore the model properties. Thus, when there is a carry-over effect or a functional temporal effect, the proposed model presents better results than the standard models. Among the theoretical properties, the solution is found to be analogous to weighted least squares. Therefore, model diagnostics can be made adapting the results from a multiple regression. The proposed methodology was implemented in the data sets of the crossover experiments that motivated the approach of this work: systolic blood pressure and insulin in rabbits.

stat.ME

A correlation structure for the analysis of Gaussian and non-Gaussian responses in crossover experimental designs with repeated measures

In this study, we propose a family of correlation structures for crossover designs with repeated measures for both, Gaussian and non-Gaussian responses using generalized estimating equations (GEE). The structure considers two matrices: one that models between-period correlation and another one that models within-period correlation. The overall correlation matrix, which is used to build the GEE, corresponds to the Kronecker between these matrices. A procedure to estimate the parameters of the correlation matrix is proposed, its statistical properties are studied and a comparison with standard models using a single correlation matrix is carried out. A simulation study showed a superior performance of the proposed structure in terms of the quasi-likelihood criterion, efficiency, and the capacity to explain complex correlation phenomena/patterns in longitudinal data from crossover designs

stat.ME