SearcharxivSearch

arXiv subjects

Victor H. Lachos

Publications and source records attributed to Victor H. Lachos.

16 recordsLinked to original sources

Scalable Bayesian Spatiotemporal Tensor Modeling for Censored and Missing Areal Data

We propose a new Bayesian approach for spatiotemporal areal data with censored and missing observations, where the response is represented as a tensor indexed by spatial units and time points. The method introduces a flexible random effect that combines the spatial dependence structures of the Simultaneous Autoregressive (SAR) and Directed Acyclic Graph Autoregressive (DAGAR) models with a temporal autoregressive component. This formulation brings both spatial models into a unified spatiotemporal framework by expressing them as Gaussian Markov random fields in innovation form, providing an interpretable representation of spatial, temporal, and joint spatiotemporal dependence while preserving the multiway structure of the data. The innovation-based structure enables scalable Bayesian implementation in \texttt{Stan} for datasets of moderate size. Simulation studies show that the proposed model outperforms common ad hoc imputation strategies for censored and missing data. We further illustrate its practical value through an application to carbon monoxide (CO) concentrations, where the proposed framework captures complex spatiotemporal dependence in the presence of censoring and missingness.

stat.ME

Semiparametric robust mixture of experts based on nonparametric maximum likelihood

The mixture of experts (MoE) model provides a flexible approach for modeling heterogeneous regression relationships by allowing covariate-dependent mixing through a gating network, but most existing MoE models rely on parametric assumptions for expert error distributions, typically Gaussian, which can lead to inefficiency and sensitivity to outliers or heavy-tailed behavior when misspecified. We propose a semiparametric MoE model in which each expert error distribution is represented as a nonparametric Gaussian scale mixture estimated via nonparametric maximum likelihood, relaxing parametric assumptions within the Gaussian scale-mixture class while preserving the interpretability and structure of the MoE framework. The resulting model adapts to complex error structures, improves robustness under contamination and heavy tails, and remains competitive under well-specified Gaussian settings, providing a practical and theoretically grounded alternative to parametric MoE formulations.

stat.ME

Multiple Heckman Selection Model

We introduce a novel matrix-variate extension of the Heckman selection model to accommodate multiple outcomes, providing a flexible and natural generalization of classical selection models for matrix-valued data. By relying on the matrix normal distribution, the proposed model captures dependencies across both rows and columns while accounting for selection bias. An Expectation/Conditional Maximization (ECM) algorithm is developed, yielding closed-form updates for all model parameters. We investigate key theoretical properties, including the connection between sample selection models and the recently developed multivariate unified skew-normal (SUN) distribution. The performance of the proposed approach is assessed through simulation studies, and its practical utility is illustrated using two real datasets. The proposed method is implemented in the R package mvHeckman.

stat.ME

Bayesian analysis of flexible Heckman selection models using Hamiltonian Monte Carlo

The Heckman selection model is widely used in econometric analysis and other social sciences to address sample selection bias in data modeling. A common assumption in Heckman selection models is that the error terms follow an independent bivariate normal distribution. However, real-world data often deviates from this assumption, exhibiting heavy-tailed behavior, which can lead to inconsistent estimates if not properly addressed. In this paper, we propose a Bayesian analysis of Heckman selection models that replace the Gaussian assumption with well-known members of the class of scale mixture of normal distributions, such as the Student's-t and contaminated normal distributions. For these complex structures, Stan's default No-U-Turn sampler is utilized to obtain posterior simulations. Through extensive simulation studies, we compare the performance of the Heckman selection models with normal, Student's-t and contaminated normal distributions. We also demonstrate the broad applicability of this methodology by applying it to medical care and labor supply data. The proposed algorithms are implemented in the R package HeckmanStan.

stat.ME

Heckman Selection Contaminated Normal Model

The Heckman selection model is one of the most well-renounced econometric models in the analysis of data with sample selection. This model is designed to rectify sample selection biases based on the assumption of bivariate normal error terms. However, real data diverge from this assumption in the presence of heavy tails and/or atypical observations. Recently, this assumption has been relaxed via a more flexible Student's t-distribution, which has appealing statistical properties. This paper introduces a novel Heckman selection model using a bivariate contaminated normal distribution for the error terms. We present an efficient ECM algorithm for parameter estimation with closed-form expressions at the E-step based on truncated multinormal distribution formulas. The identifiability of the proposed model is also discussed, and its properties have been examined. Through simulation studies, we compare our proposed model with the normal and Student's t counterparts and investigate the finite-sample properties and the variation in missing rate. Results obtained from two real data analyses showcase the usefulness and effectiveness of our model. The proposed algorithms are implemented in the R package HeckmanEM.

stat.ME

The use of the EM algorithm for regularization problems in high-dimensional linear mixed-effects models

The EM algorithm is a popular tool for maximum likelihood estimation but has not been used much for high-dimensional regularization problems in linear mixed-effects models. In this paper, we introduce the EMLMLasso algorithm, which combines the EM algorithm and the popular and efficient R package glmnet for Lasso variable selection of fixed effects in linear mixed-effects models. We compare the performance of our proposed EMLMLasso algorithm with the one implemented in the well-known R package glmmLasso through the analyses of both simulated and real-world applications. The simulations and applications demonstrated good properties, such as consistency, and the effectiveness of the proposed variable selection procedure, for both $p < n$ and $p > n$. Moreover, in all evaluated scenarios, the EMLMLasso algorithm outperformed glmmLasso. The proposed method is quite general and can be easily extended for ridge and elastic net penalties in linear mixed-effects models.

stat.ME

Spatial Censored Regression Models in R: The CensSpatial package

CensSpatial is an R package for analyzing spatial censored data through linear models. It offers a set of tools for simulating, estimating, making predictions, and performing local influence diagnostics for outlier detection. The package provides four algorithms for estimation and prediction. One of them is based on the stochastic approximation of the EM (SAEM) algorithm, which allows easy and fast estimation of the parameters of linear spatial models when censoring is present. The package provides worthy measures to perform diagnostic analysis using the Hessian matrix of the completed log-likelihood function. This work is divided into two parts. The first part discusses and illustrates the utilities that the package offers for estimating and predicting spatial censored data. The second one describes the valuable tools to perform diagnostic analysis. Several examples in spatial environmental data are also provided.

stat.ME

On moments of folded and doubly truncated multivariate extended skew-normal distributions

This paper develops recurrence relations for integrals that relate the density of multivariate extended skew-normal (ESN) distribution, including the well-known skew-normal (SN) distribution introduced by Azzalini and Dalla-Valle (1996) and the popular multivariate normal distribution. These recursions offer a fast computation of arbitrary order product moments of the multivariate truncated extended skew-normal and multivariate folded extended skew-normal distributions with the product moments as a byproduct. In addition to the recurrence approach, we realized that any arbitrary moment of the truncated multivariate extended skew-normal distribution can be computed using a corresponding moment of a truncated multivariate normal distribution, pointing the way to a faster algorithm since a less number of integrals is required for its computation which result much simpler to evaluate. Since there are several methods available to calculate the first two moments of a multivariate truncated normal distribution, we propose an optimized method that offers a better performance in terms of time and accuracy, in addition to consider extreme cases in which other methods fail. The R MomTrunc package provides these new efficient methods for practitioners.

math.ST

Finite mixture modeling of censored and missing data using the multivariate skew-normal distribution

Finite mixture models have been widely used to model and analyze data from a heterogeneous populations. Moreover, data of this kind can be missing or subject to some upper and/or lower detection limits because of the restriction of experimental apparatuses. Another complication arises when measures of each population depart significantly from normality, for instance, asymmetric behavior. For such data structures, we propose a robust model for censored and/or missing data based on finite mixtures of multivariate skew-normal distributions. This approach allows us to model data with great flexibility, accommodating multimodality and skewness, simultaneously, depending on the structure of the mixture components. We develop an analytically simple, yet efficient, EM- type algorithm for conducting maximum likelihood estimation of the parameters. The algorithm has closed-form expressions at the E-step that rely on formulas for the mean and variance of the truncated multivariate skew-normal distributions. Furthermore, a general information-based method for approximating the asymptotic covariance matrix of the estimators is also presented. Results obtained from the analysis of both simulated and real datasets are reported to demonstrate the effectiveness of the proposed method. The proposed algorithm and method are implemented in the new R package CensMFM.

stat.ME

Scale mixture of skew-normal linear mixed models with within-subject serial dependence

In longitudinal studies, repeated measures are collected over time and hence they tend to be serially correlated. In this paper we consider an extension of skew-normal/independent linear mixed models introduced by Lachos et al. (2010), where the error term has a dependence structure, such as damped exponential correlation or autoregressive correlation of order p. The proposed model provides flexibility in capturing the effects of skewness and heavy tails simultaneously when continuous repeated measures are serially correlated. For this robust model, we present an efficient EM-type algorithm for computation of maximum likelihood estimation of parameters and the observed information matrix is derived analytically to account for standard errors. The methodology is illustrated through an application to schizophrenia data and several simulation studies. The proposed algorithm and methods are implemented in the new R package skewlmm.

stat.ME

A robust nonlinear mixed-effects model for COVID-19 deaths data

The analysis of complex longitudinal data such as COVID-19 deaths is challenging due to several inherent features: (i) Similarly-shaped profiles with different decay patterns; (ii) Unexplained variation among repeated measurements within each country, these repeated measurements may be viewed as clustered data since they are taken on the same country at roughly the same time; (iii) Skewness, outliers or skew-heavy-tailed noises are possibly embodied within response variables. This article formulates a robust nonlinear mixed-effects model based in the class of scale mixtures of skew-normal distributions for modeling COVID-19 deaths, which allows the analysts to model such data in the presence of the above described features simultaneously. An efficient EM-type algorithm is proposed to carry out maximum likelihood estimation of model parameters. The bootstrap method is used to determine inherent characteristics of the nonlinear individual profiles such as confidence interval of the predicted deaths and fitted curves. The target is to model COVID-19 deaths curves from some Latin American countries since this region is the new epicenter of the disease. Moreover, since a mixed-effect framework borrows information from the population-average effects, in our analysis we include some countries from Europe and North America that are in a more advanced stage of their COVID-19 deaths curve.

stat.AP

Moments of the doubly truncated selection elliptical distributions with emphasis on the unified multivariate skew-$t$ distribution

In this paper, we compute doubly truncated moments for the selection elliptical (SE) class of distributions, which includes some multivariate asymmetric versions of well-known elliptical distributions, such as, the normal, Student's t, slash, among others. We address the moments for doubly truncated members of this family, establishing neat formulation for high order moments as well as for its first two moments. We establish sufficient and necessary conditions for the existence of these truncated moments. Further, we propose optimized methods able to deal with extreme setting of the parameters, partitions with almost zero volume or no truncation which are validated with a brief numerical study. Finally, we present some results useful in interval censoring models. All results has been particularized to the unified skew-t (SUT) distribution, a complex multivariate asymmetric heavy-tailed distribution which includes the extended skew-t (EST), extended skew-normal (ESN), skew-t (ST) and skew-normal (SN) distributions as particular and limiting cases.

math.ST

Approximate inferences for nonlinear mixed effects models with scale mixtures of skew-normal distributions

Nonlinear mixed effects models have received a great deal of attention in the statistical literature in recent years because of their flexibility in handling longitudinal studies, including human immunodeficiency virus viral dynamics, pharmacokinetic analyses, and studies of growth and decay. A standard assumption in nonlinear mixed effects models for continuous responses is that the random effects and the within-subject errors are normally distributed, making the model sensitive to outliers. We present a novel class of asymmetric nonlinear mixed effects models that provides efficient parameters estimation in the analysis of longitudinal data. We assume that, marginally, the random effects follow a multivariate scale mixtures of skew--normal distribution and that the random errors follow a symmetric scale mixtures of normal distribution, providing an appealing robust alternative to the usual normal distribution. We propose an approximate method for maximum likelihood estimation based on an EM-type algorithm that produces approximate maximum likelihood estimates and significantly reduces the numerical difficulties associated with the exact maximum likelihood estimation. Techniques for prediction of future responses under this class of distributions are also briefly discussed. The methodology is illustrated through an application to Theophylline kinetics data and through some simulating studies.

stat.ME

Objective Bayesian analysis for spatial Student-t regression models

The choice of the prior distribution is a key aspect of Bayesian analysis. For the spatial regression setting a subjective prior choice for the parameters may not be trivial, from this perspective, using the objective Bayesian analysis framework a reference is introduced for the spatial Student-t regression model with unknown degrees of freedom. The spatial Student-t regression model poses two main challenges when eliciting priors: one for the spatial dependence parameter and the other one for the degrees of freedom. It is well-known that the propriety of the posterior distribution over objective priors is not always guaranteed, whereas the use of proper prior distributions may dominate and bias the posterior analysis. In this paper, we show the conditions under which our proposed reference prior yield to a proper posterior distribution. Simulation studies are used in order to evaluate the performance of the reference prior to a commonly used vague proper prior.

math.ST

An Extended Poisson Family of Life Distribution: A Unified Approach in Competitive and Complementary Risks

In this paper, we introduce a new approach to generate flexible parametric families of distributions. These models arise on competitive and complementary risks scenario, in which the lifetime associated with a particular risk is not observable, rather, we observe only the minimum/maximum lifetime value among all risks. The latent variables have a zero truncated Poisson distribution. For the proposed family of distribution, the extra shape parameter has an important physical interpretation in the competing and complementary risks scenario. The mathematical properties and inferential procedures are discussed. The proposed approach is applied in some existing distributions in which it is fully illustrated by an important data set.

stat.AP

Robust Bayesian model selection for heavy-tailed linear regression using finite mixtures

In this paper we present a novel methodology to perform Bayesian model selection in linear models with heavy-tailed distributions. We consider a finite mixture of distributions to model a latent variable where each component of the mixture corresponds to one possible model within the symmetrical class of normal independent distributions. Naturally, the Gaussian model is one of the possibilities. This allows for a simultaneous analysis based on the posterior probability of each model. Inference is performed via Markov chain Monte Carlo - a Gibbs sampler with Metropolis-Hastings steps for a class of parameters. Simulated examples highlight the advantages of this approach compared to a segregated analysis based on arbitrarily chosen model selection criteria. Examples with real data are presented and an extension to censored linear regression is introduced and discussed.

stat.ME