Searcharxiv⌕ Search

arXiv subjects

Jeffrey L. Andrews

Publications and source records attributed to Jeffrey L. Andrews.

6 recordsLinked to original sources

An efficient EM algorithm for both element-wise and structural missingness in matrix-variate normal mixture models

Matrix-variate data with missing entries arise frequently in applications where observations are naturally organized as two-dimensional arrays. Although the matrix normal distribution provides a parsimonious model through its Kronecker covariance structure, standard EM estimation can be computationally expensive because arbitrary missingness patterns typically destroy this separability in the E-step. In this paper, we propose an efficient partial EM algorithm for matrix-variate normal data with missing entries. The proposed method updates the conditional mean and covariance of the missing component through coordinate-wise approximations, avoiding repeated inversion of pattern-specific covariance matrices and avoiding construction of the full vectorized covariance matrix. We further develop a specialized update for submatrix missingness, where the missing-block precision retains a Kronecker product structure, and the covariance update can be carried out independently in the row and column directions. Simulation studies show that the proposed methods substantially reduce computation time compared with exact EM while preserving nearly identical observed-data likelihood across a range of dimensions and missing proportions. A real-data application to hyperspectral image patches demonstrates that the proposed imputation strategy can be embedded within a matrix-variate mixture model for simultaneous imputation and clustering.

stat.ME↗

Robust model-based clustering via mixtures of multivariate pseudo-Voigt distributions

We propose a multivariate extension of the pseudo-Voigt profile-a weighted convex combination of Gaussian and Cauchy distributions-within a finite mixture modeling framework for robust model-based clustering and outlier detection. To ensure parsimony and coherence within clusters, shared location and scale parameters are imposed between the Gaussian and Cauchy components. Parameter estimation is carried out via an Expectation Maximization algorithm, with latent variables facilitating efficient likelihood-based inference. The performance of the proposed model is evaluated through simulation studies and applications to real-world data. Comparisons with established robust models, including mixtures of contaminated normal distributions, are provided to illustrate the model's clustering accuracy and outlier detection capabilities. The framework is shown to be particularly effective for data characterized by heavy-tailed behavior.

stat.ME↗

Mixtures of spatial factor analyzers for tensor-variate data

A mixture of spatial factor analyzers (MSFA) is introduced to address the challenges of clustering high-dimensional spatial data. By leveraging the underlying coordinate system, the proposed framework incorporates a flexible, spline-based spatial decay covariance structure that prevents parameter inflation as dimensionality increases. To model non-spatial dependence, matrix variate factor analyzers are employed for further dimensionality reduction. Parameter estimation is conducted via a variant of the expectation-maximization algorithm combined with a generalized least squares estimator. The proposed models are explored in the context of tensor-variate data analysis, where simulation studies and applications to Raman spectroscopy and hyperspectral texture databases demonstrate their capacity to accurately infer and differentiate distinct spatial patterns.

stat.ME↗

Spatial Covariance Constraints for Gaussian Mixture Models

Although extensive research exists in spatial modeling, few studies have addressed finite mixture model-based clustering methods for spatial data. Finite mixture models, especially Gaussian mixture models, particularly suffer from high dimensionality due to the number of free covariance parameters. This study introduces a spatial covariance constraint for Gaussian mixture models that requires only four free parameters for each component, independent of dimensionality. Using a coordinate system, the spatially constrained Gaussian mixture model enables clustering of multi-way spatial data and inference of spatial patterns. The parameter estimation is conducted by combining the expectation-maximization (EM) algorithm with the generalized least squares (GLS) estimator. Simulation studies and applications to Raman spectroscopy data are provided to demonstrate the proposed model.

stat.ME↗

Nonnegative Matrix Factorization with Group and Basis Restrictions

Nonnegative matrix factorization (NMF) is a popular method used to reduce dimensionality in data sets whose elements are nonnegative. It does so by decomposing the data set of interest, $\mathbf{X}$, into two lower rank nonnegative matrices multiplied together ($\mathbf{X} \approx \mathbf{WH}$). These two matrices can be described as the latent factors, represented in the rows of $\mathbf{H}$, and the scores of the observations on these factors that are found in the rows of $\mathbf{W}$. This paper provides an extension of this method which allows one to specify prior knowledge of the data, including both group information and possible underlying factors. This is done by further decomposing the matrix, $\mathbf{H}$, into matrices $\mathbf{A}$ and $\mathbf{S}$ multiplied together. These matrices represent an 'auxiliary' matrix and a semi-constrained factor matrix respectively. This method and its updating criterion are proposed, followed by its application on both simulated and real world examples.

stat.ME↗

Variable Selection for Clustering and Classification

As data sets continue to grow in size and complexity, effective and efficient techniques are needed to target important features in the variable space. Many of the variable selection techniques that are commonly used alongside clustering algorithms are based upon determining the best variable subspace according to model fitting in a stepwise manner. These techniques are often computationally intensive and can require extended periods of time to run; in fact, some are prohibitively computationally expensive for high-dimensional data. In this paper, a novel variable selection technique is introduced for use in clustering and classification analyses that is both intuitive and computationally efficient. We focus largely on applications in mixture model-based learning, but the technique could be adapted for use with various other clustering/classification methods. Our approach is illustrated on both simulated and real data, highlighted by contrasting its performance with that of other comparable variable selection techniques on the real data sets.

stat.CO↗