Searcharxiv⌕ Search

arXiv subjects

Geoffrey J. McLachlan

Publications and source records attributed to Geoffrey J. McLachlan.

At least 19 recordsLinked to original sources

Large Classification-Risk-Optional Label Acquisition

We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification problems in which features are observed for all sampling units while class labels can be acquired selectively. By combining the Fisher information supplied by an acquired label with the local geometry of multiclass excess risk, we derive an acquisition criterion that minimizes the leading asymptotic coefficient of expected multiclass excess risk. The resulting rule values a label according to how strongly its information is aligned with parameter directions that perturb the active Bayes decision boundary, rather than according to posterior uncertainty or global parameter information alone. We characterize the oracle acquisition design, establish its threshold structure, and derive face-specific and cost-sensitive extensions. An analytic example shows that posterior uncertainty and classification value can produce different, and even reversed, acquisition rankings. We further develop a two-stage adaptive procedure that attains the oracle leading-risk criterion under regularity conditions and provide explicit results for Gaussian discriminant analysis. Three-class QDA experiments illustrate the resulting acquisition geometry, while an application to the six-class Statlog Landsat Satellite data shows that classification-risk acquisition can differ materially from both uncertainty-based acquisition and the complete-classification-information comparator. The adaptive classification-risk design attains lower mean error than this Fisher comparator across the labeling budgets considered, although it does not uniformly outperform entropy or margin sampling and differences among the targeted strategies become small as the labeling budget increases.

stat.ML↗

Deep Skew-t Mixture Models

High-dimensional clustering is challenging when component distributions are both heavy-tailed and directionally asymmetric. We propose a deep skew-$t$ mixture model (DStMM), a hierarchical factor-analytic mixture based on the generalised-hyperbolic skew-$t$ normal mean--variance representation. A shared inverse-gamma mixing variable is propagated along each complete latent pathway, allowing heavy tails and directional asymmetry to be modelled jointly while preserving conditional Gaussianity. Each complete pathway therefore admits an exact GHST marginal representation. We formalise the reductions to symmetric deep $t$, Gaussian deep-mixture, and single-layer GHST factor-analytic models, discuss local non-identifiability and the implementation-level parameter-counting convention, and derive the conditional generalised inverse Gaussian law used for estimation. Estimation is carried out by a stochastic/Monte Carlo EM algorithm, with an explicit implementation-based parameter count for BIC architecture comparison. Simulation studies show that DStMM performs similarly to the symmetric robust model when skewness is absent but provides increasing gains as directional asymmetry becomes stronger, particularly under heavier tails; the same qualitative behaviour persists under smaller samples and unequal mixture proportions. Two real-data applications provide complementary evidence. On the UCI handwritten-digit benchmark, DStMM gives the strongest clustering performance under a common deep architecture, while on the Gas Sensor Array Drift data, DStMM improves on both deep Gaussian and deep $t$ alternatives and, under the implemented BIC criterion, selects a non-trivial second mixture layer. Together, these results support the value of propagating skewness and heavy-tail variation through a deep latent mixture while retaining an exact pathway-level likelihood.

stat.ME↗

Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models

We consider semi-supervised classification from a partially classified sample arising from a two-component Weibull mixture. The feature is observed for all data, whereas some class labels are missing. The probability of a missing label is modelled as a function of classification uncertainty, giving a feature-dependent missing-at-random (MAR) mechanism that shares parameters with the Weibull-mixture classifier. The missing-label indicators can therefore provide information about the classifier in addition to the observed features and available class labels. Under a common Weibull shape, a Bayes' rule has at most one positive decision boundary, which is unique when the rule is nonconstant; under unequal shapes, it can have two. We characterise these decision regions, derive the Fisher information for the classifier after adjustment for nuisance parameters in the missingness model, and obtain a decision-boundary expansion of the expected error rate of the plug-in sample rule relative to the Bayes error. The expansion yields classification-specific asymptotic relative efficiency formulas for the one- and two-boundary cases and shows that a positive-definite increase in Fisher information is sufficient, but not necessary, for a smaller first-order expected error rate. Numerical studies and a semi-synthetic analysis based on hard-drive failure data illustrate potential reductions in expected error rate and improvements in decision-boundary estimation from modelling feature-dependent label missingness.

stat.ML↗

Informative Label Missingness in Multiclass Classification Information Geometry and Excess Risk

Informative label missingness can change the usual efficiency ordering between completely and partially labelled classifiers because the pattern of missing labels may itself carry information about the classification model. We develop a general likelihood-based theory for this phenomenon in parametric multiclass classification. An efficient-information decomposition separates information lost through unavailable class memberships from information contributed by the missing-label mechanism. We then derive a quadratic expansion of plug-in excess risk over the active pairwise faces of the multiclass Bayes boundary, showing that classification efficiency depends on how information gains and losses align with directions that perturb the decision boundary. This yields a classification-weighted generalized-eigenvalue criterion under which informative partial classification may have smaller asymptotic classification risk without globally dominating complete classification in Fisher information. Near missing completely at random, with the marginal missing-label proportion fixed, redistribution of missing labels changes lost class-label information at first order, whereas efficient information from the missingness pattern appears only at second order. Three-class quadratic discriminant calculations, finite-sample experiments, and a semi-synthetic multiclass application illustrate the resulting regime-dependent behaviour.

stat.ML↗

Learning from Uncertainty-dependent Missing Labels for Semi-supervised Classification

Missing labels are usually regarded as a source of information loss in classification. We study a semi-supervised setting in which the probability of label missingness depends on the observed features through posterior classification uncertainty. In this setting, the missingness indicator is not only a record of an unobserved label, but also an observable signal generated by a mechanism linked to the classifier. We develop a likelihood-based information theory for such uncertainty-dependent missing labels. Under correct specification, we derive a Fisher-information decomposition that separates a partial-labeling component from a nonnegative mechanism-curvature term. Under joint misspecification of the label model and the missingness mechanism, we obtain the corresponding Godambe--Eicker--Huber--White sensitivity and sandwich-covariance partitions. We also clarify the relevant complete-data benchmark: favorable missingness can increase information relative to ordinary fully labeled or budget-matched non-informative labeling baselines, but cannot exceed the information in the augmented experiment in which labels and mechanism indicators are both observed. For plug-in classifiers, we connect the information decomposition to margin-based excess-risk bounds. In regular two-component mixture settings this yields the parametric \(n^{-1}\) excess-risk rate, with constants determined by the nuisance-adjusted information in discriminant directions. Gaussian-mixture calculations and a medical diagnosis example illustrate how uncertainty-dependent labeling mechanisms can improve estimation and classification under a fixed labeling budget.

math.ST↗

Favourable Missingness in Semi-Supervised Classification for Exponential Mixture Models

Semi-supervised classifiers are commonly trained from samples in which all features are observed but some class labels are missing. When label missingness is independent of the observed data, unavailable class memberships reduce Fisher information relative to a completely classified sample. We study a different regime in which the probability of label missingness depends on posterior classification uncertainty, so that the observed missing-label indicators can themselves carry information about the Bayes decision boundary. Building on the conditionally weighted information decomposition of Ahfock and McLachlan, we develop this phenomenon for a two-component exponential mixture. Although the exponential model is non-Gaussian, asymmetric, and supported on the positive half-line, its log-posterior odds remain linear in the feature. We derive Bayes' rule and its exact error rate, formulate entropy-logistic and squared-discriminant missingness mechanisms, and obtain the full partially classified likelihood. We then derive a decomposition of the Fisher information into the complete-data information, the conditionally weighted loss due to missing labels, and the information contributed by the missing labels. Numerical quadrature identifies regions in which the full likelihood classifier has asymptotic relative efficiency above or below one. Monte Carlo experiments with finite training samples broadly support the population calculations, with the largest departures from the asymptotic predictions occurring near the transition at which the relative efficiency crosses one.

stat.CO↗

SSLfmm: An R Package for Semi-Supervised Learning with Mixed Missingness

Partially labelled samples arise when features are observed for all data, but class labels are available for only a subset. In such settings, the mechanism governing label availability may itself contain information relevant to classification, yet it is typically left unmodelled in standard semi-supervised learning procedures. The SSLfmm package implements likelihood-based Gaussian finite-mixture classification in which the label-missingness process is modelled jointly with the class distribution. It supports complete-case, missing completely at random (MCAR), entropy-based missing at random (MAR), and mixed MCAR/MAR analyses. For the mixed mechanism, the source of a missing label may be observed or latent, allowing the same modelling framework to accommodate different forms of information about label availability. A common R interface is provided for model fitting, prediction, performance assessment, simulation, and entropy-based diagnostics. We describe the statistical formulation and software implementation, position SSLfmm relative to existing finite-mixture and semi-supervised learning software, and demonstrate its use through a reproducible simulation comparing observed- and latent-source analyses. A semi-synthetic application to the Blood Transfusion data further illustrates how alternative assumptions about label missingness can be fitted, compared, and diagnosed in practice.

stat.CO↗

Module-structured mixture factor models for molecular subtype discovery in transcriptomic data

High-throughput gene expression data exhibit high dimensionality, complex intergene dependence, and pronounced biological heterogeneity across samples, presenting major challenges for unsupervised clustering and disease subtype discovery. We introduce a module-structured mixture factor model that combines finite mixture modeling with low-rank latent factor representations defined at the gene-module level. By explicitly modeling gene modules in both the mean and covariance structure, the proposed framework decomposes expression variability into global gene-specific effects, cluster-specific module-level shifts, latent dependence within modules, and gene-specific residual noise. An Expectation--Conditional Maximization algorithm is applied for parameter estimation, allowing stable and scalable inference in high-dimensional transcriptomic settings. This framework enables interpretable unsupervised identification of disease-associated molecular subtypes and phenotypic heterogeneity across two autoimmune diseases using a large clinical transcriptomic dataset.

stat.AP↗

Robust Deep Mixture Models

We propose a robust deep mixture model based on a pathway-wise shared scale-mixture construction. Layer-specific component indicators are independently distributed according to their corresponding mixing proportions and jointly define a complete pathway through the latent hierarchy. Conditional on the selected pathway, a single gamma-distributed latent precision variable is shared across the deepest latent distribution, every intermediate latent transition, and the observation model. Integrating out this shared precision yields an exact multivariate Student-$t$ distribution for each complete pathway, allowing robustness to propagate coherently throughout the entire latent hierarchy rather than being introduced separately within individual latent layers. Model parameters are estimated using a stochastic expectation--maximisation algorithm. Complete-pathway responsibilities are evaluated analytically, whereas the shared latent precision variables and latent Gaussian variables are generated from their conditional distributions before updating the model parameters. The pathway-specific degrees-of-freedom parameters are estimated by one-dimensional numerical optimisation. Simulation studies demonstrate accurate recovery of the pathway-specific degrees-of-freedom parameters together with consistently improved clustering performance relative to the deep Gaussian mixture model under heavy-tailed and contaminated settings. Real-data applications further illustrate the ability of the proposed model to identify heterogeneous latent structures while reducing the influence of atypical observations. The proposed framework retains the hierarchical representation and parsimonious parameter-sharing structure of the deep Gaussian mixture model while providing coherent pathway-wise robustness.

stat.ME↗

Differential gene expression analysis via two-component mixture models with a semiparametric skew-normal scale mixture alternative

Two-component mixture models are particularly useful for identifying differentially expressed genes, but their performance can deteriorate markedly when the alternative distribution departs from parametric assumptions or symmetry. We propose a semiparametric mixture model in which the null component is standard normal and the alternative follows a skew-normal scale mixture with an unspecified scale mixing distribution. This formulation accommodates skewness and heavy tails, providing a flexible and computationally tractable tool for differential gene-expression analysis without restrictive distributional assumptions. We establish identifiability and consistency of the model and develop an efficient estimation algorithm that incorporates nonparametric maximum likelihood estimation of the scale distribution. Numerical studies show notable improvements over existing parametric and nonparametric approaches for modeling the alternative distribution, and applications to colon cancer and leukemia datasets demonstrate reduced false discovery and false negative rates.

stat.ME↗

Variable selection in frailty mixture cure models via penalized likelihood estimation

Variable selection naturally arises as a useful subject when faced with data with massive predictor space. In addition to the massive dimensionality, the data may be characterized by intra-subject correlation, and cure fraction, which are ubiquitous in longitudinal studies with recurrent events defining the endpoint of interest. However, variable selection methods simultaneously adjusting for intra-subject correlation, and cure fraction are rare. We propose a comprehensive variable selection method for frailty mixture cure models based on penalized least squares approximation via the generalized linear mixed model methodology. The method provides shrinkage estimation and selection of fixed effects in the incidence and the latency submodels, adjusting for intra-subject correlation using a random effect term. The random effect is shared between the incidence and the latency, incorporating a flexible choice of covariance structure, allowing intra-subject correlation to be modeled as either time-invariant or time-varying. Estimation is facilitated by a penalized semiparametric restricted maximum likelihood method using an expectation-maximization algorithm. Two penalty functions, namely the adaptive least absolute shrinkage and selection operator (adaptive lasso), and the smoothly clipped absolute deviation (SCAD) are studied in the proposed method. Simulation studies are considered, benchmarking the method against an oracle procedure to access its finite sample performance. The practical utility of the method is illustrated using data on recurrent events from a breast cancer gene expression study. In the presence of a relatively large predictor space, results show that the method yields plausible interpretability in whole, as opposed to an unpenalized model.

stat.ME↗

Informative missingness and its implications in semi-supervised learning

Semi-supervised learning (SSL) constructs classifiers using both labelled and unlabelled data. It leverages information from labelled samples, whose acquisition is often costly or labour-intensive, together with unlabelled data to enhance prediction performance. This defines an incomplete-data problem, which statistically can be formulated within the likelihood framework for finite mixture models that can be fitted using the expectation-maximisation (EM) algorithm. Ideally, one would prefer a completely labelled sample, as one would anticipate that a labelled observation provides more information than an unlabelled one. However, when the mechanism governing label absence depends on the observed features or the class labels or both, the missingness indicators themselves contain useful information. In certain situations, the information gained from modelling the missing-label mechanism can even outweigh the loss due to missing labels, yielding a classifier with a smaller expected error than one based on a completely labelled sample analysed. This improvement arises particularly when class overlap is moderate, labelled data are sparse, and the missingness is informative. Modelling such informative missingness thus offers a coherent statistical framework that unifies likelihood-based inference with the behaviour of empirical SSL methods.

stat.ML↗

Semi-supervised Gaussian mixture modelling with a missing-data mechanism in R

Semi-supervised learning is being extensively applied to estimate classifiers from training data in which not all the labels of the feature vectors are available. We present gmmsslm, an R package for estimating the Bayes' classifier from such partially classified data in the case where the feature vector has a multivariate Gaussian (normal) distribution in each of the predefined classes. Our package implements a recently proposed Gaussian mixture modelling framework that incorporates a missingness mechanism for the missing labels in which the probability of a missing label is represented via a logistic model with covariates that depend on the entropy of the feature vector. Under this framework, it has been shown that the accuracy of the Bayes' classifier formed from the Gaussian mixture model fitted to the partially classified training data can even have lower error rate than if it were estimated from the sample completely classified. This result was established in the particular case of two Gaussian classes with a common covariance matrix. Here, we focus on the effective implementation of an algorithm for multiple Gaussian classes with arbitrary covariance matrices. A strategy for initialising the algorithm is discussed and illustrated. The new package is demonstrated on some real data.

stat.CO↗

Functional Mixtures-of-Experts

We consider the statistical analysis of heterogeneous data for prediction in situations where the observations include functions, typically time series. We extend the modeling with Mixtures-of-Experts (ME), as a framework of choice in modeling heterogeneity in data for prediction with vectorial observations, to this functional data analysis context. We first present a new family of ME models, named functional ME (FME) in which the predictors are potentially noisy observations, from entire functions. Furthermore, the data generating process of the predictor and the real response, is governed by a hidden discrete variable representing an unknown partition. Second, by imposing sparsity on derivatives of the underlying functional parameters via Lasso-like regularizations, we provide sparse and interpretable functional representations of the FME models called iFME. We develop dedicated expectation--maximization algorithms for Lasso-like (EM-Lasso) regularized maximum-likelihood parameter estimation strategies to fit the models. The proposed models and algorithms are studied in simulated scenarios and in applications to two real data sets, and the obtained results demonstrate their performance in accurately capturing complex nonlinear relationships and in clustering the heterogeneous regression data.

stat.ME↗

Semi-Supervised Learning of Classifiers from a Statistical Perspective: A Brief Review

There has been increasing attention to semi-supervised learning (SSL) approaches in machine learning to forming a classifier in situations where the training data for a classifier consists of a limited number of classified observations but a much larger number of unclassified observations. This is because the procurement of classified data can be quite costly due to high acquisition costs and subsequent financial, time, and ethical issues that can arise in attempts to provide the true class labels for the unclassified data that have been acquired. We provide here a review of statistical SSL approaches to this problem, focussing on the recent result that a classifier formed from a partially classified sample can actually have smaller expected error rate than that if the sample were completely classified.

stat.ML↗

Harmless label noise and informative soft-labels in supervised classification

Manual labelling of training examples is common practice in supervised learning. When the labelling task is of non-trivial difficulty, the supplied labels may not be equal to the ground-truth labels, and label noise is introduced into the training dataset. If the manual annotation is carried out by multiple experts, the same training example can be given different class assignments by different experts, which is indicative of label noise. In the framework of model-based classification, a simple, but key observation is that when the manual labels are sampled using the posterior probabilities of class membership, the noisy labels are as valuable as the ground-truth labels in terms of statistical information. A relaxation of this process is a random effects model for imperfect labelling by a group that uses approximate posterior probabilities of class membership. The relative efficiency of logistic regression using the noisy labels compared to logistic regression using the ground-truth labels can then be derived. The main finding is that logistic regression can be robust to label noise when label noise and classification difficulty are positively correlated. In particular, when classification difficulty is the only source of label errors, multiple sets of noisy labels can supply more information for the estimation of a classification rule compared to the single set of ground-truth labels.

stat.ML↗

Data-fusion using factor analysis and low-rank matrix completion

Data-fusion involves the integration of multiple related datasets. The statistical file-matching problem is a canonical data-fusion problem in multivariate analysis, where the objective is to characterise the joint distribution of a set of variables when only strict subsets of marginal distributions have been observed. Estimation of the covariance matrix of the full set of variables is challenging given the missing-data pattern. Factor analysis models use lower-dimensional latent variables in the data-generating process, and this introduces low-rank components in the complete-data matrix and the population covariance matrix. The low-rank structure of the factor analysis model can be exploited to estimate the full covariance matrix from incomplete data via low-rank matrix completion. We prove the identifiability of the factor analysis model in the statistical file-matching problem under conditions on the number of factors and the number of shared variables over the observed marginal subsets. Additionally, we provide an EM algorithm for parameter estimation. On several real datasets, the factor model gives smaller reconstruction errors in file-matching problems than the common approaches for low-rank matrix completion.

stat.ME↗

Multi-Node EM Algorithm for Finite Mixture Models

Finite mixture models are powerful tools for modelling and analyzing heterogeneous data. Parameter estimation is typically carried out using maximum likelihood estimation via the Expectation-Maximization (EM) algorithm. Recently, the adoption of flexible distributions as component densities has become increasingly popular. Often, the EM algorithm for these models involves complicated expressions that are time-consuming to evaluate numerically. In this paper, we describe a parallel implementation of the EM-algorithm suitable for both single-threaded and multi-threaded processors and for both single machine and multiple-node systems. Numerical experiments are performed to demonstrate the potential performance gain n different settings. Comparison is also made across two commonly used platforms - R and MATLAB. For illustration, a fairly general mixture model is used in the comparison.

stat.CO↗