SearcharxivSearch

arXiv · 2609.00773

Deep Skew-t Mixture Models

Abstract

High-dimensional clustering is challenging when component distributions are both heavy-tailed and directionally asymmetric. We propose a deep skew-$t$ mixture model (DStMM), a hierarchical factor-analytic mixture based on the generalised-hyperbolic skew-$t$ normal mean--variance representation. A shared inverse-gamma mixing variable is propagated along each complete latent pathway, allowing heavy tails and directional asymmetry to be modelled jointly while preserving conditional Gaussianity. Each complete pathway therefore admits an exact GHST marginal representation. We formalise the reductions to symmetric deep $t$, Gaussian deep-mixture, and single-layer GHST factor-analytic models, discuss local non-identifiability and the implementation-level parameter-counting convention, and derive the conditional generalised inverse Gaussian law used for estimation. Estimation is carried out by a stochastic/Monte Carlo EM algorithm, with an explicit implementation-based parameter count for BIC architecture comparison. Simulation studies show that DStMM performs similarly to the symmetric robust model when skewness is absent but provides increasing gains as directional asymmetry becomes stronger, particularly under heavier tails; the same qualitative behaviour persists under smaller samples and unequal mixture proportions. Two real-data applications provide complementary evidence. On the UCI handwritten-digit benchmark, DStMM gives the strongest clustering performance under a common deep architecture, while on the Gas Sensor Array Drift data, DStMM improves on both deep Gaussian and deep $t$ alternatives and, under the implemented BIC criterion, selects a non-trivial second mixture layer. Together, these results support the value of propagating skewness and heavy-tail variation through a deep latent mixture while retaining an exact pathway-level likelihood.

Explore related subjects

Keep this discovery

BibTeXRIS

Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan. 2026-09-01. Deep Skew-t Mixture Models. https://arxiv.org/abs/2609.00773

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Surprise Reduction and Nullification in Bayesian and Inverse Bayesian Inference under Ambiguous Prediction-Error Attribution

In non-stationary environments, prediction errors may signal environmental change or transient outliers, and adaptive systems must track such changes without overreacting to outliers. We distinguish surprise reduction, which updates beliefs to fit observations, from surprise nullification, which weakens constraints imposed by the predictive structure, and formalize both within Bayesian and inverse Bayesian (BIB) inference. Belief and likelihood updates are derived from variational objectives sharing a nullification strength, determined endogenously by minimizing surprise under the candidate post-update predictive distribution. In the Gaussian case, nullification expands belief and likelihood variances by a common factor relative to standard Bayesian updating, leaving the ratio unchanged. BIB thus defers attribution of the prediction error, committing to neither latent-state change nor observation-process uncertainty. The nullification strength is carried over as a candidate and is maintained or released according to the predictive surprise of the next observation. In a mean estimation task with outliers and changepoints, no scanned parameter setting of a Sage-Husa-type adaptive Kalman filter, fixed-strength BIB variant, or belief-forgetting-only variant outperforms BIB in both changepoint tracking and post-outlier stability. An oracle-informed reduced Bayesian model tracks changepoints better but is less stable after outliers. Although BIB maintains no explicit hypotheses about changepoints or outliers, it generates event-dependent dynamics. The learning rate increases after changepoints, whereas after outliers, nullification is released, and this increase is suppressed. Deferring attribution and letting subsequent observations differentiate the responses may constitute a principle of adaptive inference in non-stationary environments.

stat.ME

Generalized Ridge Refitting for the Lasso and Prediction Improvement Bounds

We study a class of Lasso based estimators obtained by applying a quadratic correction on the Lasso equicorrelation set. The penalty matrix determines both the magnitude and geometry of the correction and contains, among other cases, the isotropic Lasso--Ridge correction, least squares refitting, Gram proportional interpolation between the Lasso and least squares, and coordinate specific penalties. We first derive a closed form representation and isolate the positive gain component of the resulting prediction improvement. We then control the remaining stochastic linear term in expectation by localizing the random signed equicorrelation model around a deterministic reference support. This yields a finite sample expectation bound that explicitly accounts for the randomness induced by Lasso model selection. The resulting decomposition provides a unified framework for understanding when Lasso based quadratic corrections can improve prediction.

stat.ME

Discretization in covariate-adaptive randomization: gains and losses

Covariate-adaptive randomization(CAR) is widely implemented in clinical trials to balance prognostic covariates across treatment arms. Continuous covariates are often discretized into strata in practice, yet their consequences are not clearly understood. This paper provides a comprehensive study of the impact of discretization on both the CAR design process and the inferential results thereafter. We establish the asymptotic properties of both imbalance measures and treatment effect estimators under discretized and non-discretized settings. Practical recommendations are given on when and how discretization should be employed. We show that discretization in design is generally recommended, as it enhances robustness against model misspecification. However, if the true model is known, the most efficient strategy is to balance covariates according to that model in the design. The theoretical results are corroborated by extensive simulation studies and an empirical application to a diabetes trial dataset. Together, the results clarify the gains and losses of discretization in CAR and pave the way for learning impact of discretization to other designs and beyond.

stat.ME