SearcharxivSearch

arXiv subjects

Safaa K. Kadhem

Publications and source records attributed to Safaa K. Kadhem.

4 recordsLinked to original sources

On WAIC for Dependent Data: A Covariance-Corrected Framework with Linear-Time Complexity

The Widely Applicable Information Criterion (WAIC) is a cornerstone of Bayesian model selection, but its conditional independence assumption renders it inappropriate for sequential and spatially correlated data - a limitation that leads to systematically underestimated model complexity and over-optimistic predictive assessments. We introduce CC-WAIC, a principled generalization of WAIC that explicitly incorporates the full posterior covariance structure of log-likelihood contributions. CC-WAIC reduces exactly to WAIC when independence holds, making it a natural extension rather than an ad-hoc modification. To overcome the prohibitive computational cost of the full covariance matrix, we develop a linear-time implementation using a banded covariance approximation that reduces complexity, with theoretical guarantees on the approximation error. We further introduce an effective sample size correction to mitigate finite-sample bias in MCMC estimation. Through extensive simulations on Hidden Markov Models and real-world applications to Old Faithful geyser data and S&P 500 volatility modelling, we demonstrate that CC-WAIC substantially outperforms WAIC, LOO-CV, iWAIC, and WAICNF, particularly under strong temporal dependence and limited sample sizes. The proposed criterion offers a computationally scalable and theoretically grounded tool for Bayesian model selection in the dependent data settings that are ubiquitous across modern science. Limitations include reliance on exponential mixing and exact conditional likelihoods; extensions to long-memory processes and approximate inference are discussed.

stat.ME

Regularized Regression by Composition: Identifiability, Structured Penalization, and Statistical Guarantees for Multi-Flow Distributional Models

Regression by composition provides a flexible framework for constructing conditional distributions through sequential group actions. However, when multiple flows act on the same distribution, the model becomes non-identifiable, leading to flat likelihood regions and unstable estimates. We introduce a structured regularization framework that resolves this issue by assigning flow-specific penalties. The resulting estimator is defined as a penalized maximum likelihood problem with heterogeneous regularization across flows. We establish theoretical properties, including identifiability under penalization, uniqueness of the minimizer via strict convexification, and asymptotic consistency. For the adaptive Lasso, we further prove the oracle property. An efficient proximal gradient algorithm handles non-smooth penalties. Extensive simulation studies evaluate performance under varying sample sizes, correlation structures, and signal-to-noise ratios, demonstrating that regularized methods (Lasso and Elastic Net) successfully break non-identifiability and achieve low estimation error with controlled false positive rates. An application to NHANES data on asthma and lead exposure illustrates the practical utility: the unregularized estimator yields implausible coefficients, whereas regularized estimators produce stable and interpretable models and automatically select the relevant risk transformation. The Labbe plots derived from regularized estimators indicate a protective effect of reducing lead exposure. The proposed framework bridges identifiability theory with penalized estimation and opens the door to high-dimensional and longitudinal extensions.

stat.ME

Balancing Efficiency and Feasibility: A Sensitivity Analysis of the Augmentation Parameter in the Finite Selection Model

This paper investigates the role of the augmentation parameter in the Finite Selection Model (FSM) and its impact on estimator performance. Through a comprehensive Monte Carlo simulation study, we analyze the sensitivity of bias, variance, and mean squared error to different values of the augmentation parameter. The results demonstrate that moderate augmentation improves covariate balance while maintaining estimation efficiency. However, excessive augmentation may increase variance and reduce estimator stability. The findings provide practical guidelines for selecting the augmentation parameter in applied experimental design settings.

stat.ME

Covariance-Corrected WAIC for Bayesian Sequential Data Models

This paper introduces and develops a theoretical extension of the widely applicable information criterion (WAIC), called the Covariance-Corrected WAIC (CC-WAIC), that applied for Bayesian sequential data models. The CC-WAIC accounts for temporal or structural dependence by incorporating the full posterior covariance structure of the log-likelihood contributions, in contrast to the classical WAIC that assumes conditional independence among data. We exploit the limitations of classical WAIC in the sequential data contexts and derive the CC-WAIC criterion under a theoretical framework. In addition, we propose a bias correction based on effective sample size to improve estimation from Markov Chain Monte Carlo (MCMC) simulations. Furthermore, we highlight the advantages of CC-WAIC in terms of stability and appropriateness for dependent data. This new criterion is supported by formal mathematical derivations, illustrative examples, and discussion of implications for model selection in both classical and modern Bayesian applications. To evaluate the reliability of CC-WAIC under varying data regimes, we conduct simulation experiments across multiple time series lengths (small, medium, and large) and different levels of temporal dependence, enabling a comprehensive performance assessment.

stat.ME