Searcharxiv⌕ Search

arXiv subjects

Haziq Jamil

Publications and source records attributed to Haziq Jamil.

8 recordsLinked to original sources

Approximate Bayesian Inference for Structural Equation Models using Integrated Nested Laplace Approximations

Markov chain Monte Carlo (MCMC) methods remain the mainstay of Bayesian estimation of structural equation models (SEM), though they often incur a high computational cost. We present a bespoke approximate Bayesian approach to SEM, drawing on ideas from the integrated nested Laplace approximation (INLA; Rue et al., 2009, J. R. Stat. Soc., B: Stat. Methodol.) framework. After analytically integrating out the latent variables, we implement a simplified Laplace approximation that profiles the posterior using an efficient gradient-based volume correction, allowing for parametric skew-normal estimation of the marginals. A variational Bayes correction further refines the centre of the joint Gaussian approximation; marginal profiling implicitly recovers the same location correction to leading order. Essential quantities, including factor scores and model-fit indices, are obtained via a Gaussian copula sampling scheme. On a benchmark SEM, inference completes in seconds with close agreement to MCMC. A simulation study stresses the approximation along sample size, identification, boundary proximity, and model dimension. Accuracy remains high across the examined sample sizes and model dimensions, with degradation concentrated in weakly identified settings and in the lower tails of variances near the zero boundary.

stat.ME↗

Deterministic Leave-One-Cluster-Out Cross-Validation for Multilevel Bayesian Structural Equation Models

We introduce a closed-form, refit-free procedure for leave-one-cluster-out (LOCO) cross-validation in multilevel Gaussian Bayesian structural equation models (SEMs), together with predictive scoring of every nested submodel. Conditional independence of clusters given the parameters expresses the cluster-deleted posterior as a functional of the full posterior. The LOCO predictive density is then a harmonic mean of the cluster likelihood, computable from a single fit in INLAvaan, the integrated nested Laplace approximation package for Bayesian SEM. We evaluate the harmonic-mean expectation in closed form through a fully exponential Laplace approximation of the reciprocal cluster likelihood under a Gaussian posterior; candidate structural restrictions follow by Gaussian conditioning of the same Laplace summary. The resulting Taylor elpd (expected log predictive density) scores are fully deterministic, requiring neither refitting nor Monte Carlo sampling, so the variance pathology of the naive harmonic-mean estimator does not arise. A direct-sum decomposition of the compound-symmetric cluster covariance makes the cost independent of cluster size, orders of magnitude below brute-force refitting. We validate against brute-force refits and Markov chain Monte Carlo in simulation, and illustrate the procedure on school-safety climate in PISA 2022 and on 16 candidate structures linking personality and well-being in the MIDUS sibling sample.

stat.ME↗

Informative Distance-Based Priors for Correlation Matrices Centred on a Target Reference

Specifying a prior over the space of correlation matrices is a persistent challenge in Bayesian analysis. The space is a curved manifold whose dimension grows quadratically with the number of variables, making substantive prior beliefs difficult to encode.\\ We propose a distance-based prior that assigns mass decaying exponentially in the Fisher arc-length distance from a user-specified reference correlation matrix, enabling shrinkage toward any target correlation structure rather than being confined to the identity matrix. Formally, this is constructed as a Penalised Complexity prior, but its interpretation shifts accordingly: unless the chosen target represents a structurally simpler state, the shrinkage penalises deviation rather than complexity in the usual sense. To accommodate conditional independence constraints, we introduce a parameterisation that constructs the correlation matrix via the Cholesky factor of the inverse correlation matrix with respect to a user-supplied graph, thereby reducing the number of free parameters from one per variable pair to one per graph edge. The prior is proper for every positive value of its rate parameter, accommodates correlations of either sign under any graph structure, and reduces to a fully unstructured prior when the graph is complete. A direct sampling algorithm is provided, enabling prior predictive checks and sensitivity analysis, implemented within the \texttt{graphpcor} package.

stat.ME↗

Implementation and Workflows for INLA-Based Approximate Bayesian Structural Equation Modelling

Bayesian structural equation modelling (BSEM) offers many advantages such as principled uncertainty quantification, small-sample regularisation, and flexible model specification. However, the Markov chain Monte Carlo (MCMC) methods on which it relies are computationally prohibitive for the iterative cycle of specification, criticism, and refinement that careful psychometric practice demands. We present INLAvaan, an R package for fast, approximate Bayesian SEM built around the Integrated Nested Laplace Approximation (INLA) framework for structural equation models developed by Jamil & Rue (2026, arXiv:2603.25690 [stat.ME]). This paper serves as a companion manuscript that describes the architectural decisions and computational strategies underlying the package. Two substantive applications -- a 256-parameter bifactor circumplex model and a multilevel mediation model with full-information missing-data handling -- demonstrate the approach on specifications where MCMC would require hours of run time and careful convergence work. In constrast, INLAvaan delivers calibrated posterior summaries in seconds.

stat.CO↗

Bias-Reduced Estimation of Structural Equation Models

Finite-sample bias is a pervasive challenge in the estimation of structural equation models (SEMs), especially when sample sizes are small or measurement reliability is low. A range of methods have been proposed to improve finite-sample bias in the SEM literature, ranging from analytic bias corrections to resampling-based techniques, with each carrying trade-offs in scope, computational burden, and statistical performance. We apply the reduced-bias M-estimation framework (RBM, Kosmidis & Lunardon, 2024, J. R. Stat. Soc. Series B Stat. Methodol.) to SEMs. The RBM framework is attractive as it requires only first- and second-order derivatives of the log-likelihood, which renders it both straightforward to implement, and computationally more efficient compared to resampling-based alternatives such as bootstrap and jackknife. It is also robust to departures from modelling assumptions. Using the same simulation setup as in Dhaene and Rosseel (2022), we illustrate that RBM estimators consistently reduce mean bias in the estimation of SEMs without inflating mean squared error. They also deliver improvements in both median bias and inference relative to maximum likelihood estimators, while maintaining robustness under non-normality. Our findings suggest that RBM offers a promising, practical, and broadly applicable tool for mitigating bias in the estimation of SEMs, particularly in small-sample research contexts.

stat.ME↗

Pairwise likelihood estimation and limited information goodness-of-fit test statistics for binary factor analysis models under complex survey sampling

This paper discusses estimation and limited information goodness-of-fit test statistics in factor models for binary data using pairwise likelihood estimation and sampling weights. The paper extends the applicability of pairwise likelihood estimation for factor models with binary data to accommodate complex sampling designs. Additionally, it introduces two key limited information test statistics: the Pearson chi-squared test and the Wald test. To enhance computational efficiency, the paper introduces modifications to both test statistics. The performance of the estimation and the proposed test statistics under simple random sampling and unequal probability sampling is evaluated using simulated data.

stat.ME↗

Additive interaction modelling using I-priors

Additive regression models with interactions are widely studied in the literature, using methods such as splines or Gaussian process regression. However, these methods can pose challenges for estimation and model selection, due to the presence of many smoothing parameters and the lack of suitable criteria. We propose to address these challenges by extending the I-prior methodology (Bergsma, 2020) to multiple covariates, which may be multidimensional. The I-prior methodology has some advantages over other methods, such as Gaussian process regression and Tikhonov regularization, both theoretically and practically. In particular, the I-prior is a proper prior, is based on minimal assumptions, yields an admissible posterior mean, and estimation of the scale (or smoothing) parameters can be done using an EM algorithm with simple E and M steps. Moreover, we introduce a parsimonious specification of models with interactions, which has two benefits: (i) it reduces the number of scale parameters and thus facilitates the estimation of models with interactions, and (ii) it enables straightforward model selection (among models with different interactions) based on the marginal likelihood.

math.ST↗

iprior: An R Package for Regression Modelling using I-priors

This is an overview of the R package iprior, which implements a unified methodology for fitting parametric and nonparametric regression models, including additive models, multilevel models, and models with one or more functional covariates. Based on the principle of maximum entropy, an I-prior is an objective Gaussian process prior for the regression function with covariance kernel equal to its Fisher information. The regression function is estimated by its posterior mean under the I-prior, and hyperparameters are estimated via maximum marginal likelihood. Estimation of I-prior models is simple and inference straightforward, while small and large sample predictive performances are comparative, and often better, to similar leading state-of-the-art models. We illustrate the use of the iprior package by analysing a simulated toy data set as well as three real-data examples, in particular, a multilevel data set, a longitudinal data set, and a dataset involving a functional covariate.

stat.ME↗