SearcharxivSearch

arXiv subjects

Siliang Zhang

Publications and source records attributed to Siliang Zhang.

12 recordsLinked to original sources

A Cumulative Ordered Spike-and-Slab Prior for Adaptive Dimension Selection in Joint Latent Space Models

Network models are increasingly vital in psychometrics for analyzing relational data, which are often accompanied by high-dimensional node attributes. Joint latent space models (JLSM) provide an elegant framework for integrating these data sources by assuming a shared underlying latent representation; however, a persistent methodological challenge is determining the dimension of the latent space, as existing methods typically require pre-specification or rely on computationally intensive post-hoc procedures. The key innovation of this work is a cumulative ordered spike-and-slab (COSS) prior, which we incorporate within a Bayesian joint latent space modeling framework. This prior enables the latent dimension to be inferred automatically and simultaneously with all model parameters. We develop an efficient Markov Chain Monte Carlo (MCMC) algorithm for posterior computation. Theoretically, we establish that the posterior distribution concentrates on the true latent dimension and that parameter estimates achieve Hellinger consistency at a near-optimal rate that adapts to the unknown dimensionality. Through extensive simulations and three real-data applications, we demonstrate the method's superior performance in both dimension recovery and parameter estimation. Our work offers a principled, computationally efficient, and theoretically grounded solution for adaptive dimension selection in psychometric network models.

stat.ME

A Latent Variable Framework for Multiple Imputation with Non-ignorable Missingness: Analyzing Perceptions of Social Justice in Europe

This paper proposes a general multiple imputation approach for analyzing large-scale data with missing values. An imputation model is derived from a joint distribution induced by a latent variable model, which can flexibly capture associations among variables of mixed types. The model also allows for missingness which depends on the latent variables and is thus non-ignorable with respect to the observed data. We develop a frequentist multiple imputation method for this framework and provide asymptotic theory that establishes valid inference for a broad class of analysis models. Simulation studies confirm the method's theoretical properties and robust practical performance. The procedure is applied to a cross-national analysis of individuals' perceptions of justice and fairness of income distributions in their societies, using data from the European Social Survey which has substantial nonresponse. The analysis demonstrates that failing to account for non-ignorable missingness can yield biased conclusions; for instance, complete-case analysis is shown to exaggerate the correlation between personal income and perceived fairness of income distributions in society. Code implementing the proposed methodology is publicly available at https://anonymous.4open.science/r/non-ignorable-missing-data-imputation-E885.

stat.ME

A Note on Ising Network Analysis with Missing Data

The Ising model has become a popular psychometric model for analyzing item response data. The statistical inference of the Ising model is typically carried out via a pseudo-likelihood, as the standard likelihood approach suffers from a high computational cost when there are many variables (i.e., items). Unfortunately, the presence of missing values can hinder the use of pseudo-likelihood, and a listwise deletion approach for missing data treatment may introduce a substantial bias into the estimation and sometimes yield misleading interpretations. This paper proposes a conditional Bayesian framework for Ising network analysis with missing data, which integrates a pseudo-likelihood approach with iterative data imputation. An asymptotic theory is established for the method. Furthermore, a computationally efficient {P{ó}lya}-Gamma data augmentation procedure is proposed to streamline the sampling of model parameters. The method's performance is shown through simulations and a real-world application to data on major depressive and generalized anxiety disorders from the National Epidemiological Survey on Alcohol and Related Conditions (NESARC).

stat.ME

Adjusting for non-confounding covariates in case-control association studies

There is a considerable literature in case-control logistic regression on whether or not non-confounding covariates should be adjusted for. However, only limited and ad hoc theoretical results are available on this important topic. A constrained maximum likelihood method was recently proposed, which appears to be generally more powerful than logistic regression methods with or without adjusting for non-confounding covariates. This note provides a theoretical clarification for the case-control logistic regression with and without covariate adjustment and the constrained maximum likelihood method on their relative performances in terms of asymptotic relative efficiencies. We show that the benefit of covariate adjustment in the case-control logistic regression depends on the disease prevalence. We also show that the constrained maximum likelihood estimator gives an asymptotically uniformly most powerful test.

math.ST

Longitudinal analysis of exchanges of support between parents and children in the UK

We consider how exchanges of support between parents and adult children vary by demographic and socio-economic characteristics and examine evidence for reciprocity in transfers and substitution between practical and financial support. Using data from the UK Household Longitudinal Study 2011-19, repeated measures of help given and received are analysed jointly using multivariate random effects probit models. Exchanges are considered from both a child and parent perspective. In the latter case, we propose a novel approach to account for correlation between mother and father reports and develop an efficient MCMC algorithm suitable for large datasets with multiple outcomes.

stat.ME

Modelling Correlation Matrices in Multivariate Dyadic Data: Latent Variable Models for Intergenerational Exchanges of Family Support

We define a model for the joint distribution of multiple continuous latent variables which includes a model for how their correlations depend on explanatory variables. This is motivated by and applied to social scientific research questions in the analysis of intergenerational help and support within families, where the correlations describe reciprocity of help between generations and complementarity of different kinds of help. We propose an MCMC procedure for estimating the model which maintains the positive definiteness of the implied correlation matrices, and describe theoretical results which justify this approach and facilitate efficient implementation of it. The model is applied to data from the UK Household Longitudinal Study to analyse exchanges of practical and financial support between adult individuals and their non-coresident parents.

stat.ME

Computation for Latent Variable Model Estimation: A Unified Stochastic Proximal Framework

Latent variable models have been playing a central role in psychometrics and related fields. In many modern applications, the inference based on latent variable models involves one or several of the following features: (1) the presence of many latent variables, (2) the observed and latent variables being continuous, discrete, or a combination of both, (3) constraints on parameters, and (4) penalties on parameters to impose model parsimony. The estimation often involves maximizing an objective function based on a marginal likelihood/pseudo-likelihood, possibly with constraints and/or penalties on parameters. Solving this optimization problem is highly non-trivial, due to the complexities brought by the features mentioned above. Although several efficient algorithms have been proposed, there lacks a unified computational framework that takes all these features into account. In this paper, we fill the gap. Specifically, we provide a unified formulation for the optimization problem and then propose a quasi-Newton stochastic proximal algorithm. Theoretical properties of the proposed algorithms are established. The computational efficiency and robustness are shown by simulation studies under various settings for latent variable model estimation.

stat.ME

Latent variable models for multivariate dyadic data with zero inflation: Analysis of intergenerational exchanges of family support

Understanding the help and support that is exchanged between family members of different generations is of increasing importance, with research questions in sociology and social policy focusing on both predictors of the levels of help given and received, and on reciprocity between them. We propose general latent variable models for analysing such data, when helping tendencies in each direction are measured by multiple binary indicators of specific types of help. The model combines two continuous latent variables, which represent the helping tendencies, with two binary latent class variables which allow for high proportions of responses where no help of any kind is given or received. This defines a multivariate version of a zero inflation model. The main part of the models is estimated using MCMC methods, with a bespoke data augmentation algorithm. We apply the models to analyse exchanges of help between adult individuals and their non-coresident parents, using survey data from the UK Household Longitudinal Study.

stat.ME

Estimation Methods for Item Factor Analysis: An Overview

Item factor analysis (IFA) refers to the factor models and statistical inference procedures for analyzing multivariate categorical data. IFA techniques are commonly used in social and behavioral sciences for analyzing item-level response data. Such models summarize and interpret the dependence structure among a set of categorical variables by a small number of latent factors. In this chapter, we review the IFA modeling technique and commonly used IFA models. Then we discuss estimation methods for IFA models and their computation, with a focus on the situation where the sample size, the number of items, and the number of factors are all large. Existing statistical softwares for IFA are surveyed. This chapter is concluded with suggestions for practical applications of IFA methods and discussions of future directions.

stat.ME

Joint Maximum Likelihood Estimation for High-dimensional Exploratory Item Response Analysis

Joint maximum likelihood (JML) estimation is one of the earliest approaches to fitting item response theory (IRT) models. This procedure treats both the item and person parameters as unknown but fixed model parameters and estimates them simultaneously by solving an optimization problem. However, the JML estimator is known to be asymptotically inconsistent for many IRT models, when the sample size goes to infinity and the number of items keeps fixed. Consequently, in the psychometrics literature, this estimator is less preferred to the marginal maximum likelihood (MML) estimator. In this paper, we re-investigate the JML estimator for high-dimensional exploratory item factor analysis, from both statistical and computational perspectives. In particular, we establish a notion of statistical consistency for a constrained JML estimator, under an asymptotic setting that both the numbers of items and people grow to infinity and that many responses may be missing. A parallel computing algorithm is proposed for this estimator that can scale to very large datasets. Via simulation studies, we show that when the dimensionality is high, the proposed estimator yields similar or even better results than those from the MML estimator, but can be obtained computationally much more efficiently. An illustrative real data example is provided based on the revised version of Eysenck's Personality Questionaire (EPQ-R).

stat.ME

A Latent Gaussian Process Model for Analyzing Intensive Longitudinal Data

Intensive longitudinal studies are becoming progressively more prevalent across many social science areas, especially in psychology. New technologies like smart-phones, fitness trackers, and the Internet of Things make it much easier than in the past for data collection in intensive longitudinal studies, providing an opportunity to look deep into the underlying characteristics of individuals under a high temporal resolution. In this paper, we introduce a new modeling framework for latent curve analysis that is more suitable for the analysis of intensive longitudinal data than existing latent curve models. Specifically, through the modeling of an individual-specific continuous-time latent process, some unique features of intensive longitudinal data are better captured, including intensive measurements in time and unequally spaced time points of observations. Technically, the continuous-time latent process is modeled by a Gaussian process model. This model can be regarded as a semi-parametric extension of the classical latent curve models and falls under the framework of structural equation modeling. Procedures for parameter estimation and statistical inference are provided under an empirical Bayes framework and evaluated by simulation studies. We illustrate the use of the proposed model through the analysis of an ecological momentary assessment dataset.

stat.ME

Structured Latent Factor Analysis for Large-scale Data: Identifiability, Estimability, and Their Implications

Latent factor models are widely used to measure unobserved latent traits in social and behavioral sciences, including psychology, education, and marketing. When used in a confirmatory manner, design information is incorporated, yielding structured (confirmatory) latent factor models. Motivated by the applications of latent factor models to large-scale measurements which consist of many manifest variables (e.g. test items) and a large sample size, we study the properties of structured latent factor models under an asymptotic setting where both the number of manifest variables and the sample size grow to infinity. Specifically, under such an asymptotic regime, we provide a definition of the structural identifiability of the latent factors and establish necessary and sufficient conditions on the measurement design that ensure the structural identifiability under a general family of structured latent factor models. In addition, we propose an estimator that can consistently recover the latent factors under mild conditions. This estimator can be efficiently computed through parallel computing. Our results shed lights on the design of large-scale measurement and have important implications on measurement validity. The properties of the proposed estimator are verified through simulation studies.

stat.ME