SearcharxivSearch

arXiv subjects

Zhihua Ma

Publications and source records attributed to Zhihua Ma.

9 recordsLinked to original sources

Modeling Nonlinear Ability Trajectories and Learner Heterogeneity in Online Learning: A Bayesian Nonparametric Dynamic IRT Framework

Online learning has amplified the need to understand how student engagement patterns influence learning outcomes, particularly given the flexibility of technology-mediated environments. To address this, we propose a Bayesian nonparametric dynamic item response theory (IRT) framework that tracks within-individual ability trajectories across instructional units. The proposed model integrates B-spline basis expansions to capture nonlinear effects of engagement behaviors on ability drift, alongside a Mixture-of-Finite-Mixtures (MFM) prior to automatically determine the number of latent learner clusters. This framework overcomes three limitations in the existing literature: (1) rigid linearity assumptions in engagement-ability relationships, (2) dependence on pre-specified cluster counts, and (3) the inability to track longitudinal ability dynamics. We apply the model to longitudinal data from 198 undergraduates completing a 9-chapter introductory statistics course on CourseKata. The model automatically identified four distinct learner profiles: struggling-declining (11\%), low-stable (23\%), mainstream-stable (55\%), and high-improving (12\%). Results indicate that ability trajectories remained remarkably stable across chapters, and engagement quantity metrics did not significantly predict ability drift. These findings suggest that in introductory online statistics education, academic ability primarily reflects a stable pre-existing characteristic rather than a dynamically malleable course outcome. Ultimately, this framework offers a flexible tool for learner profiling to inform adaptive instructional design.

stat.AP

Non-segmental Bayesian Detection of Multiple Change-points

We propose an original and general NOn-SEgmental (NOSE) approach for the detection of multiple change-points. NOSE identifies change-points by the non-negligibility of posterior estimates of the jump heights. Alternatively, under the Bayesian paradigm, NOSE treats the step-wise signal as a global infinite dimensional parameter drawn from a proposed process of atomic representation, where the random jump heights determine the locations and the number of change-points simultaneously. The random jump heights are further modeled by a Gamma-Indian buffet process shrinkage prior under the form of discrete spike-and-slab. The induced maximum a posteriori estimates of the jump heights are consistent and enjoy zerodiminishing false negative rate in discrimination under a 3-sigma rule. The success of NOSE is guaranteed by the posterior inferential results such as the minimaxity of posterior contraction rate, and posterior consistency of both locations and the number of abrupt changes. NOSE is applicable and effective to detect scale shifts, mean shifts, and structural changes in regression coefficients under linear or autoregression models. Comprehensive simulations and several real-world examples demonstrate the superiority of NOSE in detecting abrupt changes under various data settings.

stat.ME

Dependent Dirichlet Processes for Analysis of a Generalized Shared Frailty Model

Bayesian paradigm takes advantage of well fitting complicated survival models and feasible computing in survival analysis owing to the superiority in tackling the complex censoring scheme, compared with the frequentist paradigm. In this chapter, we aim to display the latest tendency in Bayesian computing, in the sense of automating the posterior sampling, through Bayesian analysis of survival modeling for multivariate survival outcomes with complicated data structure. Motivated by relaxing the strong assumption of proportionality and the restriction of a common baseline population, we propose a generalized shared frailty model which includes both parametric and nonparametric frailty random effects so as to incorporate both treatment-wise and temporal variation for multiple events. We develop a survival-function version of ANOVA dependent Dirichlet process to model the dependency among the baseline survival functions. The posterior sampling is implemented by the No-U-Turn sampler in Stan, a contemporary Bayesian computing tool, automatically. The proposed model is validated by analysis of the bladder cancer recurrences data. The estimation is consistent with existing results. Our model and Bayesian inference provide evidence that the Bayesian paradigm fosters complex modeling and feasible computing in survival analysis and Stan relaxes the posterior inference.

stat.ME

Bayesian Clustered Coefficients Regression with Auxiliary Covariates Assistant Random Effects

In regional economics research, a problem of interest is to detect similarities between regions, and estimate their shared coefficients in economics models. In this article, we propose a mixture of finite mixtures (MFM) clustered regression model with auxiliary covariates that account for similarities in demographic or economic characteristics over a spatial domain. Our Bayesian construction provides both inference for number of clusters and clustering configurations, and estimation for parameters for each cluster. Empirical performance of the proposed model is illustrated through simulation experiments, and further applied to a study of influential factors for monthly housing cost in Georgia.

stat.ME

Phase shifting, dispersion variation and defocusing suppression in wave breaking

We present an investigation of the fundamental physical processes involved in deep water wave breaking. Our motivation is to identify the underlying reason causing the deficiency of the eddy viscosity breaking model (EVBM) in predicting surface elevation for strongly nonlinear waves. Owing to the limitation of experimental methods in the provision of high-resolution flow information, we propose a numerical methodology by developing an EVBM enclosed standalone fully-nonlinear quasi-potential (FNP) flow model and a coupled FNP plus Navier-Stokes flow model. The numerical models were firstly verified with a wave train subject to modulational instability, then used to simulate a series of broad-banded focusing wave trains under non-, moderate- and strong-breaking conditions. A systematic analysis was carried out to investigate the discrepancies of numerical solutions produced by the two models in surface elevation and other important physical properties. It is found that EVBM predicts accurately the energy dissipated by breaking and the amplitude spectrum of free waves in terms of magnitude, but fails to capture accurately breaking induced phase shifting. The shift of phase grows with breaking intensity and is especially strong for high wavenumber components. This is identified as a cause of the upshift of wave dispersion relation, which increases the frequencies of large wavenumber components. Such a variation drives large-wavenumber components to propagate at nearly the same speed, which is significantly higher than the linear dispersion levels. This suppresses the instant dispersive spreading of harmonics after the focal point, prolonging the lifespan of focused waves and expanding their propagation space.

physics.flu-dyn

Geographically Weighted Regression Analysis for Spatial Economics Data: a Bayesian Recourse

The geographically weighted regression (GWR) is a well-known statistical approach to explore spatial non-stationarity of the regression relationship in spatial data analysis. In this paper, we discuss a Bayesian recourse of GWR. Bayesian variable selection based on spike-and-slab prior, bandwidth selection based on range prior, and model assessment using a modified deviance information criterion and a modified logarithm of pseudo-marginal likelihood are fully discussed in this paper. Usage of the graph distance in modeling areal data is also introduced. Extensive simulation studies are carried out to examine the empirical performance of the proposed methods with both small and large number of location scenarios, and comparison with the classical frequentist GWR is made. The performance of variable selection and estimation of the proposed methodology under different circumstances are satisfactory. We further apply the proposed methodology in analysis of a province-level macroeconomic data of 30 selected provinces in China. The estimation and variable selection results reveal insights about China's economy that are convincing and agree with previous studies and facts.

stat.AP

Bayesian Hierarchical Spatial Regression Models for Spatial Data in the Presence of Missing Covariates with Applications

In many applications, survey data are collected from different survey centers in different regions. It happens that in some circumstances, response variables are completely observed while the covariates have missing values. In this paper, we propose a joint spatial regression model for the response variable and missing covariates via a sequence of one-dimensional conditional spatial regression models. We further construct a joint spatial model for missing covariate data mechanisms. The properties of the proposed models are examined and a Markov chain Monte Carlo sampling algorithm is used to sample from the posterior distribution. In addition, the Bayesian model comparison criteria, the modified Deviance Information Criterion (mDIC) and the modified Logarithm of the Pseudo-Marginal Likelihood (mLPML), are developed to assess the fit of spatial regression models for spatial data. Extensive simulation studies are carried out to examine empirical performance of the proposed methods. We further apply the proposed methodology to analyze a real data set from a Chinese Health and Nutrition Survey (CHNS) conducted in 2011.

stat.ME

A Nonparametric Bayesian Item Response Modeling Approach for Clustering Items and Individuals Simultaneously

Item response theory (IRT) is a popular modeling paradigm for measuring subject latent traits and item properties according to discrete responses in tests or questionnaires. There are very limited discussions on heterogeneity pattern detection for both items and individuals. In this paper, we introduce a nonparametric Bayesian approach for clustering items and individuals simultaneously under the Rasch model. Specifically, our proposed method is based on the mixture of finite mixtures (MFM) model. MFM obtains the number of clusters and the clustering configurations for both items and individuals simultaneously. The performance of parameters estimation and parameters clustering under the MFM Rasch model is evaluated by simulation studies, and a real date set is applied to illustrate the MFM Rasch modeling.

stat.AP

Heterogeneous Regression Models for Clusters of Spatial Dependent Data

In economic development, there are often regions that share similar economic characteristics, and economic models on such regions tend to have similar covariate effects. In this paper, we propose a Bayesian clustered regression for spatially dependent data in order to detect clusters in the covariate effects. Our proposed method is based on the Dirichlet process which provides a probabilistic framework for simultaneous inference of the number of clusters and the clustering configurations. The usage of our method is illustrated both in simulation studies and an application to a housing cost dataset of Georgia.

econ.EM