Searcharxiv⌕ Search

arXiv subjects

Zhi Yang Tho

Publications and source records attributed to Zhi Yang Tho.

5 recordsLinked to original sources

Gradient Boosted Mixed Models: Flexible Estimation of Mean and Variance Components for Clustered Data

We introduce a novel way to combine gradient boosting with mixed effects models, whereby the mean and variance components are learned jointly as functions of covariates via likelihood-based gradients. Gradient Boosted Mixed Models (GBMixed) estimates a nonparametric fixed effects function characterizing the overall mean of the response, while also allowing the random effects covariance matrix along with the residual variance to depend on covariates in a flexible manner. We demonstrate how GBMixed facilitates covariate-dependent random effect predictions, and subsequently point predictions and prediction intervals for individual treatment effects, that can adapt between population-level and cluster-level information. Experiments and applications to two real-world datasets show that GBMixed can accurately recover complex nonlinear fixed effect functions and covariate-dependent covariances in a linear mixed model, while also improving point and probabilistic predictive performance compared with several existing approaches such as parametric linear mixed models, Natural Gradient Boosting, and Gaussian Process Boosting. In simulations where the variance components are designed to vary as a function of covariates, GBMixed reduces the mean squared error in recovering the random effects variance function by a factor of eight relative to linear mixed models and Gaussian Process Boosting.

stat.ML↗

A Proportional Random Effect Block Bootstrap for General Clustered Data

Clustered data arise naturally in many scientific and applied research settings where units are grouped within clusters. Such data are commonly analyzed using linear mixed models to account for within-cluster correlations. This article proposes a proportional random effect block bootstrap applicable to general linear mixed model settings with imbalanced cluster sizes, both random intercepts and random slopes, and autocorrelation within clusters, while allowing for non-normal random effect and error distributions. It generalizes the original random effect block bootstrap, which was developed for more restrictive settings with balanced cluster sizes, random intercepts only, and constant within-cluster correlation. The proposed bootstrap is shown to be Fisher consistent under these more general settings. Simulations demonstrate strong finite sample inferential performance relative to the original random effect block bootstrap and several existing bootstrap methods for clustered data across a variety of scenarios. Application to the Mayo Clinic primary biliary cirrhosis dataset, which contains cluster sizes ranging from 1 to 16 and exhibits evidence of within-cluster autocorrelation and non-normality, further illustrates improved bootstrap confidence intervals using the proposed method.

stat.ME↗

Bias-Adjusted Attribution Estimation for Rainfall Enhancement Trials

Model-based analyses of rainfall enhancement trial data typically involve modelling log-transformed rainfall using linear mixed models to assess the effectiveness of enhancement methods under real-world conditions. This approach improves on traditional average-based analyses by allowing explicit control for the effects of meteorological and topographical covariates that may affect precipitation amounts. However, a key issue with such analyses is the bias that arises when back-transforming the log-rainfall to the original scale for estimating attribution, defined as the additional raw-scale rainfall attributable to the enhancement method. To address this issue, we propose a new attribution estimator that incorporates theoretically justified, observation-specific bias-adjustment terms. The proposed estimator improves upon existing estimators that rely on arbitrary adjustments, and satisfies a coherence property that ensures zero estimated attribution for observations without enhancement intervention. A proportional random effect block bootstrap is further used to conduct inference on the attribution quantities. Applying both the proposed estimator and an existing estimator to the Oman rainfall enhancement trial from 2013 to 2018, we find statistically significant positive effect of the ground-based ionization technology on downwind rainfall at the 5% significance level, with our proposed estimator indicating a smaller effect than the existing method. A simulation study further support the findings based on the proposed estimator, demonstrating its superior estimation accuracy and improved inferential performance of the associated bootstrap confidence intervals.

stat.ME↗

Joint Mean and Correlation Regression Models for Multivariate Data

We propose a joint mean and correlation regression model for multivariate discrete and (semi-)continuous response data, that simultaneously regresses the mean of each response against a set of covariates, and the correlations between responses against a set of similarity/distance measures. A set of joint estimating equations are formulated to construct an estimator of both the mean regression coefficients and the correlation regression parameters. Under a general setting where the number of responses can tend to infinity, the joint estimator is demonstrated to be consistent and asymptotically normally distributed, with differing rates of convergence due to the mean regression coefficients being heterogeneous across responses. An iterative estimation procedure is developed to obtain parameter estimates in the required (constrained) parameter space. Simulations demonstrate the strong finite sample performance of the proposed estimator in terms of point estimation and inference. We apply the proposed model to a count dataset of 38 Carabidae ground beetle species sampled throughout Scotland, along with information about the environmental conditions of each site and the traits of each species. Results show the relationship between mean abundance and environmental covariates differs across the beetle species, and that beetle total length is important in driving the correlations between species.

stat.ME↗

An Ising Similarity Regression Model for Modeling Multivariate Binary Data

Understanding the dependence structure between response variables is an important component in the analysis of correlated multivariate data. This article focuses on modeling dependence structures in multivariate binary data, motivated by a study aiming to understand how patterns in different U.S. senators' votes are determined by similarities (or lack thereof) in their attributes, e.g., political parties and social network profiles. To address such a research question, we propose a new Ising similarity regression model which regresses pairwise interaction coefficients in the Ising model against a set of similarity measures available/constructed from covariates. Model selection approaches are further developed through regularizing the pseudo-likelihood function with an adaptive lasso penalty to enable the selection of relevant similarity measures. We establish estimation and selection consistency of the proposed estimator under a general setting where the number of similarity measures and responses tend to infinity. Simulation study demonstrates the strong finite sample performance of the proposed estimator, particularly compared with several existing Ising model estimators in estimating the matrix of pairwise interaction coefficients. Applying the Ising similarity regression model to a dataset of roll call voting records of 100 U.S. senators, we are able to quantify how similarities in senators' parties, businessman occupations and social network profiles drive their voting associations.

stat.ME↗