SearcharxivSearch

arXiv subjects

Yishu Xue

Publications and source records attributed to Yishu Xue.

13 recordsLinked to original sources

Heteroscedastic Growth Curve Modeling with Shape-Restricted Splines

Growth curve analysis (GCA) has a wide range of applications in various fields where growth trajectories need to be modeled. Heteroscedasticity is often present in the error term, which can not be handled with sufficient flexibility by standard linear fixed or mixed-effects models. One situation that has been addressed is where the error variance is characterized by a linear predictor with certain covariates. A frequently encountered scenario in GCA, however, is one in which the variance is a smooth function of the mean with known shape restrictions. A naive application of standard linear mixed-effects models would underestimate the variance of the fixed effects estimators and, consequently, the uncertainty of the estimated growth curve. We propose to model the variance of the response variable as a shape-restricted (increasing/decreasing; convex/concave) function of the marginal or conditional mean using shape-restricted splines. A simple iteratively reweighted fitting algorithm that takes advantage of existing software for linear mixed-effects models is developed. For inference, a parametric bootstrap procedure is recommended. Our simulation study shows that the proposed method gives satisfactory inference with moderate sample sizes. The utility of the method is demonstrated using two real-world applications.

stat.ME

Multidimensional heterogeneity learning for count value tensor data with applications to field goal attempt analysis of NBA players

We propose a multidimensional tensor clustering approach for studying how professional basketball players' shooting patterns vary over court locations and game time. Unlike most existing methods that only study continuous-valued tensors or have to assume the same cluster structure along different tensor directions, we propose a Bayesian nonparametric model that deals with count-valued tensors and projects the heterogeneity among players onto tensor dimensions while allowing cluster structures to be different over directions. Our method is fully probabilistic; hence allows simultaneous inference on both the number of clusters and the cluster configurations. We present an efficient posterior sampling method and establish the large-sample convergence properties for the posterior distribution. Simulation studies have demonstrated an excellent empirical performance of the proposed method. Finally, an application to shot chart data collected from 191 NBA players during the 2017-2018 regular season is conducted and reveals several interesting insights for basketball analytics.

stat.ME

Bayesian Clustered Coefficients Regression with Auxiliary Covariates Assistant Random Effects

In regional economics research, a problem of interest is to detect similarities between regions, and estimate their shared coefficients in economics models. In this article, we propose a mixture of finite mixtures (MFM) clustered regression model with auxiliary covariates that account for similarities in demographic or economic characteristics over a spatial domain. Our Bayesian construction provides both inference for number of clusters and clustering configurations, and estimation for parameters for each cluster. Empirical performance of the proposed model is illustrated through simulation experiments, and further applied to a study of influential factors for monthly housing cost in Georgia.

stat.ME

Bayesian Spatial Homogeneity Pursuit of Functional Data: an Application to the U.S. Income Distribution

An income distribution describes how an entity's total wealth is distributed amongst its population. A problem of interest to regional economics researchers is to understand the spatial homogeneity of income distributions among different regions. In economics, the Lorenz curve is a well-known functional representation of income distribution. In this article, we propose a mixture of finite mixtures (MFM) model as well as a Markov random field constrained mixture of finite mixtures (MRFC-MFM) model in the context of spatial functional data analysis to capture spatial homogeneity of Lorenz curves. We design efficient Markov chain Monte Carlo (MCMC) algorithms to simultaneously infer the posterior distributions of the number of clusters and the clustering configuration of spatial functional data. Extensive simulation studies are carried out to show the effectiveness of the proposed methods compared with existing methods. We apply the proposed spatial functional clustering method to state level income Lorenz curves from the American Community Survey Public Use Microdata Sample (PUMS) data. The results reveal a number of important clustering patterns of state-level income distributions across US.

stat.AP

Zero Inflated Poisson Model with Clustered Regression Coefficients: an Application to Heterogeneity Learning of Field Goal Attempts of Professional Basketball Players

Although basketball is a dynamic process sport, with 5 plus 5 players competing on both offense and defense simultaneously, learning some static information is predominant for professional players, coaches and team mangers. In order to have a deep understanding of field goal attempts among different players, we propose a zero inflated Poisson model with clustered regression coefficients to learn the shooting habits of different players over the court and the heterogeneity among them. Specifically, the zero inflated model recovers the large proportion of the court with zero field goal attempts, and the mixture of finite mixtures model learn the heterogeneity among different players based on clustered regression coefficients and inflated probabilities. Both theoretical and empirical justification through simulation studies validate our proposed method. We apply our proposed model to the National Basketball Association (NBA), for learning players' shooting habits and heterogeneity among different players over the 2017--2018 regular season. This illustrates our model as a way of providing insights from different aspects.

stat.AP

Time Fused Coefficient SIR Model with Application to COVID-19 Epidemic in the United States

In this paper, we propose a Susceptible-Infected-Removal (SIR) model with time fused coefficients. In particular, our proposed model discovers the underlying time homogeneity pattern for the SIR model's transmission rate and removal rate via Bayesian shrinkage priors. MCMC sampling for the proposed method is facilitated by the nimble package in R. Extensive simulation studies are carried out to examine the empirical performance of the proposed methods. We further apply the proposed methodology to analyze different levels of COVID-19 data in the United States.

stat.AP

An Online Updating Approach for Testing the Proportional Hazards Assumption with Streams of Survival Data

The Cox model, which remains as the first choice in analyzing time-to-event data even for large datasets, relies on the proportional hazards (PH) assumption. When survival data arrive sequentially in chunks, a fast and minimally storage intensive approach to test the PH assumption is desirable. We propose an online updating approach that updates the standard test statistic as each new block of data becomes available, and greatly lightens the computational burden. Under the null hypothesis of PH, the proposed statistic is shown to have the same asymptotic distribution as the standard version computed on the entire data stream with the data blocks pooled into one dataset. In simulation studies, the test and its variant based on most recent data blocks maintain their sizes when the PH assumption holds and have substantial power to detect different violations of the PH assumption. We also show in simulation that our approach can be used successfully with "big data" that exceed a single computer's computational resources. The approach is illustrated with the survival analysis of patients with lymphoma cancer from the Surveillance, Epidemiology, and End Results Program. The proposed test promptly identified deviation from the PH assumption that was not captured by the test based on the entire data.

stat.ME

Bayesian Group Learning for Shot Selection of Professional Basketball Players

In this paper, we develop a group learning approach to analyze the underlying heterogeneity structure of shot selection among professional basketball players in the NBA. We propose a mixture of finite mixtures (MFM) model to capture the heterogeneity of shot selection among different players based on Log Gaussian Cox process (LGCP). Our proposed method can simultaneously estimate the number of groups and group configurations. An efficient Markov Chain Monte Carlo (MCMC) algorithm is developed for our proposed model. Simulation studies have been conducted to demonstrate its performance. Ultimately, our proposed learning approach is further illustrated in analyzing shot charts of several players in the NBA's 2017-2018 regular season.

stat.AP

Geographically Weighted Regression Analysis for Spatial Economics Data: a Bayesian Recourse

The geographically weighted regression (GWR) is a well-known statistical approach to explore spatial non-stationarity of the regression relationship in spatial data analysis. In this paper, we discuss a Bayesian recourse of GWR. Bayesian variable selection based on spike-and-slab prior, bandwidth selection based on range prior, and model assessment using a modified deviance information criterion and a modified logarithm of pseudo-marginal likelihood are fully discussed in this paper. Usage of the graph distance in modeling areal data is also introduced. Extensive simulation studies are carried out to examine the empirical performance of the proposed methods with both small and large number of location scenarios, and comparison with the classical frequentist GWR is made. The performance of variable selection and estimation of the proposed methodology under different circumstances are satisfactory. We further apply the proposed methodology in analysis of a province-level macroeconomic data of 30 selected provinces in China. The estimation and variable selection results reveal insights about China's economy that are convincing and agree with previous studies and facts.

stat.AP

Heterogeneous Regression Models for Clusters of Spatial Dependent Data

In economic development, there are often regions that share similar economic characteristics, and economic models on such regions tend to have similar covariate effects. In this paper, we propose a Bayesian clustered regression for spatially dependent data in order to detect clusters in the covariate effects. Our proposed method is based on the Dirichlet process which provides a probabilistic framework for simultaneous inference of the number of clusters and the clustering configurations. The usage of our method is illustrated both in simulation studies and an application to a housing cost dataset of Georgia.

econ.EM

A comparison of Bayesian accelerated failure time models with spatially varying coefficients

The accelerated failure time (AFT) model is a commonly used tool in analyzing survival data. In public health studies, data is often collected from medical service providers in different locations. Survival rates from different locations often present geographically varying patterns. In this paper, we focus on the accelerated failure time model with spatially varying coefficients. We compare three types of the priors for spatially varying coefficients. A model selection criterion, logarithm of the pseudo-marginal likelihood (LPML), is developed to assess the fit of AFT model with different priors. Extensive simulation studies are carried out to examine the empirical performance of the proposed methods. Finally, we apply our model to SEER data on prostate cancer in Louisiana and demonstrate the existence of spatially varying effects on survival rates from prostate cancer data.

stat.AP

Spatial Weibull Regression with Multivariate Log Gamma Process and Its Applications to China Earthquake Economic Loss

Bayesian spatial modeling of heavy-tailed distributions has become increasingly popular in various areas of science in recent decades. We propose a Weibull regression model with spatial random effects for analyzing extreme economic loss. Model estimation is facilitated by a computationally efficient Bayesian sampling algorithm utilizing the multivariate Log-Gamma distribution. Simulation studies are carried out to demonstrate better empirical performances of the proposed model than the generalized linear mixed effects model. An earthquake data obtained from Yunnan Seismological Bureau, China is analyzed. Logarithm of the Pseudo-marginal likelihood values are obtained to select the optimal model, and Value-at-risk, expected shortfall, and tail-value-at-risk based on posterior predictive distribution of the optimal model are calculated under different confidence levels.

stat.AP

Geographically Weighted Cox Regression for Prostate Cancer Survival Data in Louisiana

The Cox proportional hazard model is one of the most popular tools in analyzing time-to-event data in public health studies. When outcomes observed in clinical data from different regions yield a varying pattern correlated with location, it is often of great interest to investigate spatially varying effects of covariates. In this paper, we propose a geographically weighted Cox regression model for sparse spatial survival data. In addition, a stochastic neighborhood weighting scheme is introduced at the county level. Theoretical properties of the proposed geographically weighted estimators are examined in detail. A model selection scheme based on the Takeuchi's model robust information criteria (TIC) is discussed. Extensive simulation studies are carried out to examine the empirical performance of the proposed methods. We further apply the proposed methodology to analyze real data on prostate cancer from the Surveillance, Epidemiology, and End Results cancer registry for the state of Louisiana.

stat.AP