SearcharxivSearch

arXiv subjects

Jasen Zhang

Publications and source records attributed to Jasen Zhang.

2 recordsLinked to original sources

Scaling Hawkes Processes

Hawkes processes (HP) are a large class of stochastic point process models scientists have used to analyze contagion phenomena ranging from earthquakes, infectious diseases and biological neurons to financial trading activity, memes on social media and gun violence. We introduce applications of HP to the latter before reviewing general strategies for fitting HP to data, paying attention to the influence of model structure on computational scalability considerations. We then apply a recently developed high-performance computing powered Bayesian inference strategy for the spatiotemporal HP analysis of 412,376 acts of gun violence in the U.S. between 2014 and 2024. We finish with a discussion of model fit and directions for future research.

stat.CO

Semiparametric Regression Models for Explanatory Variables with Missing Data due to Detection Limit

Detection limit (DL) has become an increasingly ubiquitous issue in statistical analyses of biomedical studies, such as cytokine, metabolite and protein analysis. In regression analysis, if an explanatory variable is left-censored due to concentrations below the DL, one may limit analyses to observed data. In many studies, additional, or surrogate, variables are available to model, and incorporating such auxiliary modeling information into the regression model can improve statistical power. Although methods have been developed along this line, almost all are limited to parametric models for both the regression and left-censored explanatory variable. While some recent work has considered semiparametric regression for the censored DL-effected explanatory variable, the regression of primary interest is still left parametric, which not only makes it prone to biased estimates, but also suffers from high computational cost and inefficiency due to maximizing an extremely complex likelihood function and bootstrap inference. In this paper, we propose a new approach by considering semiparametric generalized linear models (SPGLM) for the primary regression and parametric or semiparametric models for DL-effected explanatory variable. The semiparametric and semiparametric combination provides the most robust inference, while the semiparametric and parametric case enables more efficient inference. The proposed approach is also much easier to implement and allows for leveraging sample splitting and cross fitting (SSCF) to improve computational efficiency in variance estimation. In particular, our approach improves computational efficiency over bootstrap by 450 times. We use simulated and real study data to illustrate the approach.

stat.ME