SearcharxivSearch

arXiv subjects

Shanpeng Li

Publications and source records attributed to Shanpeng Li.

6 recordsLinked to original sources

FastJM: An R Package for Efficient Implementation of Semiparametric Joint Models for Longitudinal and Survival Data

Joint models provide a flexible framework for characterizing the association between longitudinal and time-to-event processes and have been widely applied in biomedical research. However, fitting joint models can be computationally challenging for large-scale and complex biomedical data. This paper introduces the \proglang{R} package \pkg{FastJM}, which provides computationally efficient frequentist estimation for three classes of semiparametric joint models: joint models with a single longitudinal biomarker, joint models with multiple longitudinal biomarkers, and joint models with a single longitudinal biomarker with heterogeneous within-subject (WS) variability. Within an expectation--maximization framework, \pkg{FastJM} employs customized linear-scan algorithms to efficiently update the nonparametric baseline hazards, thereby addressing a major computational bottleneck in semiparametric joint modeling. The package also supports commonly used time-dependent latent association structures by integrating these algorithms with a landmark multivariate joint modeling framework. \pkg{FastJM} provides a unified interface for model specification, estimation, inference, visualization, dynamic prediction, and prediction performance assessment, including cross-validated time-dependent accuracy measures and time-independent concordance statistics. We describe the underlying methodology and software implementation and demonstrate the main functionality of \pkg{FastJM} through reproducible examples.

stat.ME

Efficient Implementation of a Semiparametric Joint Model for Multivariate Longitudinal Biomarkers and Competing Risks Time-to-Event Data

Joint modeling has become increasingly popular for characterizing the association between one or more longitudinal biomarkers and competing risks time-to-event outcomes. However, semiparametric multivariate joint modeling for large-scale data encounter substantial statistical and computational challenges, primarily due to the high dimensionality of random effects and the complexity of estimating nonparametric baseline hazards. These challenges often lead to prolonged computation time and excessive memory usage, limiting the utility of joint modeling for biobank-scale datasets. In this article, we introduce an efficient implementation of a semiparametric multivariate joint model, supported by a normal approximation and customized linear scan algorithms within an expectation-maximization (EM) framework. Our method significantly reduces computation time and memory consumption, enabling the analysis of data from thousands of subjects. The scalability and estimation accuracy of our approach are demonstrated through two simulation studies. We also present an application to the Primary Biliary Cirrhosis (PBC) dataset involving five longitudinal biomarkers as an illustrative example. A user-friendly R package, \texttt{FastJM}, has been developed for the shared random effects joint model with efficient implementation. The package is publicly available on the Comprehensive R Archive Network: https://CRAN.R-project.org/package=FastJM.

stat.ME

PDXpower: A Power Analysis Tool for Experimental Design in Pre-clinical Xenograft Studies for Uncensored and Censored Outcomes

In cancer research, leveraging patient-derived xenografts (PDXs) in pre-clinical experiments is a crucial approach for assessing innovative therapeutic strategies. Addressing the inherent variability in treatment response among and within individual PDX lines is essential. However, the current literature lacks a user-friendly statistical power analysis tool capable of concurrently determining the required number of PDX lines and animals per line per treatment group in this context. In this paper, we present a simulation-based R package for sample size determination, named `\textbf{PDXpower}', which is publicly available at The Comprehensive R Archive Network \url{https://CRAN.R-project.org/package=PDXpower}. The package is designed to estimate the necessary number of both PDX lines and animals per line per treatment group for the design of a PDX experiment, whether for an uncensored outcome, or a censored time-to-event outcome. Our sample size considerations rely on two widely used analytical frameworks: the mixed effects ANOVA model for uncensored outcomes and Cox's frailty model for censored data outcomes, which effectively account for both inter-PDX variability and intra-PDX correlation in treatment response. Step-by-step illustrations for utilizing the developed package are provided, catering to scenarios with or without preliminary data.

stat.AP

A joint model of the individual mean and within-subject variability of a longitudinal outcome with a competing risks time-to-event outcome

Motivated by a growing body of research emphasizing the importance of modeling within-subject (WS) variability in longitudinal biomarkers and its association with health outcomes, this paper proposes a semiparametric joint model for both the mean and WS variability of a longitudinal biomarker, jointly with competing-risk time-to-event outcomes. We derive an expectation-maximization algorithm for parameter estimation and a profile-likelihood method for standard error estimation and inference, which allows time-dependent covariates and general forms of the latent association structure. Furthermore, we optimize the implementation of our joint model when the survival submodel includes only time-independent baseline covariates and shared random effects, allowing it to scale effectively to biobank-scale data involving tens of thousands of subjects. Our method demonstrates satisfactory performance in simulations, whereas classical joint models that assume homogeneous WS variability may suffer from substantial estimation bias, invalid inference, and inferior prediction when confronted with heterogeneous WS variability. We illustrate the utility of our method using the Multi-Ethnic Study of Atherosclerosis (MESA) cohort. Our analysis demonstrates that associations between WS blood pressure variability and cardiovascular outcomes, previously observed in clinical trials involving relatively homogeneous populations, extend to a more ethnically diverse and generally healthier cohort, and that explicitly modeling heterogeneous WS variability substantially enhances risk discrimination. A user-friendly R package, \textbf{JMH}, has been developed for the proposed shared random effects model with efficient implementation and is publicly available on the Comprehensive R Archive Network https://CRAN.R-project.org/package=JMH.

stat.ME

Efficient Algorithms and Implementation of a Semiparametric Joint Model for Longitudinal and Competing Risks Data: With Applications to Massive Biobank Data

Semiparametric joint models of longitudinal and competing risks data are computationally costly and their current implementations do not scale well to massive biobank data. This paper identifies and addresses some key computational barriers in a semiparametric joint model for longitudinal and competing risks survival data. By developing and implementing customized linear scan algorithms, we reduce the computational complexities from $O(n^2)$ or $O(n^3)$ to $O(n)$ in various components including numerical integration, risk set calculation, and standard error estimation, where $n$ is the number of subjects. Using both simulated and real world biobank data, we demonstrate that these linear scan algorithms generate drastic speed-up of up to hundreds of thousands fold when $n>10^4$, sometimes reducing the run-time from days to minutes. We have developed an R-package, FastJM, based on the proposed algorithms for joint modeling of longitudinal and time-to-event data with and without competing risks, and made it publicly available on the Comprehensive R Archive Network (CRAN).

stat.ME

A Flexible Joint Model for Multiple Longitudinal Biomarkers and A Time-to-Event Outcome: With Applications to Dynamic Prediction Using Highly Correlated Biomarkers

In biomedical studies it is common to collect data on multiple biomarkers during study follow-up for dynamic prediction of a time-to-event clinical outcome. The biomarkers are typically intermittently measured, missing at some event times, and may be subject to high biological variations, which cannot be readily used as time-dependent covariates in a standard time-to-event model. Moreover, they can be highly correlated if they are from in the same biological pathway. To address these issues, we propose a flexible joint model framework that models the multiple biomarkers with a shared latent reduced rank longitudinal principal component model and correlates the latent process to the event time by the Cox model for dynamic prediction of the event time. The proposed joint model for highly correlated biomarkers is more flexible than some existing methods since the latent trajectory shared by the multiple biomarkers does not require specification of a priori parametric time trend and is determined by data. We derive an Expectation-Maximization (EM) algorithm for parameter estimation, study large sample properties of the estimators, and adapt the developed method to make dynamic prediction of the time-to-event outcome. Bootstrap is used for standard error estimation and inference. The proposed method is evaluated using simulations and illustrated on a lung transplant data to predict chronic lung allograft dysfunction (CLAD) using chemokines measured in bronchoalveolar lavage fluid of the patients.

stat.ME