SearcharxivSearch

arXiv subjects

Xinping Cui

Publications and source records attributed to Xinping Cui.

5 recordsLinked to original sources

FastJM: An R Package for Efficient Implementation of Semiparametric Joint Models for Longitudinal and Survival Data

Joint models provide a flexible framework for characterizing the association between longitudinal and time-to-event processes and have been widely applied in biomedical research. However, fitting joint models can be computationally challenging for large-scale and complex biomedical data. This paper introduces the \proglang{R} package \pkg{FastJM}, which provides computationally efficient frequentist estimation for three classes of semiparametric joint models: joint models with a single longitudinal biomarker, joint models with multiple longitudinal biomarkers, and joint models with a single longitudinal biomarker with heterogeneous within-subject (WS) variability. Within an expectation--maximization framework, \pkg{FastJM} employs customized linear-scan algorithms to efficiently update the nonparametric baseline hazards, thereby addressing a major computational bottleneck in semiparametric joint modeling. The package also supports commonly used time-dependent latent association structures by integrating these algorithms with a landmark multivariate joint modeling framework. \pkg{FastJM} provides a unified interface for model specification, estimation, inference, visualization, dynamic prediction, and prediction performance assessment, including cross-validated time-dependent accuracy measures and time-independent concordance statistics. We describe the underlying methodology and software implementation and demonstrate the main functionality of \pkg{FastJM} through reproducible examples.

stat.ME

Efficient Implementation of a Semiparametric Joint Model for Multivariate Longitudinal Biomarkers and Competing Risks Time-to-Event Data

Joint modeling has become increasingly popular for characterizing the association between one or more longitudinal biomarkers and competing risks time-to-event outcomes. However, semiparametric multivariate joint modeling for large-scale data encounter substantial statistical and computational challenges, primarily due to the high dimensionality of random effects and the complexity of estimating nonparametric baseline hazards. These challenges often lead to prolonged computation time and excessive memory usage, limiting the utility of joint modeling for biobank-scale datasets. In this article, we introduce an efficient implementation of a semiparametric multivariate joint model, supported by a normal approximation and customized linear scan algorithms within an expectation-maximization (EM) framework. Our method significantly reduces computation time and memory consumption, enabling the analysis of data from thousands of subjects. The scalability and estimation accuracy of our approach are demonstrated through two simulation studies. We also present an application to the Primary Biliary Cirrhosis (PBC) dataset involving five longitudinal biomarkers as an illustrative example. A user-friendly R package, \texttt{FastJM}, has been developed for the shared random effects joint model with efficient implementation. The package is publicly available on the Comprehensive R Archive Network: https://CRAN.R-project.org/package=FastJM.

stat.ME

Principles of Conditionality and Layering of Error Rates with Application to Platform Trials

There has been a misconception that only one type of error rate control is necessary in clinical trials, leading to debates over whether to prioritize Familywise Error Rate (FWER) or False Discovery Rate (FDR). This misconception has led to misleading statements about FWER control and proposals to shift towards FDR control, which could be manipulated by the industry. In reality, since the early 2000s, biopharmaceutical statistics have implicitly applied two layers of Type I error rate control. This aligns with Tukey's 1953 invention of Error Rate per Family (ERpF) for controlling error across studies, while FWER applies within each study. Our paper clarifies this layering, using Platform trials to demonstrate the verifiable conditions needed across studies for the FDA to fulfill its regulatory mission. We show that controlling FWER within a study at $5\%$ inherently controls ERpF across studies at 5-per-100, regardless of study correlations. This supports current regulatory practices that protect public health while fostering innovation. We also address concerns about ERpF stability in Platform trials, where shared controls introduce dependencies. By applying the Conditionality Principle and utilizing an innovative Shiny app, we explore how correlations impact ERpF variability, providing deeper insights for informed decision-making. Our findings, supported by principles like Layering of Error Rate Controls and the Conditionality Principle, are particularly relevant as Platform trials gain popularity for their efficiency in testing multiple treatments simultaneously.

stat.ME

A Machine Learning Approach to Galaxy-LSS Classification I: Imprints on Halo Merger Trees

The cosmic web plays a major role in the formation and evolution of galaxies and defines, to a large extent, their properties. However, the relation between galaxies and environment is still not well understood. Here we present a machine learning approach to study imprints of environmental effects on the mass assembly of haloes. We present a galaxy-LSS machine learning classifier based on galaxy properties sensitive to the environment. We then use the classifier to assess the relevance of each property. Correlations between galaxy properties and their cosmic environment can be used to predict galaxy membership to void/wall or filament/cluster with an accuracy of $93\%$. Our study unveils environmental information encoded in properties of haloes not normally considered directly dependent on the cosmic environment such as merger history and complexity. Understanding the physical mechanism by which the cosmic web is imprinted in a halo can lead to significant improvements in galaxy formation models. This is accomplished by extracting features from galaxy properties and merger trees, computing feature scores for each feature and then applying support vector machine to different feature sets. To this end, we have discovered that the shape and depth of the merger tree, formation time and density of the galaxy are strongly associated with the cosmic environment. We describe a significant improvement in the original classification algorithm by performing LU decomposition of the distance matrix computed by the feature vectors and then using the output of the decomposition as input vectors for support vector machine.

astro-ph.GA

Constrained Nonlinear and Mixed Effects Differential Equation Models for Dynamic Cell Polarity Signaling

The key of tip growth in eukaryotes is the polarized distribution on plasma membrane of a particle named ROP1. This distribution is the result of a positive feedback loop, whose mechanism can be described by a Differential Equation parametrized by two meaningful parameters kpf and knf . We introduce a mechanistic Integro-Differential Equation (IDE) derived from a spatiotemporal model of cell polarity and we show how this model can be fitted to real data, i.e., ROP1 intensities measured on pollen tubes. At first, we provide an existence and uniqueness result for the solution of our IDE model under certain conditions. Interestingly, this analysis gives a tractable expression for the likelihood, and our approach can be seen as the estimation of a constrained nonlinear model. Moreover, we introduce a population variability by a constrained nonlinear mixed model. We then propose a constrained Least Squares method to fit the model for the single pollen tube case, and two methods, constrained Methods of Moments and constrained Restricted Maximum Likelihood (REML) to fit the model for the multiple pollen tubes case. The performances of all three methods are studied through simulations and are used on an in-house multiple pollen tubes dataset generated at UC Riverside.

stat.ME