SearcharxivSearch

arXiv subjects

Taehwa Choi

Publications and source records attributed to Taehwa Choi.

5 recordsLinked to original sources

Censored broken adaptive ridge rank regression via induced smoothing

Broken adaptive ridge (BAR) penalty approximates $L_0$-regularization through iterative reweighting of L2 penalties. This penalty enjoys both the oracle property and the grouping effect for highly correlated covariates, making it particularly attractive for penalized regression with complex dependence among predictors. In this paper, we develop a BAR-penalized linear rank regression method for the semiparametric accelerated failure time model with right-censored data. Computational tractability is achieved by applying induced smoothing to the nonsmooth Gehan-type rank estimating function, yielding a more stable framework for estimation and inference. For scalable penalization, we develop a cyclic coordinate descent algorithm that minimizes the penalized objective function, and estimates the regression coefficients in a coordinate-wise manner. We further extend the proposed method to more complex survival endpoints, such as multivariate partly interval-censored (PIC) data. Under mild conditions, the proposed estimator satisfies both the oracle property and the grouping effect, and the variance estimator of the informative coefficients can be derived in analytic form. Numerical studies using synthetic data compare our approach to several well-known penalties, and demonstrate its superior selection accuracy and estimation efficiency across various scenarios. Furthermore, applications to right-censored outcomes from primary biliary cirrhosis, and correlated PIC outcomes from colorectal cancer further illustrate the practical utility of the proposed method. The R package aftPenCDA for implementing the method is available on R CRAN.

stat.ME

Cox Model Predicting Covariate Subject to Right Censoring

Time-to-event endpoints are frequently used as outcomes in oncology and other disease areas where the outcome of interest may not be observed within a predetermined period. Although many analytical methods address the challenges of censoring in outcomes, limited research has focused on censored covariates. Conventional methods such as the complete case (CC) analysis, where data from patients with censored covariates are discarded, suffer from efficiency loss and potential bias due to reduced sample size. Alternatively, imputing censored covariates with a constant value can underestimate variability. Recognizing these limitations, novel estimation procedures within the generalized linear model framework have been proposed, with some research emerging in time-to-event outcomes. In this paper, we investigate the association between progression-free survival and overall survival using a semi-parametric Cox model framework. We modify the Cox model's partial likelihood function to account for censored covariates by replacing the relative risk associated with censored covariates with a weighted average of patients with observed covariates. The performance of the proposed method is demonstrated through simulations and applications to two oncology clinical trials. Results indicate that the proposed method offers improved estimation efficiency and better utilization of available data compared to other approaches.

stat.ME

Rank estimation for the accelerated failure time model with partially interval-censored data

This paper presents a unified rank-based inferential procedure for fitting the accelerated failure time model to partially interval-censored data. A Gehan-type monotone estimating function is constructed based on the idea of the familiar weighted log-rank test, and an extension to a general class of rank-based estimating functions is suggested. The proposed estimators can be obtained via linear programming and are shown to be consistent and asymptotically normal via standard empirical process theory. Unlike common maximum likelihood-based estimators for partially interval-censored regression models, our approach can directly provide a regression coefficient estimator without involving a complex nonparametric estimation of the underlying residual distribution function. An efficient variance estimation procedure for the regression coefficient estimator is considered. Moreover, we extend the proposed rank-based procedure to the linear regression analysis of multivariate clustered partially interval-censored data. The finite-sample operating characteristics of our approach are examined via simulation studies. Data example from a colorectal cancer study illustrates the practical usefulness of the method.

stat.ME

Interval-censored linear quantile regression

Censored quantile regression has emerged as a prominent alternative to classical Cox's proportional hazards model or accelerated failure time model in both theoretical and applied statistics. While quantile regression has been extensively studied for right-censored survival data, methodologies for analyzing interval-censored data remain limited in the survival analysis literature. This paper introduces a novel local weighting approach for estimating linear censored quantile regression, specifically tailored to handle diverse forms of interval-censored survival data. The estimation equation and the corresponding convex objective function for the regression parameter can be constructed as a weighted average of quantile loss contributions at two interval endpoints. The weighting components are nonparametrically estimated using local kernel smoothing or ensemble machine learning techniques. To estimate the nonparametric distribution mass for interval-censored data, a modified EM algorithm for nonparametric maximum likelihood estimation is employed by introducing subject-specific latent Poisson variables. The proposed method's empirical performance is demonstrated through extensive simulation studies and real data analyses of two HIV/AIDS datasets.

stat.ME

Towards Flexible Time-to-event Modeling: Optimizing Neural Networks via Rank Regression

Time-to-event analysis, also known as survival analysis, aims to predict the time of occurrence of an event, given a set of features. One of the major challenges in this area is dealing with censored data, which can make learning algorithms more complex. Traditional methods such as Cox's proportional hazards model and the accelerated failure time (AFT) model have been popular in this field, but they often require assumptions such as proportional hazards and linearity. In particular, the AFT models often require pre-specified parametric distributional assumptions. To improve predictive performance and alleviate strict assumptions, there have been many deep learning approaches for hazard-based models in recent years. However, representation learning for AFT has not been widely explored in the neural network literature, despite its simplicity and interpretability in comparison to hazard-focused methods. In this work, we introduce the Deep AFT Rank-regression model for Time-to-event prediction (DART). This model uses an objective function based on Gehan's rank statistic, which is efficient and reliable for representation learning. On top of eliminating the requirement to establish a baseline event time distribution, DART retains the advantages of directly predicting event time in standard AFT models. The proposed method is a semiparametric approach to AFT modeling that does not impose any distributional assumptions on the survival time distribution. This also eliminates the need for additional hyperparameters or complex model architectures, unlike existing neural network-based AFT models. Through quantitative analysis on various benchmark datasets, we have shown that DART has significant potential for modeling high-throughput censored time-to-event data.

cs.LG