Searcharxiv⌕ Search

arXiv subjects

Vladimir Svetnik

Publications and source records attributed to Vladimir Svetnik.

7 recordsLinked to original sources

Distribution-Free Selection of Low-Risk Oncology Patients for Survival Beyond a Time Horizon

We study how to select a subset of patients who are unlikely to experience an adverse event within a given time horizon, by calibrating a screening rule based on the output of any survival model. We consider two complementary frameworks. The first extends the classical idea of estimating the event rate among selected patients using a hold-out dataset, integrating it with the Learn-Then-Test method. This provides approximate high-probability guarantees that are comparable to those obtainable from simultaneous confidence bands while being often less conservative. The second takes a different perspective by reformulating the problem in terms of multiple hypotheses testing, enabling false discovery rate (FDR) control via the Benjamini-Hochberg procedure applied to selective conformal p-values. This provides approximate guarantees in expectation. We clarify the theoretical relationship between these approaches, explain how both can handle right-censored data and be made doubly robust via augmented inverse probability of censoring weighting, and compare them empirically using simulations and oncology data from the Flatiron Health Research Database. Our results reveal a practical trade-off: FDR-based screening is typically more powerful, while high-probability calibration is more conservative but offers stronger guarantees, especially when few patients are selected. We also provide practical guidance on implementation and tuning.

stat.AP↗

Conformal Prediction for Regression with Clipped Outcomes

We study conformal prediction for regression using calibration data with outcomes that are doubly censored (clipped) at known fixed thresholds. We show that existing methods are unsatisfactory in this setting, as they yield intervals that may have higher marginal coverage than desired and yet lose conditional coverage precisely for the easier-to-predict cases whose outcomes are typically fully observed. This reveals that marginal coverage, the usual target of conformal prediction, may not be the ideal goal under clipping. We address this challenge by introducing a new nonconformity score and calibration methods at both ends of this trade-off: one for tight marginal coverage, and a two-step method that prioritizes conditional coverage. We characterize their finite-sample coverage and oracle-like asymptotic behavior under suitable consistency of the underlying model, and we compare them to more direct adaptations of existing approaches.

stat.ME↗

Conformal Survival Bands for Risk Screening under Right-Censoring

We propose a method to quantify uncertainty around individual survival distribution estimates using right-censored data, compatible with any survival model. Unlike classical confidence intervals, the survival bands produced by this method offer predictive rather than population-level inference, making them useful for personalized risk screening. For example, in a low-risk screening scenario, they can be applied to flag patients whose survival band at 12 months lies entirely above 50\%, while ensuring that at least half of flagged individuals will survive past that time on average. Our approach builds on recent advances in conformal inference and integrates ideas from inverse probability of censoring weighting and multiple testing with false discovery rate control. We provide asymptotic guarantees and show promising performance in finite samples with both simulated and real data.

stat.ME↗

Doubly Robust Conformalized Survival Analysis with Right-Censored Data

We present a conformal inference method for constructing lower prediction bounds for survival times from right-censored data, extending recent approaches designed for more restrictive type-I censoring scenarios. The proposed method imputes unobserved censoring times using a machine learning model, and then analyzes the imputed data using a survival model calibrated via weighted conformal inference. This approach is theoretically supported by an asymptotic double robustness property. Empirical studies on simulated and real data demonstrate that our method leads to relatively informative predictive inferences and is especially robust in challenging settings where the survival model may be inaccurate.

stat.ME↗

Challenges in Variable Importance Ranking Under Correlation

Variable importance plays a pivotal role in interpretable machine learning as it helps measure the impact of factors on the output of the prediction model. Model agnostic methods based on the generation of "null" features via permutation (or related approaches) can be applied. Such analysis is often utilized in pharmaceutical applications due to its ability to interpret black-box models, including tree-based ensembles. A major challenge and significant confounder in variable importance estimation however is the presence of between-feature correlation. Recently, several adjustments to marginal permutation utilizing feature knockoffs were proposed to address this issue, such as the variable importance measure known as conditional predictive impact (CPI). Assessment and evaluation of such approaches is the focus of our work. We first present a comprehensive simulation study investigating the impact of feature correlation on the assessment of variable importance. We then theoretically prove the limitation that highly correlated features pose for the CPI through the knockoff construction. While we expect that there is always no correlation between knockoff variables and its corresponding predictor variables, we prove that the correlation increases linearly beyond a certain correlation threshold between the predictor variables. Our findings emphasize the absence of free lunch when dealing with high feature correlation, as well as the necessity of understanding the utility and limitations behind methods in variable importance estimation.

stat.ML↗

Development and Evaluation of Conformal Prediction Methods for QSAR

The quantitative structure-activity relationship (QSAR) regression model is a commonly used technique for predicting biological activities of compounds using their molecular descriptors. Predictions from QSAR models can help, for example, to optimize molecular structure; prioritize compounds for further experimental testing; and estimate their toxicity. In addition to the accurate estimation of the activity, it is highly desirable to obtain some estimate of the uncertainty associated with the prediction, e.g., calculate a prediction interval (PI) containing the true molecular activity with a pre-specified probability, say 70%, 90% or 95%. The challenge is that most machine learning (ML) algorithms that achieve superior predictive performance require some add-on methods for estimating uncertainty of their prediction. The development of these algorithms is an active area of research by statistical and ML communities but their implementation for QSAR modeling remains limited. Conformal prediction (CP) is a promising approach. It is agnostic to the prediction algorithm and can produce valid prediction intervals under some weak assumptions on the data distribution. We proposed computationally efficient CP algorithms tailored to the most advanced ML models, including Deep Neural Networks and Gradient Boosting Machines. The validity and efficiency of proposed conformal predictors are demonstrated on a diverse collection of QSAR datasets as well as simulation studies.

q-bio.BM↗

The Reciprocal Bayesian LASSO

A reciprocal LASSO (rLASSO) regularization employs a decreasing penalty function as opposed to conventional penalization approaches that use increasing penalties on the coefficients, leading to stronger parsimony and superior model selection relative to traditional shrinkage methods. Here we consider a fully Bayesian formulation of the rLASSO problem, which is based on the observation that the rLASSO estimate for linear regression parameters can be interpreted as a Bayesian posterior mode estimate when the regression parameters are assigned independent inverse Laplace priors. Bayesian inference from this posterior is possible using an expanded hierarchy motivated by a scale mixture of double Pareto or truncated normal distributions. On simulated and real datasets, we show that the Bayesian formulation outperforms its classical cousin in estimation, prediction, and variable selection across a wide range of scenarios while offering the advantage of posterior inference. Finally, we discuss other variants of this new approach and provide a unified framework for variable selection using flexible reciprocal penalties. All methods described in this paper are publicly available as an R package at: https://github.com/himelmallick/BayesRecipe.

stat.ME↗