SearcharxivSearch

arXiv subjects

Mi-Ok Kim

Publications and source records attributed to Mi-Ok Kim.

5 recordsLinked to original sources

External Risk Prediction Informed Bayesian Survival Analysis

Prognostic factor evaluation and prediction model development are central to precision oncology, enabling patient risk stratification and individualized treatment selection. Unified predictions that synthesize information from existing models are valuable for comprehensive and consistent risk assessment. Many studies also seek to evaluate the incremental value of new biomarkers beyond established prognostic factors. However, such efforts are often constrained by small-to-moderate sample sizes. Motivated by these challenges, we consider Cox regression analysis in settings where individualized risk predictions from existing models are externally available without a transparent or interpretable structure, for example, through online calculators. We develop a Bayesian discretized survival time inference framework in which individualized predictions from potentially multiple external sources are integrated through a formulation based on Kullback-Leibler divergence, yielding informative priors. The divergence-based formulation serves as a surrogate for the external information likelihood, enabling principled incorporation of individualized predictions without requiring knowledge of the underlying external prediction models. Theoretical results show that the resulting posterior mean estimators are asymptotically more efficient than their internal-only maximum likelihood counterparts. However, using the divergence-based surrogate in place of the unavailable external likelihood renders posterior variance-based inference conservative. We propose a correction to address this overcoverage. We demonstrate the performance of the proposed approach through simulations and an application to prostate cancer trial data.

stat.ME

Decomposition of Longitudinal Disparities: an Application to the Fetal Growth-Singletons Study

Addressing health disparities among different demographic groups is a key challenge in public health. Despite many efforts, there is still a gap in understanding how these disparities unfold over time. Our paper focuses on this overlooked longitudinal aspect, which is crucial in both clinical and public health settings. In this paper, we introduce a longitudinal disparity decomposition method that decomposes disparities into three components: the explained disparity linked to differences in the exploratory variables' conditional distribution when the modifier distribution is identical between majority and minority groups, the explained disparity that emerges specifically from the unequal distribution of the modifier and its interaction with covariates, and the unexplained disparity. The proposed method offers a dynamic alternative to the traditional Peters-Belson decomposition approach, tackling both the potential reduction in disparity if the covariate distributions of minority groups matched those of the majority group and the evolving nature of disparity over time. We apply the proposed approach to a fetal growth study to gain insights into disparities between different race/ethnicity groups in fetal developmental progress throughout the course of pregnancy.

stat.AP

Designing monitoring strategies for deployed machine learning algorithms: navigating performativity through a causal lens

After a machine learning (ML)-based system is deployed, monitoring its performance is important to ensure the safety and effectiveness of the algorithm over time. When an ML algorithm interacts with its environment, the algorithm can affect the data-generating mechanism and be a major source of bias when evaluating its standalone performance, an issue known as performativity. Although prior work has shown how to validate models in the presence of performativity using causal inference techniques, there has been little work on how to monitor models in the presence of performativity. Unlike the setting of model validation, there is much less agreement on which performance metrics to monitor. Different monitoring criteria impact how interpretable the resulting test statistic is, what assumptions are needed for identifiability, and the speed of detection. When this choice is further coupled with the decision to use observational versus interventional data, ML deployment teams are faced with a multitude of monitoring options. The aim of this work is to highlight the relatively under-appreciated complexity of designing a monitoring strategy and how causal reasoning can provide a systematic framework for choosing between these options. As a motivating example, we consider an ML-based risk prediction algorithm for predicting unplanned readmissions. Bringing together tools from causal inference and statistical process control, we consider six monitoring procedures (three candidate monitoring criteria and two data sources) and investigate their operating characteristics in simulation studies. Results from this case study emphasize the seemingly simple (and obvious) fact that not all monitoring systems are created equal, which has real-world impacts on the design and documentation of ML monitoring systems.

cs.LG

A Brief Tutorial on Sample Size Calculations for Fairness Audits

In fairness audits, a standard objective is to detect whether a given algorithm performs substantially differently between subgroups. Properly powering the statistical analysis of such audits is crucial for obtaining informative fairness assessments, as it ensures a high probability of detecting unfairness when it exists. However, limited guidance is available on the amount of data necessary for a fairness audit, lacking directly applicable results concerning commonly used fairness metrics. Additionally, the consideration of unequal subgroup sample sizes is also missing. In this tutorial, we address these issues by providing guidance on how to determine the required subgroup sample sizes to maximize the statistical power of hypothesis tests for detecting unfairness. Our findings are applicable to audits of binary classification models and multiple fairness metrics derived as summaries of the confusion matrix. Furthermore, we discuss other aspects of audit study designs that can increase the reliability of audit results.

stat.AP

Quantile regression with varying coefficients

Quantile regression provides a framework for modeling statistical quantities of interest other than the conditional mean. The regression methodology is well developed for linear models, but less so for nonparametric models. We consider conditional quantiles with varying coefficients and propose a methodology for their estimation and assessment using polynomial splines. The proposed estimators are easy to compute via standard quantile regression algorithms and a stepwise knot selection algorithm. The proposed Rao-score-type test that assesses the model against a linear model is also easy to implement. We provide asymptotic results on the convergence of the estimators and the null distribution of the test statistic. Empirical results are also provided, including an application of the methodology to forced expiratory volume (FEV) data.

math.ST