SearcharxivSearch

arXiv subjects

Ting-Yu Chen

Publications and source records attributed to Ting-Yu Chen.

3 recordsLinked to original sources

Latent Confidence Alignment for LLM Self-Assessment

Confidence calibration in large language models (LLMs) is commonly evaluated by comparing predicted confidence with observed accuracy. However, such approaches do not model item difficulty, making it difficult to interpret discrepancies and to determine whether model confidence reflects genuine self-assessment or is merely a byproduct of the response generation process. To address this, we adopt a Rasch model-based latent ability framework and a metacognitive perspective, and propose Latent Confidence Alignment Error (LCAE) to measure the consistency between model self-assessment and the latent error probability implied by model ability and item difficulty. We further incorporate item difficulty as an external signal with a reasoning mechanism. Experiments on a medical-domain dataset with 20 models show that the proposed approach improves self-assessment quality without affecting model ability, and reveals an association between reliability and inference cost.

cs.CY

Towards Interpretable Renal Health Decline Forecasting via Multi-LMM Collaborative Reasoning Framework

Accurate and interpretable prediction of estimated glomerular filtration rate (eGFR) is essential for managing chronic kidney disease (CKD) and supporting clinical decisions. Recent advances in Large Multimodal Models (LMMs) have shown strong potential in clinical prediction tasks due to their ability to process visual and textual information. However, challenges related to deployment cost, data privacy, and model reliability hinder their adoption. In this study, we propose a collaborative framework that enhances the performance of open-source LMMs for eGFR forecasting while generating clinically meaningful explanations. The framework incorporates visual knowledge transfer, abductive reasoning, and a short-term memory mechanism to enhance prediction accuracy and interpretability. Experimental results show that the proposed framework achieves predictive performance and interpretability comparable to proprietary models. It also provides plausible clinical reasoning processes behind each prediction. Our method sheds new light on building AI systems for healthcare that combine predictive accuracy with clinically grounded interpretability.

cs.LG

Multiplicity for a Group Sequential Trial with Biomarker Subpopulations

Biomarker subpopulations have become increasingly important for drug development in targeted therapies. The use of biomarkers has the potential to facilitate more effective outcomes by guiding patient selection appropriately, thus enhancing the benefit-risk profile and improving trial power. Studying a broad population simultaneously with a more targeted one allows the trial to determine the population for which a treatment is effective and allows a goal of making approved regulatory labeling as inclusive as is appropriate. We examine new methods accounting for the complete correlation structure in group sequential designs with hypotheses in nested subgroups. The designs provide full control of family-wise Type I error rate. This extension of previous methods accounting for either group sequential design or correlation between subgroups improves efficiency (power or sample size) over a typical Bonferroni approach for testing nested populations.

stat.ME