SearcharxivSearch

arXiv subjects

James A. Diao

Publications and source records attributed to James A. Diao.

9 recordsLinked to original sources

Kidney function and kidney failure prediction in a large multiethnic population

Background: Patients with chronic kidney disease (CKD) experience worsening kidney function and develop subsequent kidney failure at different rates. Accurate estimates of a patient's current and future kidney function are needed for optimal clinical decision-making. Methods: To compare current and previously recommended equations for estimating kidney function and risk of kidney failure across a large multiethnic population, we conducted a retrospective multicenter cohort study of primary care, acute care, and hospital settings. Our study population comprised 1,909,042 adults with at least one recorded serum creatinine measurement during 2012-2014, with follow-up until January 2025. Our primary outcomes were the area under the receiver operating characteristic curve and prevalence of CKD by stage across regions of origin. Findings: GFR estimates from the two race-stratified equations (2006 MDRD and 2009 CKD-EPI) were similarly calibrated. GFR estimates from the two race-neutral equations diverged, with 2021 EKFC producing lower GFR estimates and 2021 CKD-EPI producing higher GFR estimates compared to prior equations. The oldest equation, 2006 MDRD, was the most discriminative of kidney failure within 5 years while the newer European equation, 2021 EKFC, was the least discriminative, with AUROCs of 0.862 (95% CI, 0.855-0.869) and 0.846 (95% CI, 0.838-0.853), respectively. The age-adjusted prevalence of CKD varied by the choice of eGFR equation (ranging from 8.5% for 2021 CKD-EPI to 10.8% for EKFC) and across regions of birth (ranging from 8.9% for East Sub-Saharan Africa to 15.3% for South Asia). Interpretation: Using 10 years of follow-up for nearly 2 million individuals, our study demonstrates how newly developed US and European equations diverge from previously recommended equations with broad clinical and epidemiological implications.

q-bio.QM

Laboratory Trajectories Improve Kidney Failure Risk Estimation

Accurate kidney failure risk assessment is critical to timely intervention in chronic kidney disease (CKD). Existing equations (e.g. Kidney Failure Risk Equation; KFRE) rely on single laboratory measurements to estimate short- and long-term kidney failure risk, leaving longitudinal laboratory patterns unused. Here we introduce Clalit Longitudinal Assessment of Risk of Kidney Failure (CLARK), an interpretable longitudinal extension of latest-value methods which incorporates routinely collected repeat laboratory measures. We develop CLARK using data from 5.4 million individuals, identifying 270,009 patients with CKD to create one of the largest longitudinal CKD cohorts to date, with 12,087 kidney replacement therapy initiation events and a median follow-up of 10.4 years. Across laboratory configurations and prediction horizons, CLARK demonstrated improved discrimination over static models (e.g., 2-year average precision 0.541 vs 0.516 in the eGFR-only setting). At intervention thresholds, trajectory-based models improved identification of high-risk patients, especially for longer-term prediction, suggesting that interpretable longitudinal laboratory features may enhance kidney failure risk assessment through improved identification of patients most likely to benefit from timely intervention.

q-bio.QM

Teaching large language models to reason like expert diagnosticians

Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferences (CPCs), published continuously since 1923, feature expert physicians who demonstrate diagnostic reasoning to peers, and have been used for decades to evaluate AI. However, prior AI evaluations have largely focused on final diagnostic accuracy rather than nuanced clinical reasoning. Here, we introduce Dr. CaBot, an agentic AI system that emulates an expert diagnostician by generating written and narrated slide-based presentations from an initial case description alone. CaBot recently generated the first AI diagnosis published in the 100+ year history of the NEJM CPCs. In blinded evaluations, physicians misclassified the source of the differential (CaBot vs. physician-written) in 46/62 (74%) of trials and rated them favorably across quality dimensions. When tasked with solving cases for 72 patients with undiagnosed disease from the NIH Undiagnosed Diseases Network, CaBot identified the working diagnosis in 50/72 (69%) of cases from referral notes alone. To promote transparency and research, we also developed CPC-Bench, a physician-validated benchmark based on 7,102 CPCs and 47,648 questions across 10 tasks. We show that CaBot outperforms frontier models on CPC-Bench, and release both CaBot and CPC-Bench publicly to foster progress in clinical AI.

cs.AI

Learning Normal Representations for Blood Biomarkers

Blood-based biomarkers underpin clinical diagnosis and management, yet their interpretation relies largely on fixed population reference intervals that ignore stable, intra-patient variability. As such, population-based interpretation can mask meaningful deviation from an individual's baseline, risking delayed disease detection. To remedy this, there have been increasing efforts to personalize blood biomarker interpretation using individual testing histories. However, these methods may overfit to sparse data, inflating false-positive rates and unnecessary follow-up, and can also unwittingly include unrecognized or subclinical disease. Here, we leverage nearly 2 billion longitudinal laboratory measurements from over 1.6 million individuals across North America, the Middle East, and East Asia, to show that while laboratory values are highly individual, purely personalized intervals routinely overfit, classifying up to 68% of measurements as abnormal, without corresponding associations with adverse clinical outcomes. We then introduce NORMA, a conditional transformer-based framework that generates reference intervals by conditioning on both a patient's history and population-level data about "normal" variation. NORMA-derived intervals achieve higher precision for predicting outcomes, including mortality, acute kidney injury, and chronic disease. These findings caution against over-personalization in laboratory medicine and demonstrate that anchoring individual trajectories to population-level priors outperforms either approach alone. To promote transparency, we publicly release the model, code, and an interactive user interface for accessible, individualized laboratory interpretation.

cs.LG

Statin Recommendations among US Adults with the 2026 Dyslipidemia Guidelines

Importance: The 2026 multisociety dyslipidemia guideline recommended the PREVENT equations in place of the PCE equations, introduced 30-year risk assessment as a new treatment pathway, and lowered risk-based treatment thresholds. The net population impact of these concurrent changes on statin recommendations is unknown. Objective: To estimate changes in statin recommendations under 2026 PREVENT-based dyslipidemia guidelines compared with 2018 PCE-based guidelines. Design and Participants: Cross-sectional analysis of pooled data from NHANES, spanning 2011-2023 and comprising 24,199 participants aged 30-79 years. Main Outcomes and Measures: Number and proportion of US adults receiving or recommended for statin therapy. Results: At the class 1 threshold, the number of US adults receiving or recommended for statin therapy decreased by an estimated 3.0 million (95% CI, 2.3 million to 3.6 million), with larger reductions among Black adults (-4.2 percentage points [pp]), men (-4.0pp), and adults aged 50-69 years (-5.6pp). At the class 2 threshold--which additionally recommends statins for adults aged 30-59 years based on 30-year risk--the number of adults recommended increased by an estimated 20.8 million (95% CI, 19.6 million to 22.0 million), or +11.6pp. The increase was largest among adults aged 50-59 years (+19.7pp) and 40-49 years (+14.8pp). Conclusions: The net population impact of the 2026 dyslipidemia guidelines depends critically on which recommendation class is applied. At the class 1 threshold, statin recommendations decreased modestly; at the class 2 threshold, inclusion of 30-year risk assessment substantially expanded recommendations, particularly among younger adults. These divergent effects underscore the importance of the 30-year risk criterion as a major driver of new eligibility and the need for outcomes and equity monitoring during guideline implementation.

q-bio.OT

Disentangling Proxies of Demographic Adjustments in Clinical Equations

The use of coarse demographic adjustments in clinical equations has been increasingly scrutinized. In particular, adjustments for race have sparked significant debate with several medical professional societies recommending race-neutral equations in recent years. However, current approaches to remove race from clinical equations do not address the underlying causes of observed differences. Here, we present ARC (Approach for identifying pRoxies of demographic Correction), a framework to identify explanatory factors of group-level differences, which may inform the development of more accurate and precise clinical equations. We apply ARC to spirometry tests across two observational cohorts, CDC NHANES and UK Biobank, comprising 159,893 participants. Cross-sectional sociodemographic or exposure measures did not explain differences in reference lung function across race groups beyond those already explained by age, sex, and height. By contrast, sitting height accounted for up to 26% of the remaining differences in lung volumes between healthy Black and White adults. We then demonstrate how pulmonary function test (PFT) reference equations can incorporate these factors in a new set of equations called $ARC_{PFT}$, surpassing the predictive performance of the race-neutral GLI-Global equation recommended by major pulmonary societies. When compared to GLI-Global, inclusion of sitting height and waist circumference in $ARC_{PFT}$ decreased mean absolute error by 13% among Black participants in the UK Biobank and by 24% in NHANES. $ARC_{PFT}$ also had reduced vulnerability to domain shift compared to race-based methods, with mean absolute error 19.3% and 35.6% lower than race-stratified models in out-of-sample Asian and Hispanic populations, respectively. This approach provides a path for understanding the proxies of imprecise demographic adjustments and developing personalized clinical equations.

q-bio.QM

Multimodal Foundation Models Exploit Text to Make Medical Image Predictions

Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text, remain poorly understood. Here, using a diverse collection of 1014 multimodal medical cases, we evaluate the unimodal and multimodal image interpretation abilities of proprietary (GPT-4, Gemini Pro 1.0) and open-source (Llama-3.2-90B, LLaVA-Med-v1.5) multimodal foundational models with and without the use of text descriptions. Across all models, image predictions were largely driven by exploiting text, with accuracy increasing monotonically with the amount of informative text. By contrast, human performance on medical image interpretation did not improve with informative text. Exploitation of text is a double-edged sword; we show that even mild suggestions of an incorrect diagnosis in text diminishes image-based classification, reducing performance dramatically in cases the model could previously answer with images alone. Finally, we conducted a physician evaluation of model performance on long-form medical cases, finding that the provision of images either reduced or had no effect on model performance when text is already highly informative. Our results suggest that multimodal AI models may be useful in medical diagnostic reasoning but that their accuracy is largely driven, for better and worse, by their exploitation of text.

cs.CV

Augmenting medical image classifiers with synthetic data from latent diffusion models

While hundreds of artificial intelligence (AI) algorithms are now approved or cleared by the US Food and Drugs Administration (FDA), many studies have shown inconsistent generalization or latent bias, particularly for underrepresented populations. Some have proposed that generative AI could reduce the need for real data, but its utility in model development remains unclear. Skin disease serves as a useful case study in synthetic image generation due to the diversity of disease appearance, particularly across the protected attribute of skin tone. Here we show that latent diffusion models can scalably generate images of skin disease and that augmenting model training with these data improves performance in data-limited settings. These performance gains saturate at synthetic-to-real image ratios above 10:1 and are substantially smaller than the gains obtained from adding real images. As part of our analysis, we generate and analyze a new dataset of 458,920 synthetic images produced using several generation strategies. Our results suggest that synthetic data could serve as a force-multiplier for model development, but the collection of diverse real-world data remains the most important step to improve medical AI algorithms.

cs.CV

Improving dermatology classifiers across populations using images generated by large diffusion models

Dermatological classification algorithms developed without sufficiently diverse training data may generalize poorly across populations. While intentional data collection and annotation offer the best means for improving representation, new computational approaches for generating training data may also aid in mitigating the effects of sampling bias. In this paper, we show that DALL$\cdot$E 2, a large-scale text-to-image diffusion model, can produce photorealistic images of skin disease across skin types. Using the Fitzpatrick 17k dataset as a benchmark, we demonstrate that augmenting training data with DALL$\cdot$E 2-generated synthetic images improves classification of skin disease overall and especially for underrepresented groups.

eess.IV