SearcharxivSearch

arXiv subjects

Riya Nagar

Publications and source records attributed to Riya Nagar.

3 recordsLinked to original sources

Patterns in Individual Blood Count Trajectories in the UK Biobank Characterise Disease-Specific Signatures and Anticipate Pan-Cancer Risk

We investigate the longitudinal behaviour of blood markers from common haematological tests as a marker of disease and as a function of disease progression in a variety of conditions including cancer, cardiovascular disease, and infections. We study confounding and non-confounding factors to allow for the earlier detection of disease and conditions based on their longitudinal signatures from biomarker patterns commonly measured in popular and scalable common blood tests across routine clinical tests, in particular the Complete Blood Count (CBC or FBC). Our analysis with normalised temporal profiles and machine learning techniques even before any symptoms appear demonstrates that analyte-group patterns found in blood testing are disease sensitive and disease specific. We demonstrate that CBC markers contribute to the majority of the predictive signal, while biochemistry and other blood panels provide only a modest additional gain mostly associated to very the individual disease for which the test was designed (e.g. CRP, liver enzymes, blood sugar). Our results demonstrate how regular monitoring, computational intelligence, and machine learning applied to longitudinal CBC data can converge to uncover disease patterns, advancing the potential for precision healthcare and predictive medicine on a mass scale leveraging an existing and pervasive blood test.

q-bio.QM

XGBoost-Powered Digital Twins Leverage Routine Blood Tests for Early Detection of Cancer and Cardiovascular Disease

Early detection of cancer and cardiovascular diseases is fundamental to improving patient outcomes and reducing healthcare expenditure. Current cancer screening programs are targeted towards specific cancers and are often inaccessible to large parts of the population, particularly in remote regions. This project aimed to develop digital blood twins: machine learning models that leverage routinely collected blood test data, demographics, comorbidities, and prescribed medications, for scalable and cost-effective disease screening. Digital blood twins were constructed using the UK Biobank dataset (n = 373,269). Using age, sex, comorbidities, medication profiles, and blood test z-scores, three iterations of XGBoost classifiers were trained for broad cancer, colorectal cancer, and cardiovascular disease prediction. Model interpretability was achieved through SHAP and dimensionality reduction analyses (UMAP, t-SNE). Broad-category cancer models achieved ROC-AUC = 0.607-0.706. Colorectal cancer prediction demonstrated excellent discrimination (ROC-AUC = 0.816-0.993), and cardiovascular models showed clinical utility, notably for hypertension (ROC-AUC = 0.813, F1 = 0.861). SHAP revealed consistent importance of age, sex, basophil count, and cystatin C. Immune digital blood twins as an agnostic tool demonstrate proof-of-concept feasibility for accessible, low-cost, and scalable screening of cancer and cardiovascular diseases, supporting future integration into predictive and preventive healthcare.

q-bio.OT

Exhaustive Investigation of CBC-Derived Biomarker Ratios for Clinical Outcome Prediction: The RDW-to-MCHC Ratio as a Novel Mortality Predictor in Critical Care

Ratios of common biomarkers and blood analytes are well established for early detection and predictive purposes. Early risk stratification in critical care is often limited by the delayed availability of complex severity scores. Complete blood count (CBC) parameters, available within hours of admission, may enable rapid prognostication. We conducted an exhaustive and systematic evaluation of CBC-derived ratios for mortality prediction to identify robust, accessible, and generalizable biomarkers. We generated all feasible two-parameter CBC ratios with unit checks and plausibility filters on more than 90,000 ICU admissions (MIMIC-IV). Discrimination was assessed via cross-validated and external AUC, calibration via isotonic regression, and clinical utility with decision-curve analysis. Retrospective validation was performed on eICU-CRD (n = 156530) participants. The ratio of Red Cell Distribution Width (RDW) to Mean Corpuscular Hemoglobin Concentration (MCHC), denoted RDW:MCHC, emerged as the top biomarker (AUC = 0.699 discovery; 0.662 validation), outperforming RDW and NLR. It achieved near-universal availability (99.9\% vs.\ 35.0\% for NLR), excellent calibration (Hosmer--Lemeshow $p = 1.0$; $\mathrm{ECE} < 0.001$), and preserved performance across diagnostic groups, with only modest attenuation in respiratory cases. Expressed as a logistic odds ratio, each one standard deviation increase in RDW:MCHC nearly quadrupled 30-day mortality odds (OR = 3.81, 95\% CI [3.70, 3.95]). Decision-curve analysis showed positive net benefit at high-risk triage thresholds. A simple, widely available CBC-derived feature (RDW:MCHC) provides consistent, externally validated signal for early mortality risk. While not a substitute for multivariable scores, it offers a pragmatic adjunct for rapid triage when full scoring is impractical.

q-bio.QM