Searcharxiv⌕ Search

arXiv subjects

Liat Antwarg Friedman

Publications and source records attributed to Liat Antwarg Friedman.

3 recordsLinked to original sources

Agentic discovery of blood biomarker from distilled private health records

Routine complete blood counts (CBCs) could yield new biomarkers, but the private records needed to evaluate candidates cannot be shared with frontier language model agents that excel at discovery. We distilled the evidence held in the Clalit Health Services panel of over 5.4 million patients into a released scoring tool: for each of 13 immune-mediated diseases, a graph attention network was trained inside the data boundary to predict the case-control AUC of candidate CBC expressions, and only the trained weights were released. The tool grounds an agent's propose-score-refine loop in real-world data without exposing any patient data. In external validation, agent-discovered expressions improved on their literature-seeded starting points by a median of 4.18 AUC percentage points, and across three independent cohorts, reranking the candidates of three frontier research tools improved on their first choices in most comparisons, with gains that varied by cohort. The released scorer supports privacy-preserving biomarker hypothesis generation.

cs.AI↗

EveryQuery: A Promptable Foundation Model for Clinical Prediction Tasks over Electronic Health Records

Autoregressive foundation models for electronic health records can make predictions for diverse clinical tasks without finetuning. These methods work by generating possible "synthetic futures" for patients, and then inferring downstream predictions over these simulated trajectories. Despite their strong capabilities, these methods have several flaws. (1) Inference is extremely computationally expensive, as each prediction requires simulating many trajectories, each with many observations. (2) Their performance is noisy due to the variance inherent to simulation, which in particular reduces efficacy on rare events, which are often of high importance clinically. (3) They are not promptable, as their only input is the patient's medical history, requiring indirect aggregation at inference time. We introduce EveryQuery, a promptable foundation model for structured EHR data. Building on the insight that clinical prediction tasks can be written in a structured query language, EveryQuery takes a query as a prompt alongside the patient's history and estimates the outcome directly. This allows EveryQuery to make predictions across diverse clinical tasks, such as classification and survival modeling, without simulation or finetuning. Across three EHR datasets, EveryQuery outperforms a competitive autoregressive baseline on clinically relevant classification and time-to-event tasks, with mean AUROC gains of 14.3% and 12.7% on MIMIC-IV and a large academic medical center and comparable performance on the smaller NWICU dataset. Its advantage is largest for rare outcomes, and it requires 580x less inference compute.

cs.AI↗

Laboratory Trajectories Improve Kidney Failure Risk Estimation

Accurate kidney failure risk assessment is critical to timely intervention in chronic kidney disease (CKD). Existing equations (e.g. Kidney Failure Risk Equation; KFRE) rely on single laboratory measurements to estimate short- and long-term kidney failure risk, leaving longitudinal laboratory patterns unused. Here we introduce Clalit Longitudinal Assessment of Risk of Kidney Failure (CLARK), an interpretable longitudinal extension of latest-value methods which incorporates routinely collected repeat laboratory measures. We develop CLARK using data from 5.4 million individuals, identifying 270,009 patients with CKD to create one of the largest longitudinal CKD cohorts to date, with 12,087 kidney replacement therapy initiation events and a median follow-up of 10.4 years. Across laboratory configurations and prediction horizons, CLARK demonstrated improved discrimination over static models (e.g., 2-year average precision 0.541 vs 0.516 in the eGFR-only setting). At intervention thresholds, trajectory-based models improved identification of high-risk patients, especially for longer-term prediction, suggesting that interpretable longitudinal laboratory features may enhance kidney failure risk assessment through improved identification of patients most likely to benefit from timely intervention.

q-bio.QM↗