SearcharxivSearch

arXiv subjects

Sarah Pungitore

Publications and source records attributed to Sarah Pungitore.

4 recordsLinked to original sources

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks

Although computational phenotyping is a central informatics activity with resulting cohorts supporting a wide variety of applications, it is time-intensive because of manual data review. We previously assessed the ability of LLMs to perform computational phenotyping tasks using computable phenotypes for ARF respiratory support therapies. They successfully performed concept classification and classification of single-therapy phenotypes but underperformed on multi-therapy phenotypes. To better understand issues with these complex tasks, we expanded PHEONA, a generalizable framework for evaluation of LLMs, to include methods specifically for evaluating faulty reasoning. We assessed the responses of two lightweight non-reasoning LLMs (Mistral Small 24 billion and Phi-4 14 billion) and one lightweight reasoning LLM (Qwen-distilled DeepSeek-r1 32 billion) both with and without prompt modifications to identify explanation correctness errors and unfaithfulness errors during phenotyping. For experiments without prompt modifications, both errors were present in responses from all models. For experiments with prompt modifications, we measured the mean absolute change in accuracy relative to the unbiased prompt across biasing conditions. Adding specific few-shot examples aligned with an incorrect phenotype reduced accuracy by at least 5% and up to 10% depending on the model and CoT type. Since reasoning errors were ubiquitous across models, our enhancement of PHEONA to include a component for assessing faulty reasoning provides a practical framework for evaluating LLM reasoning and empirical evidence that reasoning errors occur during complex computational phenotyping.

q-bio.QM

SHREC: A Framework for Advancing Next-Generation Computational Phenotyping with Large Language Models

Computational phenotyping is a central informatics activity with resulting cohorts supporting a wide variety of applications. However, it is time-intensive because of manual data review and limited automation. Since LLMs have demonstrated promising capabilities for text classification, comprehension, and generation, we posit they will perform well at repetitive manual review tasks traditionally performed by human experts. To support next-generation computational phenotyping, we developed SHREC, a framework for integrating LLMs into end-to-end phenotyping pipelines. We applied and tested three lightweight LLMs (Gemma2 27 billion, Mistral Small 24 billion, and Phi-4 14 billion) to classify concepts and phenotype patients using phenotypes for ARF respiratory support therapies. All models performed well on concept classification, with the best (Mistral) achieving an AUROC of 0.896. For phenotyping, models demonstrated near-perfect specificity for all phenotypes with the top-performing model (Mistral) achieving an average AUROC of 0.853 for single-therapy phenotypes. In conclusion, lightweight LLMs can assist researchers with resource-intensive phenotyping tasks. Several advantages of LLMs included their ability to adapt to new tasks with prompt engineering alone and their ability to incorporate raw EHR data. Future steps include determining optimal strategies for integrating biomedical data and understanding reasoning errors

q-bio.QM

PHEONA: An Evaluation Framework for Large Language Model-based Approaches to Computational Phenotyping

Computational phenotyping is essential for biomedical research but often requires significant time and resources, especially since traditional methods typically involve extensive manual data review. While machine learning and natural language processing advancements have helped, further improvements are needed. Few studies have explored using Large Language Models (LLMs) for these tasks despite known advantages of LLMs for text-based tasks. To facilitate further research in this area, we developed an evaluation framework, Evaluation of PHEnotyping for Observational Health Data (PHEONA), that outlines context-specific considerations. We applied and demonstrated PHEONA on concept classification, a specific task within a broader phenotyping process for Acute Respiratory Failure (ARF) respiratory support therapies. From the sample concepts tested, we achieved high classification accuracy, suggesting the potential for LLM-based methods to improve computational phenotyping processes.

cs.CL

Computable Phenotypes for Post-acute sequelae of SARS-CoV-2: A National COVID Cohort Collaborative Analysis

Post-acute sequelae of SARS-CoV-2 (PASC) is an increasingly recognized yet incompletely understood public health concern. Several studies have examined various ways to phenotype PASC to better characterize this heterogeneous condition. However, many gaps in PASC phenotyping research exist, including a lack of the following: 1) standardized definitions for PASC based on symptomatology; 2) generalizable and reproducible phenotyping heuristics and meta-heuristics; and 3) phenotypes based on both COVID-19 severity and symptom duration. In this study, we defined computable phenotypes (or heuristics) and meta-heuristics for PASC phenotypes based on COVID-19 severity and symptom duration. We also developed a symptom profile for PASC based on a common data standard. We identified four phenotypes based on COVID-19 severity (mild vs. moderate/severe) and duration of PASC symptoms (subacute vs. chronic). The symptoms groups with the highest frequency among phenotypes were cardiovascular and neuropsychiatric with each phenotype characterized by a different set of symptoms.

q-bio.QM