arXiv · 2507.21176
Toward Revealing Nuanced Biases in Medical LLMs
Abstract
Large language models (LLMs) used in medical applications are known to be prone to exhibiting biased and unfair patterns. Prior to deploying these in clinical decision-making, it is crucial to identify such bias patterns to enable effective mitigation and minimize negative impacts. In this study, we present a novel framework combining knowledge graphs (KGs) with auxiliary (agentic) LLMs to systematically reveal complex bias patterns in medical LLMs. The proposed approach integrates adversarial perturbation (red teaming) techniques to identify subtle bias patterns and adopts a customized multi-hop characterization of KGs to enhance the systematic evaluation of target LLMs. It aims not only to generate more effective red-teaming questions for bias evaluation but also to utilize those questions more effectively in revealing complex biases. Through a series of comprehensive experiments on three datasets, six LLMs, and five bias types, we demonstrate that our proposed framework exhibits a noticeably greater ability and scalability in revealing complex biased patterns of medical LLMs compared to other common approaches.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Farzana Islam Adiba, Rahmatollah Beheshti. 2025-07-26. Toward Revealing Nuanced Biases in Medical LLMs. https://arxiv.org/abs/2507.21176
Cite the original work for its findings. Save a collection to share your selection of sources.