TY - RPRT TI - Medical Large Language Model Benchmarks Should Prioritize Construct Validity AU - Ahmed Alaa AU - Thomas Hartvigsen AU - Niloufar Golchini AU - Shiladitya Dutta AU - Frances Dean AU - Inioluwa Deborah Raji AU - Travis Zack PY - 2025 UR - https://arxiv.org/abs/2503.10694 ID - 2503.10694 ER -