TY - RPRT TI - K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations AU - Laura M. Vowels AU - Matthew J. Vowels AU - Shivali Sharma AU - Apoorv Jha AU - Rehnuma Choudhury AU - Wasseem El Sarraj AU - Rachel Francois-Walcott AU - Aruba Hussain AU - Sarah Ingram AU - Angela Loulopoulou AU - Adva Segal AU - Elena Volkova PY - 2026 UR - https://arxiv.org/abs/2609.15855 ID - 2609.15855 ER -