K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
K-Bench tests LLMs on protected, clinician-calibrated multi-turn mental health scenarios where safety failures carry high stakes.
The benchmark covers 125 model configurations from 33 base models across suicide, self-harm, domestic violence, substance misuse, and no-risk vignettes. Its synthetic patient conversations substantially overlapped with real human-AI conversations, according to the paper. A frozen GPT-4o judge matched clinician consensus 94.2% of the time across eligible comparisons. Stronger models scored above 95 on combined-risk measures, while lower-performing configurations varied sharply when asked to explore risk. ArXiv · AI/CL/LG's note
The benchmark covers 125 model configurations from 33 base models across suicide, self-harm, domestic violence, substance misuse, and no-risk vignettes. Its synthetic patient conversations substantially overlapped with real human-AI conversations, according to the paper. A frozen GPT-4o judge matched clinician consensus 94.2% of the time across eligible comparisons. Stronger models scored above 95 on combined-risk measures, while lower-performing configurations varied sharply when asked to explore risk. ArXiv · AI/CL/LG's note
score 6