Megadose AI progress, ranked and analyzed.

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

· ArXiv · AI/CL/LG ·
K-Bench tests LLMs on protected, clinician-calibrated multi-turn mental health scenarios where safety failures carry high stakes.

The benchmark covers 125 model configurations from 33 base models across suicide, self-harm, domestic violence, substance misuse, and no-risk vignettes. Its synthetic patient conversations substantially overlapped with real human-AI conversations, according to the paper. A frozen GPT-4o judge matched clinician consensus 94.2% of the time across eligible comparisons. Stronger models scored above 95 on combined-risk measures, while lower-performing configurations varied sharply when asked to explore risk. ArXiv · AI/CL/LG's note

score 6

Categories: Research