HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication
The paper’s core warning is that high overall scores can hide under-triage failures in multilingual women’s-health queries.
HerHealthEval tests whether models understand the same clinical concern across English, French, and Modern Standard Arabic, and across clinical, lay, hedged, emotional, and under-specified phrasing. The under-specified version is designed to see whether a model asks for clarification instead of forcing an answer. In the reported tests, one multilingual adaptation model showed severe under-triage in French and Arabic under language-asymmetric risk supervision. A controlled re-adaptation using language-invariant risk labels reduced those failures, but did not erase the need for explicit register and uncertainty testing. ArXiv · AI/CL/LG's note
HerHealthEval tests whether models understand the same clinical concern across English, French, and Modern Standard Arabic, and across clinical, lay, hedged, emotional, and under-specified phrasing. The under-specified version is designed to see whether a model asks for clarification instead of forcing an answer. In the reported tests, one multilingual adaptation model showed severe under-triage in French and Arabic under language-asymmetric risk supervision. A controlled re-adaptation using language-invariant risk labels reduced those failures, but did not erase the need for explicit register and uncertainty testing. ArXiv · AI/CL/LG's note
score 4