Megadose Built for builders and researchers.

InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

· ArXiv · AI/CL/LG ·
The benchmark tests whether legal LLMs notice missing facts before giving advice.

InsufficiencyBench covers 202 items across six legal domains and 24 U.S. jurisdictions, with deficient variants annotated by practising attorneys. The paper says ten frontier models performed poorly at identifying legally material missing elements, with none above F2 = 0.46 and median recall at 0.44. The authors found models tended either to hedge broadly or answer under unstated assumptions. ArXiv · AI/CL/LG's note

score 5

Categories: Research