ContractScrub: A benchmark for final review of legal contracts
Frontier models still miss too much in lawyer-built contract scrubbing tests.
ContractScrub evaluates final contract review tasks such as defined-term misuse, bad references, and inconsistent language. The benchmark uses contracts hand-crafted by experienced lawyers across multiple error categories. The authors say no formal LLM evaluation for contract scrubbing had previously been conducted. In their tests, only one frontier model reached 0.75 macro average recall, despite stronger results on related general benchmarks. ArXiv · AI/CL/LG's note
ContractScrub evaluates final contract review tasks such as defined-term misuse, bad references, and inconsistent language. The benchmark uses contracts hand-crafted by experienced lawyers across multiple error categories. The authors say no formal LLM evaluation for contract scrubbing had previously been conducted. In their tests, only one frontier model reached 0.75 macro average recall, despite stronger results on related general benchmarks. ArXiv · AI/CL/LG's note
score 5