GANDR: Claim Auditing for Verifiable Legal Answer Generation
GANDR’s main claim is that legal AI answers should be judged claim by claim, not just by final correctness.
The paper describes a two-agent setup: a Drafter writes structured legal answers, while a Critic checks each claim against the cited retrieved passage. On a 185-item legal benchmark, GANDR reached 70.8% strict accuracy and beat the strongest baseline by 11.3 points. Removing its protocol-anchored commit rule cut strict accuracy by 22.7 points. The authors say the audit flags under-supported claims well as a binary detector, though its finer verdict labels agree only weakly with law-trained annotators. ArXiv · AI/CL/LG's note
The paper describes a two-agent setup: a Drafter writes structured legal answers, while a Critic checks each claim against the cited retrieved passage. On a 185-item legal benchmark, GANDR reached 70.8% strict accuracy and beat the strongest baseline by 11.3 points. Removing its protocol-anchored commit rule cut strict accuracy by 22.7 points. The authors say the audit flags under-supported claims well as a binary detector, though its finer verdict labels agree only weakly with law-trained annotators. ArXiv · AI/CL/LG's note
score 5