Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case
The paper tests whether LLM investigators can tell when the evidence is still too thin to close a case.
Its benchmark, Nautil, contains 731 audited investigation cases across transport, safety, vehicle-defect and server-incident reports. The authors say untrained and frontier models often overstate what the record supports, even when they identify the right cause. Fine-tuning on teacher trajectories reduced overstatement in a 9B model from 97% to 35%, while raising correct, non-overstated conclusions from 3% to 43%. A closure-only reinforcement step improved balanced accuracy to 83.3, but weakened evidence dependence. ArXiv · AI/CL/LG's note
Its benchmark, Nautil, contains 731 audited investigation cases across transport, safety, vehicle-defect and server-incident reports. The authors say untrained and frontier models often overstate what the record supports, even when they identify the right cause. Fine-tuning on teacher trajectories reduced overstatement in a 9B model from 97% to 35%, while raising correct, non-overstated conclusions from 3% to 43%. A closure-only reinforcement step improved balanced accuracy to 83.3, but weakened evidence dependence. ArXiv · AI/CL/LG's note
score 4