Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
A prior audit-and-repair exchange made LLM checkers less likely to raise false alarms on the same verification task.
The paper reports that false alarms fell in all 15 tested model-and-wording combinations, by 2.8 to 11.5 percentage points versus a length-matched control. The authors say the shift came from a looser decision threshold, not better discrimination. An audit episode that had found an error pushed leniency further, contrary to the negativity-asymmetry expectation cited in the paper. A hand audit of 50 false alarms found 82% were simply wrong, so the authors argue this operating-point shift was not necessarily harmful. HF Daily Papers' note
The paper reports that false alarms fell in all 15 tested model-and-wording combinations, by 2.8 to 11.5 percentage points versus a length-matched control. The authors say the shift came from a looser decision threshold, not better discrimination. An audit episode that had found an error pushed leniency further, contrary to the negativity-asymmetry expectation cited in the paper. A hand audit of 50 false alarms found 82% were simply wrong, so the authors argue this operating-point shift was not necessarily harmful. HF Daily Papers' note
score 4