Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
Lenient keyword benchmarks credited tool use that one small model was not actually performing.
The paper reports a false positive across two Spanish security models with shared architecture details but different training histories. A 1.1B model with tool-SFT matched a 661.6M model on lenient metrics, yet failed strict reproduction checks and assigned near-zero probability to the tool-call trigger token. A targeted SFT run repaired valid tool-call emission, raising it to 0.959 on the corpus rows and outperforming the 600M model on unseen prompts. The author says these cheap diagnostics should gate tool-use claims for small models. ArXiv · AI/CL/LG's note
The paper reports a false positive across two Spanish security models with shared architecture details but different training histories. A 1.1B model with tool-SFT matched a 661.6M model on lenient metrics, yet failed strict reproduction checks and assigned near-zero probability to the tool-call trigger token. A targeted SFT run repaired valid tool-call emission, raising it to 0.959 on the corpus rows and outperforming the 600M model on unseen prompts. The author says these cheap diagnostics should gate tool-use claims for small models. ArXiv · AI/CL/LG's note
score 4