Megadose Built for builders and researchers.

Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

· HF Daily Papers ·
The paper says lenient keyword benchmarks credited a small model with tool use it was not actually performing.

Santillana tests two Spanish security models with shared architecture details and finds near-identical scores on loose tool-use metrics, despite one model failing stricter checks. The 1.1B model’s failure is traced to a very low prior on the `<|tool_call|>` token after web-heavy training. A targeted SFT run repairs the format in about 3.3 GPU-hours, lifting valid emission from 0.100 to 0.959 on the corpus rows. The paper still flags over-triggering in both models, with negative prompts often getting tool calls anyway. HF Daily Papers' note

score 4

Categories: Research