Megadose Built for builders and researchers.

When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs

· HF Daily Papers ·
The paper argues that steering a tool-using model’s internals can improve decisions while still damaging many decisions that were already correct.

SAKIKO audits whether internal interventions actually “repair” pre-tool choices such as calling a tool, asking for clarification, answering, or declining. Across seven LLMs on When2Call and MetaTool, the authors report direction-specific net gains in five models, with random directions failing to match calibrated target gains in sealed evaluations. The central warning is that aggregate gains can hide bad destinations: one intervention posted a +55 net gain while corrupting more than half of the baseline-correct decisions it touched. The authors also decline some apparently promising results because the samples were too uncertain to license the claim. Source: HF Daily Papers' note

score 4

Categories: Research