Megadose AI progress, ranked and analyzed.

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

· HF Daily Papers ·
The paper tests whether models can use helpful workflows without getting trapped by bad ones.

The authors introduce Box²-Bench, varying workflow reliability while keeping the model and task fixed. Frontier models often improve with reliable guidance but remain exposed when that guidance misleads or degrades. Training on bad workflows made two open-weight models more robust, while reinforcement learning pushed models toward using helpful workflows more. The same selective-reliance behavior also appeared in peer correction and corrupted-memory settings. HF Daily Papers' note

score 4

Categories: Research