CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents
CAST trains agents to catch bad tool actions before they become irreversible failures.
The paper turns sparse task outcomes into action-level critique supervision, then uses those critiques to improve the policy model. Its target is long-horizon, stateful tool use where mistakes like acting on the wrong purchase can end a task. Fine-tuned Qwen3-family models improved reliability on dynamic benchmarks, including a reported 10%+ pass^4 gain over GPT-OSS-120B on Retail tasks and a 9% out-of-domain gain on Telehealth. HF Daily Papers' note
The paper turns sparse task outcomes into action-level critique supervision, then uses those critiques to improve the policy model. Its target is long-horizon, stateful tool use where mistakes like acting on the wrong purchase can end a task. Fine-tuned Qwen3-family models improved reliability on dynamic benchmarks, including a reported 10%+ pass^4 gain over GPT-OSS-120B on Retail tasks and a 9% out-of-domain gain on Telehealth. HF Daily Papers' note
score 5