AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
AdviSD trains a small advisor to improve frozen frontier models by selectively learning from the advice that actually changes outcomes.
The method combines outcome-based reinforcement learning with self-distillation from a feedback-conditioned advisor copy. It keeps supervision targets when the advisor’s issued advice meaningfully changes the recorded executor response, avoiding corrections the paper argues can dilute learning. In experiments, Qwen3-8B advisors improved Gemini and Claude executors over advisor-GRPO on BFCL-v3 and EnvScaler. The paper also reports out-of-domain generalization and transfer across executor versions and model families. HF Daily Papers' note
The method combines outcome-based reinforcement learning with self-distillation from a feedback-conditioned advisor copy. It keeps supervision targets when the advisor’s issued advice meaningfully changes the recorded executor response, avoiding corrections the paper argues can dilute learning. In experiments, Qwen3-8B advisors improved Gemini and Claude executors over advisor-GRPO on BFCL-v3 and EnvScaler. The paper also reports out-of-domain generalization and transfer across executor versions and model families. HF Daily Papers' note
score 4