Megadose AI progress, ranked and analyzed.

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

· HF Daily Papers ·
AdviSD trains a small advisor to improve frozen frontier models by selectively learning from the advice that actually changes outcomes.

The method combines outcome-based reinforcement learning with self-distillation from a feedback-conditioned advisor copy. It keeps supervision targets when the advisor’s issued advice meaningfully changes the recorded executor response, avoiding corrections the paper argues can dilute learning. In experiments, Qwen3-8B advisors improved Gemini and Claude executors over advisor-GRPO on BFCL-v3 and EnvScaler. The paper also reports out-of-domain generalization and transfer across executor versions and model families. HF Daily Papers' note

score 4

Categories: Research