FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
FlowBalance tries to make self-training use a model’s own reasoning without letting its own confidence become the teacher.
The method keeps verifier outcomes in control, using them to retain, reverse, or disable same-model guidance depending on rollout results. It fits a normalized target distribution over complete responses rather than adding a separate token-level imitation loss. In math reasoning tests, the paper reports gains over FlowRL on Qwen3-4B and Qwen3-8B, along with better speed, stability, and strategy diversity. HF Daily Papers' note
The method keeps verifier outcomes in control, using them to retain, reverse, or disable same-model guidance depending on rollout results. It fits a normalized target distribution over complete responses rather than adding a separate token-level imitation loss. In math reasoning tests, the paper reports gains over FlowRL on Qwen3-4B and Qwen3-8B, along with better speed, stability, and strategy diversity. HF Daily Papers' note
score 5