Megadose AI progress, ranked and analyzed.

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

· HF Daily Papers ·
FlowBalance tries to make self-training use a model’s own reasoning without letting its own confidence become the teacher.

The method keeps verifier outcomes in control, using them to retain, reverse, or disable same-model guidance depending on rollout results. It fits a normalized target distribution over complete responses rather than adding a separate token-level imitation loss. In math reasoning tests, the paper reports gains over FlowRL on Qwen3-4B and Qwen3-8B, along with better speed, stability, and strategy diversity. HF Daily Papers' note

score 5

Categories: Research