Asymmetric Capacity Allocation in Self-Refinement Pipelines
The study finds self-refinement systems should spend model capacity on generation and revision, not evenly across all three stages.
Across five benchmarks, larger generators and refiners generally improved results, while too-small refiners could hurt performance. Critic size mattered far less, though even a small critic beat removing critique entirely. The authors tested Qwen3 and Gemma 3 models at multiple sizes and argue that stage-wise scaling can make multi-stage LLM systems more efficient. ArXiv · AI/CL/LG's note
Across five benchmarks, larger generators and refiners generally improved results, while too-small refiners could hurt performance. Critic size mattered far less, though even a small critic beat removing critique entirely. The authors tested Qwen3 and Gemma 3 models at multiple sizes and argue that stage-wise scaling can make multi-stage LLM systems more efficient. ArXiv · AI/CL/LG's note
score 4