Rufus-Air: An Open LLM Post-Training Recipe
Rufus-Air lays out an eight-stage, reproducible post-training pipeline for GLM-4.5-Air-Base.
The paper documents the data, rewards, infrastructure, ordering, and stagewise results behind the recipe. Its stages move from SFT through several RL and agent-training phases, ending with RLHF. The authors say the recipe uses open-source components and public data, largely without new human annotation or an in-house distillation teacher. They report that Rufus-Air beats the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models. HF Daily Papers' note
The paper documents the data, rewards, infrastructure, ordering, and stagewise results behind the recipe. Its stages move from SFT through several RL and agent-training phases, ending with RLHF. The authors say the recipe uses open-source components and public data, largely without new human annotation or an in-house distillation teacher. They report that Rufus-Air beats the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models. HF Daily Papers' note
score 6