Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
The paper says composed rewards beat a simple weighted blend for post-training text-to-image models.
The authors combine a human-preference reward with rubric rewards for prompt faithfulness and other checks. They argue the setup helps avoid reward hacking while still improving visual preference. Their RL-trained Flux2dev is reported 69 Elo points above its base model on the Arena text-to-image leaderboard. They also say post-trained Ideogram-4 reached 1223.5 Elo and surpassed every open-source model in the September 4, 2026 leaderboard snapshot. HF Daily Papers' note
The authors combine a human-preference reward with rubric rewards for prompt faithfulness and other checks. They argue the setup helps avoid reward hacking while still improving visual preference. Their RL-trained Flux2dev is reported 69 Elo points above its base model on the Arena text-to-image leaderboard. They also say post-trained Ideogram-4 reached 1223.5 Elo and surpassed every open-source model in the September 4, 2026 leaderboard snapshot. HF Daily Papers' note
score 5