WorldReward: Reward Modeling for Camera-Conditioned World Models
WorldReward judges whether camera commands actually show up in the video, without separating that from visual quality.
The model breaks paired videos into action-aligned chunks, turns each chunk into structured visual evidence, and votes those judgments into video-level preferences. The authors trained it on a reasoning-augmented preference dataset generated by a frontier VLM, then audited with tools and targeted human review. On WorldReward-Bench, it beats GPT-5.5 agreement with human preferences across action consistency, appearance quality, and motion quality. Used for RL post-training of HY-WorldPlay 1.5, it improves both action execution and visual quality across short and long horizons. HF Daily Papers' note
The model breaks paired videos into action-aligned chunks, turns each chunk into structured visual evidence, and votes those judgments into video-level preferences. The authors trained it on a reasoning-augmented preference dataset generated by a frontier VLM, then audited with tools and targeted human review. On WorldReward-Bench, it beats GPT-5.5 agreement with human preferences across action consistency, appearance quality, and motion quality. Used for RL post-training of HY-WorldPlay 1.5, it improves both action execution and visual quality across short and long horizons. HF Daily Papers' note
score 4