Megadose Built for builders and researchers.

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

· HF Daily Papers ·
The paper argues world-model benchmarks need inspectable reasoning, not just a final score.

HarnessEval-W evaluates generated visual rollouts by breaking each case into subproblems and assigning specialized agents to gather evidence. A parent agent checks that evidence and produces a verdict, leaving a traceable “evidence tree” behind the score. The authors tested it on 18 world models across 330 cases and report close alignment with human preferences. HF Daily Papers' note

score 5

Categories: Research