HarnessEval-W: Agentifying the Evaluation of Visual Worlds
The paper argues world-model benchmarks need inspectable reasoning, not just a final score.
HarnessEval-W evaluates generated visual rollouts by breaking each case into subproblems and assigning specialized agents to gather evidence. A parent agent checks that evidence and produces a verdict, leaving a traceable “evidence tree” behind the score. The authors tested it on 18 world models across 330 cases and report close alignment with human preferences. HF Daily Papers' note
HarnessEval-W evaluates generated visual rollouts by breaking each case into subproblems and assigning specialized agents to gather evidence. A parent agent checks that evidence and produces a verdict, leaving a traceable “evidence tree” behind the score. The authors tested it on 18 world models across 330 cases and report close alignment with human preferences. HF Daily Papers' note
score 5