Megadose AI progress, ranked and analyzed.

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

· ArXiv · AI/CL/LG ·
Paraphrased robot goals can make the same behavior score as both success and failure.

The paper says current vision-language reward models often change their progress judgments when only the wording of the instruction changes. It introduces ROBORMBENCH, with 2,390 real-robot trajectories and 21,673 verified paraphrases, to measure that failure mode. The instability appears across proprietary and open-source models, worsens with more divergent rewrites, and is not reliably fixed by model scale or explicit reasoning. Reward models trained with trajectory-grounded supervision were substantially more stable. ArXiv · AI/CL/LG's note

score 5

Categories: Research