Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models
Paraphrased robot goals can make the same behavior score as both success and failure.
The paper says current vision-language reward models often change their progress judgments when only the wording of the instruction changes. It introduces ROBORMBENCH, with 2,390 real-robot trajectories and 21,673 verified paraphrases, to measure that failure mode. The instability appears across proprietary and open-source models, worsens with more divergent rewrites, and is not reliably fixed by model scale or explicit reasoning. Reward models trained with trajectory-grounded supervision were substantially more stable. ArXiv · AI/CL/LG's note
The paper says current vision-language reward models often change their progress judgments when only the wording of the instruction changes. It introduces ROBORMBENCH, with 2,390 real-robot trajectories and 21,673 verified paraphrases, to measure that failure mode. The instability appears across proprietary and open-source models, worsens with more divergent rewrites, and is not reliably fixed by model scale or explicit reasoning. Reward models trained with trajectory-grounded supervision were substantially more stable. ArXiv · AI/CL/LG's note
score 5