ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context
Progress reward models failed mainly when the missing ingredient was task context, not model capacity.
The paper introduces ContextProgress-Bench, a 24-task benchmark testing whether embodied agents can judge progress when the current frame is ambiguous. In paired tests, adding the right context cut progress-estimation error by 77-82% across five models. The authors then propose ProgressCompass, an agentic loop that supplies context to a frozen PRM, reducing error by 63% and improving rank agreement by 76%. Source: HF Daily Papers' note
The paper introduces ContextProgress-Bench, a 24-task benchmark testing whether embodied agents can judge progress when the current frame is ambiguous. In paired tests, adding the right context cut progress-estimation error by 77-82% across five models. The authors then propose ProgressCompass, an agentic loop that supplies context to a frozen PRM, reducing error by 63% and improving rank agreement by 76%. Source: HF Daily Papers' note
score 4