RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
The paper argues that timestamp-derived temporal distance can replace preference labels as supervision for robot reward models.
RynnValue is trained on more than 7,000 hours and about 3 million instruction-conditioned clips by estimating directed cost-to-go from an observation to a language-specified goal. The authors say this avoids task-specific anchors like preferences or normalized progress, which transfer poorly across robot embodiments and datasets. On RBM-EVAL-OOD, it reports a Kendall’s tau_a of 0.675, above a fully preference-supervised baseline at 0.655. Used as dense rewards, it raises reported real-world policy success to 72.5% online and 82.5% offline. ArXiv · AI/CL/LG's note
RynnValue is trained on more than 7,000 hours and about 3 million instruction-conditioned clips by estimating directed cost-to-go from an observation to a language-specified goal. The authors say this avoids task-specific anchors like preferences or normalized progress, which transfer poorly across robot embodiments and datasets. On RBM-EVAL-OOD, it reports a Kendall’s tau_a of 0.675, above a fully preference-supervised baseline at 0.655. Used as dense rewards, it raises reported real-world policy success to 72.5% online and 82.5% offline. ArXiv · AI/CL/LG's note
score 6