RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
RynnValue trains robot reward signals from timestamp-derived temporal distance instead of preference labels.
The paper says this lets the model scale across more than 7,000 hours and about 3 million instruction-conditioned clips. Its training setup is designed to avoid shortcuts that miss failures or regressions. On RBM-EVAL-OOD, it reports a Kendall’s tau_a of 0.675, above a fully preference-supervised baseline at 0.655. Used as dense rewards, it improved reported real-world policy success both online and offline. HF Daily Papers' note
The paper says this lets the model scale across more than 7,000 hours and about 3 million instruction-conditioned clips. Its training setup is designed to avoid shortcuts that miss failures or regressions. On RBM-EVAL-OOD, it reports a Kendall’s tau_a of 0.675, above a fully preference-supervised baseline at 0.655. Used as dense rewards, it improved reported real-world policy success both online and offline. HF Daily Papers' note
score 6