WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning
WorldToken compresses each robot timestep into a single fused “world token,” then models the sequence over time.
The paper reports 59.45% mean closed-loop success across 23 RoboCasa tasks with an 85.3M-parameter policy trained mostly from scratch. Its sweep found more target-domain data helped consistently, while larger models showed diminishing returns past moderate size. Truncating visible history hurt success across all tested RoboCasa policies, and RMBench performance fell from 95% to 28% when history dropped from 146 to 8 seconds. The authors say the results show feasibility and temporal-context sensitivity, not superiority over other sequence layouts. HF Daily Papers' note
The paper reports 59.45% mean closed-loop success across 23 RoboCasa tasks with an 85.3M-parameter policy trained mostly from scratch. Its sweep found more target-domain data helped consistently, while larger models showed diminishing returns past moderate size. Truncating visible history hurt success across all tested RoboCasa policies, and RMBench performance fell from 95% to 28% when history dropped from 146 to 8 seconds. The authors say the results show feasibility and temporal-context sensitivity, not superiority over other sequence layouts. HF Daily Papers' note
score 5