Keeping JEPA World Models Plannable When Little of the Frame Moves
A JEPA planner failed when small object motion barely changed the image, and an inverse-dynamics loss made its latent state usable again.
The authors introduce SLIM, a pushing benchmark with small objects and paired visual and language goals. A LeWM model that works on PushT solved under 1% of SLIM trials because its encoder latent was nearly insensitive to actions and could not expose pusher or object positions. Adding a shared inverse-dynamics auxiliary loss lifted success from 0.003 to 0.35 and restored the diagnostic probes. With the repaired latent, a small language-goal head could plan from sentences for navigation and staged pushing without retraining the world model. ArXiv · AI/CL/LG's note
The authors introduce SLIM, a pushing benchmark with small objects and paired visual and language goals. A LeWM model that works on PushT solved under 1% of SLIM trials because its encoder latent was nearly insensitive to actions and could not expose pusher or object positions. Adding a shared inverse-dynamics auxiliary loss lifted success from 0.003 to 0.35 and restored the diagnostic probes. With the repaired latent, a small language-goal head could plan from sentences for navigation and staged pushing without retraining the world model. ArXiv · AI/CL/LG's note
score 4