JEPA-TTT: Persistent Test-Time Training of Latent World Models for Planning under Dynamics Shifts
JEPA-TTT keeps adapting a pretrained latent world model during deployment, improving planning when environment dynamics shift.
The method updates only the latent dynamics predictor, leaving the visual encoder and reward head fixed. Its self-supervised training accumulates across test-time episodes using dense replay from a growing buffer. The authors report gains across eight dynamics shifts in four continuous-control environments, including an 83% average reduction in latent prediction error after 500 episodes. Planning performance rose 153% over the frozen JEPA world model, without needing a goal image or online reward. HF Daily Papers' note
The method updates only the latent dynamics predictor, leaving the visual encoder and reward head fixed. Its self-supervised training accumulates across test-time episodes using dense replay from a growing buffer. The authors report gains across eight dynamics shifts in four continuous-control environments, including an 83% average reduction in latent prediction error after 500 episodes. Planning performance rose 153% over the frozen JEPA world model, without needing a goal image or online reward. HF Daily Papers' note
score 5