H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning
H-JEPA reports a jump from 18% to 73% success on Visual AntMaze by planning through a learned hierarchy.
The paper trains action-conditioned JEPA models at multiple levels, with each level predicting farther ahead in its own latent space. Planning runs top-down, turning higher-level predictions into subgoals for lower levels. The authors say the hierarchy drops fast, unpredictable visual detail while keeping slower task-relevant state when the data separates that way. Tests cover four simulated navigation and manipulation environments, plus offline planning on DROID robot videos with inverse-dynamics supervision. ArXiv · AI/CL/LG's note
The paper trains action-conditioned JEPA models at multiple levels, with each level predicting farther ahead in its own latent space. Planning runs top-down, turning higher-level predictions into subgoals for lower levels. The authors say the hierarchy drops fast, unpredictable visual detail while keeping slower task-relevant state when the data separates that way. Tests cover four simulated navigation and manipulation environments, plus offline planning on DROID robot videos with inverse-dynamics supervision. ArXiv · AI/CL/LG's note
score 6