FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
The paper tests JEPA-style world models on dense Global South street scenes and claims factorized futures hold up better than monolithic latents.
The authors introduce DENSEWORLD, a 1,000-hour video dataset from drive-through, walk-through, and aerial footage across 22 cities. FactorJEPA splits prediction into layout, entities, and interactions, with visibility gating meant to handle occlusion and partially observed agents. They report gains on future-frame accuracy, causal prediction sensitivity, and robustness when visual evidence is reduced. Method rankings held across 2B and 1B V-JEPA 2.1 backbones. Source: HF Daily Papers' note.
The authors introduce DENSEWORLD, a 1,000-hour video dataset from drive-through, walk-through, and aerial footage across 22 cities. FactorJEPA splits prediction into layout, entities, and interactions, with visibility gating meant to handle occlusion and partially observed agents. They report gains on future-frame accuracy, causal prediction sensitivity, and robustness when visual evidence is reduced. Method rankings held across 2B and 1B V-JEPA 2.1 backbones. Source: HF Daily Papers' note.
score 5