Hierarchical Denoising For Multi-Step Visual Reasoning
HDR raises multi-step video reasoning success from 34.22 to 60.29 while keeping streaming latency low.
The paper proposes a tree-structured latent hierarchy for causal video generation, letting models plan coarsely before refining visual states. Its benchmark covers six reasoning tasks, including maze navigation, Tower of Hanoi, Sokoban, and water pouring, with out-of-distribution cases. The authors report 0.70 seconds per latent and 54.2x faster inference than bidirectional diffusion. They also say HDR keeps 82.9% of full-data performance using only 2% of the training data. HF Daily Papers' note
The paper proposes a tree-structured latent hierarchy for causal video generation, letting models plan coarsely before refining visual states. Its benchmark covers six reasoning tasks, including maze navigation, Tower of Hanoi, Sokoban, and water pouring, with out-of-distribution cases. The authors report 0.70 seconds per latent and 54.2x faster inference than bidirectional diffusion. They also say HDR keeps 82.9% of full-data performance using only 2% of the training data. HF Daily Papers' note
score 5