ProAR: Learning Prospective Reasoning with Autoregressive Video Models
ProAR turns autoregressive video generation into a goal-guided reasoning process, instead of only predicting the next chunk.
The paper adds goal-frame prediction inside the autoregressive loop so intermediate states are generated with the target outcome in view. It also aligns current hidden states with future representations to improve short-range transitions. The authors say these mechanisms improve performance across visual reasoning benchmarks with modest compute cost. ProAR reportedly beats fully trained standard AR baselines using 25% of the training steps, with possible use in embodied reasoning. HF Daily Papers' note
The paper adds goal-frame prediction inside the autoregressive loop so intermediate states are generated with the target outcome in view. It also aligns current hidden states with future representations to improve short-range transitions. The authors say these mechanisms improve performance across visual reasoning benchmarks with modest compute cost. ProAR reportedly beats fully trained standard AR baselines using 25% of the training steps, with possible use in embodied reasoning. HF Daily Papers' note
score 4