ShotPlan: Cinematic Video Generation with Learnable Planning Token
ShotPlan adds explicit, frame-level shot planning to video diffusion for multi-shot cinematic generation.
The paper says current models are strong at single shots but weaker at coherent multi-shot sequences. Its planning tokens capture transition cues and control when shot changes happen. FRoPE lets those transitions be modeled at the frame level. The authors report stronger inter-shot consistency and more flexible shot management than existing cinematic video generation methods. HF Daily Papers' note
The paper says current models are strong at single shots but weaker at coherent multi-shot sequences. Its planning tokens capture transition cues and control when shot changes happen. FRoPE lets those transitions be modeled at the frame level. The authors report stronger inter-shot consistency and more flexible shot management than existing cinematic video generation methods. HF Daily Papers' note
score 5