LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
LIFT lets image-to-video users pin the layout of a future view, including what should appear in the final frame.
The paper frames this as a control gap: camera paths set motion, while text prompts do not fix spatial content in newly revealed areas. LIFT uses the last-frame layout as an explicit condition for large viewpoint changes. Its on-policy self-distillation trains a student model from a dense-layout teacher so sparse future-layout guidance can work. The authors also introduce LIFT-Vista and report gains in video quality, future-layout control, and camera control. HF Daily Papers' note
The paper frames this as a control gap: camera paths set motion, while text prompts do not fix spatial content in newly revealed areas. LIFT uses the last-frame layout as an explicit condition for large viewpoint changes. Its on-policy self-distillation trains a student model from a dense-layout teacher so sparse future-layout guidance can work. The authors also introduce LIFT-Vista and report gains in video quality, future-layout control, and camera control. HF Daily Papers' note
score 4