4DAnyone: Create Anyone in 4D from a Casual Monocular Video
The paper’s core claim is that stable 4D human reconstruction needs many novel views that stay consistent, not just plausible.
4DAnyone generates multiview-consistent videos from an uncalibrated single-camera clip, then lifts them into 4D Gaussian Splatting. The authors say prior camera-controlled diffusion models drift when scaled to the tens of target views needed for reconstruction. Their fixes are Reference Context Packing, which keeps conditioning fixed-size, and Target Context Routing, which lets target-view groups share context during denoising. They report stronger novel-view video quality and 4DGS reconstruction on DNA-Rendering and DyMVHumans, with in-the-wild generalization. HF Daily Papers' note
4DAnyone generates multiview-consistent videos from an uncalibrated single-camera clip, then lifts them into 4D Gaussian Splatting. The authors say prior camera-controlled diffusion models drift when scaled to the tens of target views needed for reconstruction. Their fixes are Reference Context Packing, which keeps conditioning fixed-size, and Target Context Routing, which lets target-view groups share context during denoising. They report stronger novel-view video quality and 4DGS reconstruction on DNA-Rendering and DyMVHumans, with in-the-wild generalization. HF Daily Papers' note
score 5