DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
DreamTraj predicts object motion from diffusion-model internals instead of rendering video first.
The paper introduces MOVE, a dataset of 5,038 egocentric object trajectories paired with fine-grained language instructions. DreamTraj takes a single RGB image and a task instruction, then decodes relative 6-DoF poses from early latent representations inside a frozen image-to-video diffusion model. The authors say it beats prior forecasters on translation and rotation, including systems using multi-frame or privileged inputs, and runs 4.6x faster than generate-then-extract pipelines. HF Daily Papers' note
The paper introduces MOVE, a dataset of 5,038 egocentric object trajectories paired with fine-grained language instructions. DreamTraj takes a single RGB image and a task instruction, then decodes relative 6-DoF poses from early latent representations inside a frozen image-to-video diffusion model. The authors say it beats prior forecasters on translation and rotation, including systems using multi-frame or privileged inputs, and runs 4.6x faster than generate-then-extract pipelines. HF Daily Papers' note
score 5