Megadose Built for builders and researchers.

Does Physics Live in the Activations? Localizing Physical Quantities in Video Diffusion Models

· ArXiv · AI/CL/LG ·
The paper says video diffusion models build readable physical variables during denoising, then store them in localized object-token activations.

The authors probe video DiT internals using simulator-derived ground truth for motion and rigid-body dynamics. They report that physical quantities are linearly decodable early in denoising, beyond what can be recovered from the model’s noised latents. Multi-frame quantities can be read from single latent frames, and the probes partly transfer outside their training setups. The same activation directions can also steer generated outputs when fitted at full resolution. ArXiv · AI/CL/LG's note

score 5

Categories: Research