PhiZero: A World Model Built Around Physical Language
PhiZero turns future video prediction into a reason-then-render step using “physical language.”
The paper describes a compact discrete representation for world-state transitions, learned from in-the-wild videos through self-supervision. Instead of predicting future frames directly in pixel space, PhiZero first infers a sequence of physical-language transitions, then renders them back into video. The authors report results across generation and understanding benchmarks, and point to uses in interactive world modeling, action-conditioned simulation, and zero-shot motion transfer. HF Daily Papers' note
The paper describes a compact discrete representation for world-state transitions, learned from in-the-wild videos through self-supervision. Instead of predicting future frames directly in pixel space, PhiZero first infers a sequence of physical-language transitions, then renders them back into video. The authors report results across generation and understanding benchmarks, and point to uses in interactive world modeling, action-conditioned simulation, and zero-shot motion transfer. HF Daily Papers' note
score 5