SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
SG-WAM reports stronger robot task success by predicting future scene dynamics inside the policy’s own geometry-aware latent space.
The paper says the model uses learnable dynamics tokens and a self-guided predictor conditioned on robot actions. Its targets come from an exponential-moving-average copy of the same policy backbone, keeping supervision inside the action model’s representation family. The authors report 98.5% average success on LIBERO and 73% on LIBERO-Plus using a 0.9B model without large-scale embodied pretraining. HF Daily Papers' note
The paper says the model uses learnable dynamics tokens and a self-guided predictor conditioned on robot actions. Its targets come from an exponential-moving-average copy of the same policy backbone, keeping supervision inside the action model’s representation family. The authors report 98.5% average success on LIBERO and 73% on LIBERO-Plus using a 0.9B model without large-scale embodied pretraining. HF Daily Papers' note
score 5