AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models
AnyStep-WAM cuts denoising work by choosing cheaper generation only when the action chunk can tolerate it.
The paper trains world-action models to operate across budgets, from one-step prediction to multi-step refinement. A scheduler uses a one-step preview to estimate task difficulty and pick the smallest denoising budget expected to preserve fidelity. Across Motus, FastWAM, and LingBotVA on RoboTwin 2.0, it reduces average denoising steps by 60.2%, 49.8%, and 85.28% while maintaining baseline success rates. The authors also report better one-step success rates and validation on six real-world manipulation tasks. HF Daily Papers' note
The paper trains world-action models to operate across budgets, from one-step prediction to multi-step refinement. A scheduler uses a one-step preview to estimate task difficulty and pick the smallest denoising budget expected to preserve fidelity. Across Motus, FastWAM, and LingBotVA on RoboTwin 2.0, it reduces average denoising steps by 60.2%, 49.8%, and 85.28% while maintaining baseline success rates. The authors also report better one-step success rates and validation on six real-world manipulation tasks. HF Daily Papers' note
score 5