OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
OpenWAM turns world-action model pretraining into a modular testbed, then ships the resulting stack.
The paper separates the generative backbone, visual representation, architecture, information flow, inference, and data choices so they can be tested independently. Its study argues that WAMs need a capable video prior, compact latent space, dedicated action capacity, explicit world-to-action flow, and synchronized joint denoising. The authors build OpenWAM-alpha on about 6,400 hours of egocentric human and robot data. They report strong results across eight simulation benchmarks and real-robot tests spanning single-arm, bimanual, and dexterous-hand settings. HF Daily Papers' note
The paper separates the generative backbone, visual representation, architecture, information flow, inference, and data choices so they can be tested independently. Its study argues that WAMs need a capable video prior, compact latent space, dedicated action capacity, explicit world-to-action flow, and synchronized joint denoising. The authors build OpenWAM-alpha on about 6,400 hours of egocentric human and robot data. They report strong results across eight simulation benchmarks and real-robot tests spanning single-arm, bimanual, and dexterous-hand settings. HF Daily Papers' note
score 5