Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Capek 0.5 trains embodied VLM abilities as execution roles, then merges them into one runtime model.
The paper groups robot-agent capabilities into Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification. Each starts as a reinforcement-learned specialist with verifiable rewards from a shared backbone, then is consolidated through weight-space merging and routed policy-space distillation. The authors report 2B and 35B-A3B versions, including tests on Capek-StateBench and closed-loop simulated embodied environments. They say the unified model improves most matched benchmark rows over its initialization and retains all four specialist capabilities with measured losses. HF Daily Papers' note
The paper groups robot-agent capabilities into Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification. Each starts as a reinforcement-learned specialist with verifiable rewards from a shared backbone, then is consolidated through weight-space merging and routed policy-space distillation. The authors report 2B and 35B-A3B versions, including tests on Capek-StateBench and closed-loop simulated embodied environments. They say the unified model improves most matched benchmark rows over its initialization and retains all four specialist capabilities with measured losses. HF Daily Papers' note
score 5