Megadose Built for builders and researchers.

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

· HF Daily Papers ·
Dream4ACT turns robot actions into multiview visual “action views” so one video model can work across different robot bodies.

The paper says joint-space action vectors are hard to share across embodiments because their dimensions and meanings differ. Dream4ACT instead renders target joint configurations from four virtual cameras using URDF-based forward kinematics. A jointly trained video autoencoder and diffusion transformer then handle observation and action sequences together. The authors report 88.98% average success on RoboTwin 2.0 and a 65.66 overall score on TriWorldBench. HF Daily Papers' note

score 5

Categories: Research