SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
SyncWorld uses a short visual calibration episode to make robot actions readable across new camera views, environments, and embodiments.
The paper frames the problem as action labels breaking down in pixel space: the same numerical command can look different when the setup changes. SyncWorld conditions its rollouts on paired frames and actions that reveal the setup’s action-to-visual mapping, then uses that context to simulate outcomes without extra training. The authors say it can also draw on interaction history when explicit calibration is missing. In experiments, they report accurate rollouts in unseen settings and test-time policy improvement without training. HF Daily Papers' note
The paper frames the problem as action labels breaking down in pixel space: the same numerical command can look different when the setup changes. SyncWorld conditions its rollouts on paired frames and actions that reveal the setup’s action-to-visual mapping, then uses that context to simulate outcomes without extra training. The authors say it can also draw on interaction history when explicit calibration is missing. In experiments, they report accurate rollouts in unseen settings and test-time policy improvement without training. HF Daily Papers' note
score 5