WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
WorldDiT pairs robot action generation with future-frame prediction in one diffusion transformer.
The paper says the model generates continuous action chunks while also predicting normalized RGB patch targets from future camera views. Its reported results cover four LIBERO simulation suites. The authors frame the result as a strong sub-billion-parameter baseline, reaching the reported Pareto frontier for parameter count and mean success among methods that report all four suites. HF Daily Papers' note
The paper says the model generates continuous action chunks while also predicting normalized RGB patch targets from future camera views. Its reported results cover four LIBERO simulation suites. The authors frame the result as a strong sub-billion-parameter baseline, reaching the reported Pareto frontier for parameter count and mean success among methods that report all four suites. HF Daily Papers' note
score 5