Megadose AI progress, ranked and analyzed.

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

· HF Daily Papers ·
Lucida shifts precision to the end of real-to-sim reconstruction, using a VLM placement policy to align generated objects in closed loop.

The paper keeps the parse-generate-place pipeline but changes what each stage is expected to know from cluttered real captures. It parses video into a scene graph with per-object multi-view evidence, then generates complete assets for each instance. Placement is handled by GizmoAct, which treats alignment as multi-turn GUI manipulation and stops when it judges the object is aligned. The authors report gains on detection, pose estimation, and reconstruction benchmarks, including [email protected] rising from 57.8% to 83.4% on CA-1M. HF Daily Papers' note

score 5

Categories: Research