FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
The paper turns origami videos into executable, physically plausible folding programs.
FoldingAgent uses a pretrained vision-language model with tools for geometry simulation, plausibility checks, visual comparison, and self-evaluation. It works step by step in a parametric folding space, with the ability to re-plan when earlier actions would compound errors. The authors evaluate it on PurelandFold, a new benchmark of Pureland origami videos with ground-truth geometry and action labels. Accepted to SIGGRAPH ASIA 2026. HF Daily Papers' note
FoldingAgent uses a pretrained vision-language model with tools for geometry simulation, plausibility checks, visual comparison, and self-evaluation. It works step by step in a parametric folding space, with the ability to re-plan when earlier actions would compound errors. The authors evaluate it on PurelandFold, a new benchmark of Pureland origami videos with ground-truth geometry and action labels. Accepted to SIGGRAPH ASIA 2026. HF Daily Papers' note
score 4