Megadose AI progress, ranked and analyzed.

AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos

· ArXiv · AI/CL/LG ·
AgentSTAR tracks objects by building a structured 3D model first, then using it to refine motion over time.

The paper says the method avoids the usual dense pixel-correspondence step. A VLM agent iteratively adjusts shape or generalized pose through a render-and-compare loop, paired with numerical pose optimization. The authors say that lets it handle large motion, articulation, and heavy occlusion. They report stronger results than evaluated 3D point-tracking baselines on ARCTIC and rigid-object tracking baselines on HOT3D. ArXiv · AI/CL/LG's note

score 4

Categories: Research