Megadose AI progress, ranked and analyzed.

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

· HF Daily Papers ·
ReferTrack grounds embodied visual tracking in explicit bounding-box selection before planning the robot’s path.

The model uses a single forward-facing camera to choose the referred target from indexed boxes, then predicts tracking waypoints from that image-grounded decision. It keeps a sliding window of prior selected boxes and feeds their geometry back into the visual history with TVBI tokens. On EVT-Bench, the paper reports success rates of 89.4%, 73.3%, and 74.1% across single-target, distracted, and ambiguity splits. The authors also report real-world tests on legged and humanoid robots. HF Daily Papers' note

score 4

Categories: Research