GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery
GrabVG narrows UAV visual grounding by turning crowded aerial scenes into a language-guided object graph.
The paper says UAV imagery is hard because small, similar objects cluster densely and create spatial ambiguity. GrabVG first filters the scene into a compact set of object hypotheses, then uses graph attention to bind visual cues with topological relationships between candidates. On AerialVG and AerialSense, it reports 67.31% and 80.34% [email protected], beating the cited baselines by 10.55 and 8.76 points. ArXiv · AI/CL/LG's note
The paper says UAV imagery is hard because small, similar objects cluster densely and create spatial ambiguity. GrabVG first filters the scene into a compact set of object hypotheses, then uses graph attention to bind visual cues with topological relationships between candidates. On AerialVG and AerialSense, it reports 67.31% and 80.34% [email protected], beating the cited baselines by 10.55 and 8.76 points. ArXiv · AI/CL/LG's note
score 4