Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
The paper finds the weak point is not drone navigation, but whether the model follows the action protocol and stops correctly.
DroneCATS tests multimodal LLMs as swappable drone-control agents across approaching, tracking, searching, and multi-drone commanding. The authors say small open models can often reach the target radius, but lose by declaring arrival too early or failing to declare it. Frontier models are not presented as having solved the setting either. In fleet control, smaller models also collapse distinct views into copied coordinates. HF Daily Papers' note
DroneCATS tests multimodal LLMs as swappable drone-control agents across approaching, tracking, searching, and multi-drone commanding. The authors say small open models can often reach the target radius, but lose by declaring arrival too early or failing to declare it. Frontier models are not presented as having solved the setting either. In fleet control, smaller models also collapse distinct views into copied coordinates. HF Daily Papers' note
score 4