Can Edge-Deployable Vision-Language Models Identify Species?
Small VLMs can spot species above chance, but field camera-trap images remain the hard part.
The paper tests Qwen3-VL 2B/4B/8B and Gemma3 4B against BioCLIP on a 96-species task. Every model drops sharply when moving from clean iNaturalist photos to camera-trap imagery, with domain gaps of 9.6 to 26.6 points. BioCLIP beats all tested VLMs by wide margins despite being much smaller, pointing to specialized training data rather than scale. Open-set prompting also produced nonexistent species names in 5.9% to 9.6% of responses. ArXiv · AI/CL/LG's note
The paper tests Qwen3-VL 2B/4B/8B and Gemma3 4B against BioCLIP on a 96-species task. Every model drops sharply when moving from clean iNaturalist photos to camera-trap imagery, with domain gaps of 9.6 to 26.6 points. BioCLIP beats all tested VLMs by wide margins despite being much smaller, pointing to specialized training data rather than scale. Open-set prompting also produced nonexistent species names in 5.9% to 9.6% of responses. ArXiv · AI/CL/LG's note
score 4