FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
The test asks whether models can tell friends from strangers using only how two people interact.
FriendBench uses 20-second ice-breaker clips from 96 balanced pairs, with each pair answering the same kind of prompt. The authors compare 26 multimodal models with matched human panels across text, audio, and video. The best model and the human crowd are statistically tied on accuracy, but the models skew toward guessing “stranger” while humans stay more balanced. Only humans improve from visible behavior beyond speech. ArXiv · AI/CL/LG's note
FriendBench uses 20-second ice-breaker clips from 96 balanced pairs, with each pair answering the same kind of prompt. The authors compare 26 multimodal models with matched human panels across text, audio, and video. The best model and the human crowd are statistically tied on accuracy, but the models skew toward guessing “stranger” while humans stay more balanced. Only humans improve from visible behavior beyond speech. ArXiv · AI/CL/LG's note
score 4