MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings
MeetingToM tests whether multimodal models can read hidden disagreement and social cues in meetings.
The benchmark targets theory-of-mind reasoning across speech and behavior in multi-party settings. It evaluates models at subject, dyadic, and group levels, including mental state prediction, addressee understanding, and consensus reasoning. The authors say representative MLLMs still struggle with non-verbal cues, hidden attitudes, and telling real consensus from pseudo-consensus. ArXiv · AI/CL/LG's note
The benchmark targets theory-of-mind reasoning across speech and behavior in multi-party settings. It evaluates models at subject, dyadic, and group levels, including mental state prediction, addressee understanding, and consensus reasoning. The authors say representative MLLMs still struggle with non-verbal cues, hidden attitudes, and telling real consensus from pseudo-consensus. ArXiv · AI/CL/LG's note
score 4