paper-with-me

홈 › Papers

MimeQA: Towards Socially-Intelligent Nonverbal Foundation Models

2025-02-23 · Hengzhi Li, Megan Tjandrasuwita, Yi R. Fung, Armando Solar-Lezama, Paul Pu Liang

Socially intelligent AI that can understand and interact seamlessly with humans in daily lives is increasingly important as AI becomes more closely integrated with peoples' daily activities. However, current works in artificial social reasoning all rely on language-only, or language-dominant approaches to benchmark and training models, resulting in systems that are improving in verbal communication but struggle with nonverbal social understanding. To address this limitation, we tap into a novel source of data rich in nonverbal and social interactions -- mime videos. Mimes refer to the art of expression through gesture and movement without spoken words, which presents unique challenges and opportunities in interpreting non-verbal social communication. We contribute a new dataset called MimeQA, obtained by sourcing 221 videos from YouTube, through rigorous annotation and verification, resulting in a benchmark with 101 videos and 806 question-answer pairs. Using MimeQA, we evaluate state-of-the-art video large language models (vLLMs) and find that their overall accuracy ranges from 15-30%. Our analysis reveals that vLLMs often fail to ground imagined objects and over-rely on the text prompt while ignoring subtle nonverbal interactions. Our data resources are released at https://github.com/MIT-MI/MimeQA to inspire future work in foundation models that embody true social intelligence capable of interpreting non-verbal human interactions.

📄 PDF Abstract BibTeX arXiv:2502.16671

Code (1)

mit-mi/mimeqa 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Using Vision-Language Models as Proxies for Social Intelligence in Human-Robot Interaction

2025-12-08 · Fanjun Bu, Melina Tsai, Audrey Tjokro, Tapomayukh Bhattacharjee 외 arxiv

Robots operating in everyday environments must often decide when and whether to engage with people, yet such decisions often hinge on subtle nonverbal cues that unfold over time and are difficult to model explicitly. Dra…

Robotic Backchanneling in Online Conversation Facilitation: A Cross-Generational Study

2024-09-25 · Sota Kobuki, Katie Seaborn, Seiki Tokunaga, Kosuke Fukumori 외

Japan faces many challenges related to its aging society, including increasing rates of cognitive decline in the population and a shortage of caregivers. Efforts have begun to explore solutions using artificial intellige…

SalsaAgent: A multimodal embodied language model for interactive dance generation

2026-05-28 · Payam Jome Yazdian, Zoe Stanley, Angelica Lim arxiv

Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive virtual agents, we present SalsaAgent, a language model that generat…

Towards a Unifying Model of Rationality in Multiagent Systems

2023-05-29 · Robert Loftin, Mustafa Mert Çelikok, Frans A. Oliehoek

Multiagent systems deployed in the real world need to cooperate with other agents (including humans) nearly as effectively as these agents cooperate with one another. To design such AI, and provide guarantees of its effe…

Socially-Aware Animated Intelligent Personal Assistant Agent

2016-09-01 · WS 2016 9 · Yoichi Matsuyama, Arjun Bhardwaj, Ran Zhao, Oscar Romeo 외
Speech Recognition