paper-with-me

홈 › Papers

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions

2026-05-15 · Junho Kim, Xu Cao, Houze Yang, Bikram Boote, Ana Jojic, Fiona Ryan, Bolin Lai, Sangmin Lee, James M. Rehg arxiv

Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with whom in multi-person videos. We introduce GRASP, a large-scale social reasoning dataset that connects high-level social QA with fine-grained gaze and deictic gesture events. GRASP contains 290K question--answer pairs over 46K videos totaling 749 hours, organized by a 16-category taxonomy spanning gaze, gesture, and joint gaze--gesture reasoning, together with GRASP-Bench for evaluation. Unlike prior resources that focus on either isolated cues or high-level social QA, GRASP builds questions from identity-consistent gaze trajectories, deictic gestures, and their joint compositions into social events. Moreover, we propose Social Grounding Reward (SGR), a learning signal that uses these social events to encourage models to reason about the participants involved in each interaction. Experiments show that SGR improves performance on GRASP-Bench while maintaining zero-shot performance on related social video QA benchmarks.

📄 PDF Abstract BibTeX arXiv:2605.15764

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reasoning-Based Personalized Generation for Users with Sparse Data

2026-01-31 · Bo Ni, Branislav Kveton, Samyadeep Basu, Subhojyoti Mukherjee 외 arxiv

Large Language Model (LLM) personalization holds great promise for tailoring responses by leveraging personal context and history. However, real-world users usually possess sparse interaction histories with limited perso…

Text Generation

Graph-Based Social Relation Reasoning

2020-07-15 · ECCV 2020 8 · Wanhua Li, Yueqi Duan, Jiwen Lu, Jianjiang Feng 외

Human beings are fundamentally sociable -- that we generally organize our social lives in terms of relations with other people. Understanding social relations from an image has great potential for intelligent systems suc…

RelationRelational ReasoningVisual Social Relationship Recognition

Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning

2025-09-09 · Houjian Yu, Zheming Zhou, Min Sun, Omid Ghasemalizadeh 외 arxiv

Enabling robots to grasp objects specified through natural language is essential for effective human-robot interaction, yet it remains a significant challenge. Existing approaches often struggle with open-form language e…

Spatial ReasoningRobotic Grasping

Obstruction reasoning for robotic grasping

2025-11-28 · Runyu Jiao, Matteo Bortolon, Francesco Giuliari, Alice Fasoli 외 arxiv

Successful robotic grasping in cluttered environments not only requires a model to visually ground a target object but also to reason about obstructions that must be cleared beforehand. While current vision-language embo…

Robotic Grasping

A Roadmap for Embodied and Social Grounding in LLMs

2024-09-25 · Sara Incao, Carlo Mazzola, Giulia Belgiovine, Alessandra Sciutti

The fusion of Large Language Models (LLMs) and robotic systems has led to a transformative paradigm in the robotic field, offering unparalleled capabilities not only in the communication domain but also in skills like mu…