paper-with-me

홈 › Papers

EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception

2025-10-15 · Xijun Wang, Tanay Sharma, Achin Kulshrestha, Abhimitra Meka, Aveek Purohit, Dinesh Manocha arxiv

As AR/VR technologies become integral to daily life, there's a growing need for AI that understands human social dynamics from an egocentric perspective. However, current LLMs often lack the social awareness to discern when to intervene as AI assistant. This leads to constant, socially unaware responses that may disrupt natural conversation and negatively impact user focus. To address these limitations, we introduce EgoSocial, a large-scale egocentric dataset with 13,500 social video-question pairs, specifically designed to benchmark intervention in social interaction perception. We also present an in-depth analysis of current omnimodal LLMs (OLLMs) to assess their effectiveness in detecting diverse social contextual cues. Experiments show that OLLMs still struggle to detect the intervention timing (14.4% for Gemini 2.5 Pro). We also propose EgoSoD (EgoSocial Detection), an end-to-end method for robustly discerning social dynamics. Informed by our OLLM analysis, EgoSoD integrates multimodal contextual cues (e.g., audio and visual cues) into a social thinking graph, dynamically modeling participants and interactions. Our method proactively detects intervention timing and social interactions, precisely determining when to intervene. Our EgoSoD improves Phi-4 by 45.6% and Gemini 2.5 Pro by 9.9% on Intervention Timing performance, and improves Phi-4 by 20.4% and Gemini 2.5 Pro by 6.9% on overall Social Interaction performance. We will release the dataset and code soon.

📄 PDF Abstract BibTeX arXiv:2510.13105

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

2026-05-26 · Xudong Lu, Xueying Li, Annan Wang, Yang Bo 외 arxiv

We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline video understanding or text-prompted stre…

Mathematical Reasoning

OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answering

2026-02-03 · Yifan Zhu, Xinyu Mu, Tao Feng, Zhonghong Ou 외 arxiv

Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA still suffers from costly dense encoding…

Video Question Answering

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

2026-05-11 · Yeongtak Oh, Dongwook Lee, Sangkwon Park, Heeseung Kim 외 arxiv

While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, with unified omnimodal benchmarking that jointly covers text, image, an…

Visual Grounding

Entering Real Social World! Benchmarking the Social Intelligence of Large Language Models from a First-person Perspective

2024-10-08 · Guiyang Hou, Wenqi Zhang, Yongliang Shen, Zeqi Tan 외

Social intelligence is built upon three foundational pillars: cognitive intelligence, situational intelligence, and behavioral intelligence. As large language models (LLMs) become increasingly integrated into our social …

AttributeBenchmarkingcounterfactualNavigate

ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models

2026-03-19 · Thomas De Min, Subhankar Roy, Stéphane Lathuilière, Elisa Ricci 외 arxiv

Effective collaboration begins with knowing when to ask for help. For example, when trying to identify an occluded object, a human would ask someone to remove the obstruction. Can MLLMs exhibit a similar "proactive" beha…

Reinforcement Learning