paper-with-me

홈 › Papers

Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views

2025-10-26 · Anna Deichler, Jonas Beskow arxiv

We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and stationary cameras, we recorded synchronized gaze, speech, and video as 25 participants instructed a partner to identify ingredients in a kitchen. Combined with 3D scene reconstructions, this setup provides a benchmark for evaluating how different spatial representations (2D vs. 3D; ego vs. exo) affect multimodal grounding. The dataset contains 3.67 hours of recordings, including 2,707 richly annotated referential expressions, and is designed to advance the development of embodied agents that can understand and engage in situated dialogue.

📄 PDF Abstract BibTeX arXiv:2510.22672

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Syntax-Guided Transformers: Elevating Compositional Generalization and Grounding in Multimodal Environments

2023-11-07 · Danial Kamali, Parisa Kordjamshidi

Compositional generalization, the ability of intelligent models to extrapolate understanding of components to novel compositions, is a fundamental yet challenging facet in AI research, especially within multimodal enviro…

Compositional Generalization (AVG)Dependency Parsing

CodeMMR: Bridging Natural Language, Code, and Image for Unified Retrieval

2026-04-17 · Jiahui Geng, Qing Li, Fengyu Cai, Fakhri Karray arxiv

Code search, framed as information retrieval (IR), underpins modern software engineering and increasingly powers retrieval-augmented generation (RAG), improving code discovery, reuse, and the reliability of LLM-based cod…

Information RetrievalVisual GroundingCode GenerationCode Search

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

2026-08-04 · Zihan Wang, Tong Liu, Zhiwei Wang, Tao Huang 외 arxiv

Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general-purpose grounding models…

Visual Grounding

Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational Agents

2024-07-01 · Mehdi Arjmand, Farnaz Nouraei, Ian Steenstra, Timothy Bickmore

We introduce the concept of "empathic grounding" in conversational agents as an extension of Clark's conceptualization of grounding in conversation in which the grounding criterion includes listener empathy for the speak…

Emotional IntelligenceEmotion ClassificationHuman Interaction RecognitionLanguage Modelling+4

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

2026-05-21 · Deshui Miao, Xingsen Huang, Yameng Gu, Xin Li 외 arxiv

Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearance. Recent multimodal models incorporate geometric information throu…

Spatial Reasoning