paper-with-me

Papers

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

2026-05-20 · Anna Deichler, Jim O'Regan, Fethiye Irmak Dogan, Lubos Marcinek, Anna Klezovich, Iolanda Leite, Jonas Beskow arxiv

Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to resolve ambiguous expressions in spontaneous, multi-turn dialogue. We address this gap by introducing (1) a benchmark for referential communication in dynamic 3D environments, built from 6.7 hours of egocentric VR interaction with synchronized speech, motion, gaze, and 3D scene geometry, and (2) a two-stage grounding pipeline that explicitly resolves conversational ambiguity before visual localization. The benchmark includes over 4,200 manually verified referring expressions spanning full, partitive, and pronominal types. Our contextual rewriting approach improves grounding performance by 11-22 percentage points on average, with a pure detector (GroundingDINO) reaching 56.7% on pronominals after rewriting, nearly double the best end-to-end baseline. Results demonstrate that decoupling linguistic reasoning from visual perception is more effective than end-to-end approaches for conversational grounding.

📄 PDF Abstract BibTeX arXiv:2605.21796

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Localization

Similar Papers 제목 키워드 기반

Region-Level Context-Aware Multimodal Understanding

2025-08-17 · Hongliang Wei, Xianqi Zhang, Xingtao Wang, Xiaopeng Fan 외 arxiv

Despite significant progress, existing research on Multimodal Large Language Models (MLLMs) mainly focuses on general visual understanding, overlooking the ability to integrate textual context associated with objects for…

SURE: Synergistic Uncertainty-aware Reasoning for Multimodal Emotion Recognition in Conversations

2026-04-02 · Yiqiang Cai, Chengyan Wu, Bolei Ma, Bo Chen 외 arxiv

Multimodal emotion recognition in conversations (MERC) requires integrating multimodal signals while being robust to noise and modeling contextual reasoning. Existing approaches often emphasize fusion but overlook uncert…

Multimodal Emotion RecognitionMultimodal Reasoning

CaMML: Context-Aware Multimodal Learner for Large Models

2024-01-06 · Yixin Chen, Shuai Zhang, Boran Han, Tong He 외

In this work, we introduce Context-Aware MultiModal Learner (CaMML), for tuning large multimodal models (LMMs). CaMML, a lightweight module, is crafted to seamlessly integrate multimodal contextual samples into large mod…

Visual Question Answering

MCDubber: Multimodal Context-Aware Expressive Video Dubbing

2024-08-21 · Yuan Zhao, Zhenqi Jia, Rui Liu, De Hu 외

Automatic Video Dubbing (AVD) aims to take the given script and generate speech that aligns with lip motion and prosody expressiveness. Current AVD models mainly utilize visual information of the current sentence to enha…

Sentence

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

2025-05-30 · Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky 외

Cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey sufficient context to capture region-specific meani…

BenchmarkingMachine TranslationMultimodal Machine TranslationTranslation