paper-with-me

Papers

A Multimodal Depth-Aware Method For Embodied Reference Understanding

2025-10-09 · Fevziye Irem Eyiokur, Dogucan Yaman, Hazım Kemal Ekenel, Alexander Waibel arxiv

Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in ambiguous scenarios where multiple candidate objects exist in the scene. To address these challenges, we propose a novel ERU framework that jointly leverages LLM-based data augmentation, depth-map modality, and a depth-aware decision module. This design enables robust integration of linguistic and embodied cues, improving disambiguation in complex or cluttered environments. Experimental results on two datasets demonstrate that our approach significantly outperforms existing baselines, achieving more accurate and reliable referent detection.

📄 PDF Abstract BibTeX arXiv:2510.08278

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationObject Detection

Similar Papers 제목 키워드 기반

YouRefIt: Embodied Reference Understanding with Language and Gesture

2021-09-08 · ICCV 2021 10 · Yixin Chen, Qing Li, Deqian Kong, Yik Lun Kei 외

We study the understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding mul…

Embodied Multimodal Agents to Bridge the Understanding Gap

2021-04-01 · EACL (HCINLP) 2021 4 · Nikhil Krishnaswamy, Nada Alalyani

In this paper we argue that embodied multimodal agents, i.e., avatars, can play an important role in moving natural language processing toward “deep understanding.” Fully-featured interactive agents, model encounters bet…

CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding

2025-07-29 · Fevziye Irem Eyiokur, Dogucan Yaman, Hazım Kemal Ekenel, Alexander Waibel arxiv

We address Embodied Reference Understanding, the task of predicting the object a person in the scene refers to through pointing gesture and language. This requires multimodal reasoning over text, visual pointing cues, an…

Multimodal Reasoning

PhysBrain 1.0 Technical Report

2026-05-14 · Shijie Lian, Bin Yu, Xiaopeng Lin, Changti Wu 외 arxiv

Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a complementary route: converting large-scale hu…

FlyAwareV2: A Multimodal Cross-Domain UAV Dataset for Urban Scene Understanding

2025-10-15 · Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh arxiv

The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting…

Monocular Depth EstimationSemantic SegmentationScene UnderstandingDomain Adaptation