paper-with-me

Papers

DSM: Building A Diverse Semantic Map for 3D Visual Grounding

2025-04-11 · Qinghongbing Xie, Zijian Liang, Long Zeng

In recent years, with the growing research and application of multimodal large language models (VLMs) in robotics, there has been an increasing trend of utilizing VLMs for robotic scene understanding tasks. Existing approaches that use VLMs for 3D Visual Grounding tasks often focus on obtaining scene information through geometric and visual information, overlooking the extraction of diverse semantic information from the scene and the understanding of rich implicit semantic attributes, such as appearance, physics, and affordance. The 3D scene graph, which combines geometry and language, is an ideal representation method for environmental perception and is an effective carrier for language models in 3D Visual Grounding tasks. To address these issues, we propose a diverse semantic map construction method specifically designed for robotic agents performing 3D Visual Grounding tasks. This method leverages VLMs to capture the latent semantic attributes and relations of objects within the scene and creates a Diverse Semantic Map (DSM) through a geometry sliding-window map construction strategy. We enhance the understanding of grounding information based on DSM and introduce a novel approach named DSM-Grounding. Experimental results show that our method outperforms current approaches in tasks like semantic segmentation and 3D Visual Grounding, particularly excelling in overall metrics compared to the state-of-the-art. In addition, we have deployed this method on robots to validate its effectiveness in navigation and grasping tasks.

📄 PDF Abstract BibTeX arXiv:2504.08307

Code (0)

등록된 구현이 없습니다.

Tasks

3D visual groundingScene UnderstandingSemantic SegmentationVisual Grounding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

VisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought

2026-03-12 · Eunsoo Lee, Jeongwoo Lee, Minki Hong, Jangho Choi 외 arxiv

Large vision-language models (LVLMs) struggle to reliably detect visual primitives in charts and align them with semantic representations, which severely limits their performance on complex visual reasoning. This lack of…

Visual Question AnsweringVisual GroundingVisual Reasoning

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

2025-05-08 · Henry Zheng, Hao Shi, Qihang Peng, Yong Xien Chng 외

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D vi…

3D visual groundingcross-modal alignmentVisual Grounding

StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles

2026-02-25 · Daniel Oliveira, David Martins de Matos arxiv

Visual storytelling models that correctly ground entities in images may still hallucinate semantic relationships, generating incorrect dialogue attribution, character interactions, or emotional states. We introduce Story…

Visual StorytellingVisual Grounding

Visual Word2Vec (vis-w2v): Learning Visually Grounded Word Embeddings Using Abstract Scenes

2015-11-22 · CVPR 2016 6 · Satwik Kottur, Ramakrishna Vedantam, José M. F. Moura, Devi Parikh

We propose a model to learn visually grounded word embeddings (vis-w2v) to capture visual notions of semantic relatedness. While word embeddings trained using text have been extremely successful, they cannot uncover noti…

Common Sense ReasoningImage RetrievalRetrievalVisual Grounding+1

Learning Cross-modal Context Graph for Visual Grounding

2020-02-13 · AAAI-2020 2020 2 · Yongfei Liu; Bo Wan; Xiaodan Zhu; Xuming He

Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the res…

Graph MatchingGraph Neural NetworkLanguage ModellingNatural Language Visual Grounding+2