paper-with-me

Papers

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

2025-05-08 · Henry Zheng, Hao Shi, Qihang Peng, Yong Xien Chng, Rui Huang, Yepeng Weng, Zhongchao shi, Gao Huang

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D visual grounding, where agents locate target objects in real-world 3D spaces based on verbal descriptions. However, this task faces two significant challenges: (1) loss of fine-grained visual semantics due to sparse fusion of point clouds with ego-centric multi-view images, (2) limited textual semantic context due to arbitrary language descriptions. We propose DenseGrounding, a novel approach designed to address these issues by enhancing both visual and textual semantics. For visual features, we introduce the Hierarchical Scene Semantic Enhancer, which retains dense semantics by capturing fine-grained global scene features and facilitating cross-modal alignment. For text descriptions, we propose a Language Semantic Enhancer that leverages large language models to provide rich context and diverse language descriptions with additional context during model training. Extensive experiments show that DenseGrounding significantly outperforms existing methods in overall accuracy, with improvements of 5.81% and 7.56% when trained on the comprehensive full dataset and smaller mini subset, respectively, further advancing the SOTA in egocentric 3D visual grounding. Our method also achieves 1st place and receives the Innovation Award in the CVPR 2024 Autonomous Grand Challenge Multi-view 3D Visual Grounding Track, validating its effectiveness and robustness.

📄 PDF Abstract BibTeX arXiv:2505.04965

Code (0)

등록된 구현이 없습니다.

Tasks

3D visual groundingcross-modal alignmentVisual Grounding

Similar Papers 제목 키워드 기반

Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language Tracking

2023-11-28 · Jiawei Ge, Xiangmei Chen, Jiuxin Cao, Xuelin Zhu 외

Single object tracking aims to locate one specific target in video sequences, given its initial state. Classical trackers rely solely on visual cues, restricting their ability to handle challenges such as appearance vari…

Object TrackingRepresentation Learning

Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model

2025-09-09 · Zhuoxu Huang, Mingqi Gao, Jungong Han arxiv

3D object segmentation with Large Language Models (LLMs) has become a prevailing paradigm due to its broad semantics, task flexibility, and strong generalization. However, this paradigm is hindered by representation misa…

Object SegmentationPoint Clouds

Reasoning-Augmented Representations for Multimodal Retrieval

2026-02-06 · Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han, Soochahn Lee 외 arxiv

Universal Multimodal Retrieval (UMR) seeks any-to-any search across text and vision, yet modern embedding models remain brittle when queries require latent reasoning (e.g., resolving underspecified references or matching…

SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling

2025-01-01 · CVPR 2025 1 · Qi Zhu, Jiangwei Lao, Deyi Ji, Junwei Luo 외

Open-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancemen…

Language ModelingLanguage Modelling

Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos

2023-08-19 · ICCV 2023 1 · Rui Qian, Shuangrui Ding, Xian Liu, Dahua Lin

Self-supervised methods have shown remarkable progress in learning high-level semantics and low-level temporal correspondence. Building on these results, we take one step further and explore the possibility of integratin…

ObjectObject DiscoverySemantic Segmentation