paper-with-me

홈 › Papers

Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding

2025-11-10 · Yuzhen Li, Min Liu, Zhaoyang Li, Yuan Bian, Xueping Wang, Erbo Zhai, Yaonan Wang arxiv

Monocular 3D Visual Grounding (Mono3DVG) is an emerging task that locates 3D objects in RGB images using text descriptions with geometric cues. However, existing methods face two key limitations. Firstly, they often over-rely on high-certainty keywords that explicitly identify the target object while neglecting critical spatial descriptions. Secondly, generalized textual features contain both 2D and 3D descriptive information, thereby capturing an additional dimension of details compared to singular 2D or 3D visual features. This characteristic leads to cross-dimensional interference when refining visual features under text guidance. To overcome these challenges, we propose Mono3DVG-EnSD, a novel framework that integrates two key components: the CLIP-Guided Lexical Certainty Adapter (CLIP-LCA) and the Dimension-Decoupled Module (D2M). The CLIP-LCA dynamically masks high-certainty keywords while retaining low-certainty implicit spatial descriptions, thereby forcing the model to develop a deeper understanding of spatial relationships in captions for object localization. Meanwhile, the D2M decouples dimension-specific (2D/3D) textual features from generalized textual features to guide corresponding visual features at same dimension, which mitigates cross-dimensional interference by ensuring dimensionally-consistent cross-modal interactions. Through comprehensive comparisons and ablation studies on the Mono3DRefer dataset, our method achieves state-of-the-art (SOTA) performance across all metrics. Notably, it improves the challenging Far(Acc@0.5) scenario by a significant +13.54%.

📄 PDF Abstract BibTeX arXiv:2511.06908

Code (0)

등록된 구현이 없습니다.

Tasks

Object LocalizationVisual Grounding

Similar Papers 제목 키워드 기반

On Conditional Stochastic Interpolation for Generative Nonlinear Sufficient Dimension Reduction

2025-12-22 · Shuntuo Xu, Zhou Yu, Jian Huang arxiv

Identifying low-dimensional sufficient structures in nonlinear sufficient dimension reduction (SDR) has long been a fundamental yet challenging problem. Most existing methods lack theoretical guarantees of exhaustiveness…

LensDFF: Language-enhanced Sparse Feature Distillation for Efficient Few-Shot Dexterous Manipulation

2025-03-05 · Qian Feng, David S. Martinez Lema, Jianxiang Feng, Zhaopeng Chen 외

Learning dexterous manipulation from few-shot demonstrations is a significant yet challenging problem for advanced, human-like robotic systems. Dense distilled feature fields have addressed this challenge by distilling r…

NeRFNeural Rendering

OpenSDI: Spotting Diffusion-Generated Images in the Open World

2025-03-25 · CVPR 2025 1 · Yabin Wang, Zhiwu Huang, Xiaopeng Hong

This paper identifies OpenSDI, a challenge for spotting diffusion-generated images in open-world settings. In response to this challenge, we define a new benchmark, the OpenSDI dataset (OpenSDID), which stands out from e…

OpenSD: Unified Open-Vocabulary Segmentation and Detection

2023-12-10 · Shuai Li, Minghan Li, Pengfei Wang, Lei Zhang

Recently, a few open-vocabulary methods have been proposed by employing a unified architecture to tackle generic segmentation and detection tasks. However, their performance still lags behind the task-specific models due…

DecoderPrompt LearningSegmentationZero Shot Segmentation

Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation

2026-06-23 · Tianyu Zhu, Yingping Liang, Hesong Li, Ying Fu arxiv

Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing models are typically trained on 2D image or video datasets with naive s…

Referring Video Object SegmentationZero-shot GeneralizationImage SegmentationVideo Segmentation