paper-with-me

Papers

Grounding Beyond Detection: Enhancing Contextual Understanding in Embodied 3D Grounding

2025-06-05 · Yani Zhang, Dongming Wu, Hao Shi, Yingfei Liu, Tiancai Wang, Haoqiang Fan, Xingping Dong

Embodied 3D grounding aims to localize target objects described in human instructions from ego-centric viewpoint. Most methods typically follow a two-stage paradigm where a trained 3D detector's optimized backbone parameters are used to initialize a grounding model. In this study, we explore a fundamental question: Does embodied 3D grounding benefit enough from detection? To answer this question, we assess the grounding performance of detection models using predicted boxes filtered by the target category. Surprisingly, these detection models without any instruction-specific training outperform the grounding models explicitly trained with language instructions. This indicates that even category-level embodied 3D grounding may not be well resolved, let alone more fine-grained context-aware grounding. Motivated by this finding, we propose DEGround, which shares DETR queries as object representation for both DEtection and Grounding and enables the grounding to benefit from basic category classification and box detection. Based on this framework, we further introduce a regional activation grounding module that highlights instruction-related regions and a query-wise modulation module that incorporates sentence-level semantic into the query representation, strengthening the context-aware understanding of language instructions. Remarkably, DEGround outperforms state-of-the-art model BIP3D by 7.52% at overall accuracy on the EmbodiedScan validation set. The source code will be publicly available at https://github.com/zyn213/DEGround.

📄 PDF Abstract BibTeX arXiv:2506.05199

Code (1)

zyn213/deground 공식 구현

Similar Papers 제목 키워드 기반

Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs

2024-11-28 · Anirudh Phukan, Divyansh, Harshit Kumar Morj, Vaishnavi 외

The rapid development of Large Multimodal Models (LMMs) has significantly advanced multimodal understanding by harnessing the language abilities of Large Language Models (LLMs) and integrating modality-specific encoders.…

AttributeHallucinationOptical Character Recognition (OCR)Question Answering+2

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

2025-08-23 · Leilei Guo, Antonio Carlos Rivera, Peiyu Tang, Haoxuan Ren 외 arxiv

Large Language Models (LLMs) and Vision-Language Large Models (LVLMs) have achieved remarkable progress in natural language processing and multimodal understanding. Despite their impressive generalization capabilities, c…

Referring ExpressionVisual Reasoning

Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding

2025-01-28 · Akash Kumar, Zsolt Kira, Yogesh Singh Rawat

In this work, we focus on Weakly Supervised Spatio-Temporal Video Grounding (WSTVG). It is a multimodal task aimed at localizing specific subjects spatio-temporally based on textual queries without bounding box supervisi…

object-detectionObject DetectionPhrase GroundingScene Understanding+2

Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations

2025-09-16 · Jinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu 외 arxiv

The detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progress, they suffer from misalignment artifa…

Domain-Specific Lexical Grounding in Noisy Visual-Textual Documents

2020-10-30 · EMNLP 2020 11 · Gregory Yauney, Jack Hessel, David Mimno

Images can give us insights into the contextual meanings of words, but current image-text grounding approaches require detailed annotations. Such granular annotation is rare, expensive, and unavailable in most domain-spe…

Clusteringobject-detectionObject DetectionSentence