paper-with-me

Papers

GROUNDHOG: Grounding Large Language Models to Holistic Segmentation

2024-02-26 · CVPR 2024 1 · Yichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah, Qiaozi Gao, Joyce Chai

Most multimodal large language models (MLLMs) learn language-to-object grounding through causal language modeling where grounded objects are captured by bounding boxes as sequences of location tokens. This paradigm lacks pixel-level representations that are important for fine-grained visual understanding and diagnosis. In this work, we introduce GROUNDHOG, an MLLM developed by grounding Large Language Models to holistic segmentation. GROUNDHOG incorporates a masked feature extractor and converts extracted features into visual entity tokens for the MLLM backbone, which then connects groundable phrases to unified grounding masks by retrieving and merging the entity masks. To train GROUNDHOG, we carefully curated M3G2, a grounded visual instruction tuning dataset with Multi-Modal Multi-Grained Grounding, by harvesting a collection of segmentation-grounded datasets with rich annotations. Our experimental results show that GROUNDHOG achieves superior performance on various language grounding tasks without task-specific fine-tuning, and significantly reduces object hallucination. GROUNDHOG also demonstrates better grounding towards complex forms of visual input and provides easy-to-understand diagnosis in failure cases.

📄 PDF Abstract BibTeX arXiv:2402.16846

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Language ModelingGeneralized Referring Expression SegmentationHallucinationLanguage ModelingLanguage ModellingObject HallucinationReferring Expression Segmentation

Similar Papers 제목 키워드 기반

ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation

2024-12-12 · CVPR 2025 1 · Ali Athar, Xueqing Deng, Liang-Chieh Chen

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body…

Phrase GroundingQuestion AnsweringSegmentationSemantic Segmentation+2

GLaMM: Pixel Grounding Large Multimodal Model

2023-11-06 · CVPR 2024 1 · Hanoona Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman Shaker 외

Large Multimodal Models (LMMs) extend Large Language Models to the vision domain. Initial LMMs used holistic images and text prompts to generate ungrounded textual responses. Recently, region-level LMMs have been used to…

Conversational Question AnsweringImage CaptioningmodelReferring Expression+4

InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual Referring

2021-03-01 · ICCV 2021 10 · Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang 외

Compared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer, to achieve a superior 3…

3D visual groundingAttributeObject LocalizationPanoptic Segmentation+1

Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs

2026-05-13 · Jincai Huang, Shihao Zou, Yuchen Guo, Jingjing Li 외 arxiv

Surgical scene understanding is a cornerstone of computer-assisted intervention. While recent advances, particularly in surgical image segmentation, have driven progress, real-world clinical applications require a more h…

Scene UnderstandingImage SegmentationVisual Grounding

Groundhog DAG: Representing Semantic Repetition in Literary Narratives

2013-06-01 · WS 2013 6 · Greg Lessard, Michael Levison