paper-with-me

홈 › Papers

RegionGPT: Towards Region Understanding Vision Language Model

2024-03-04 · CVPR 2024 1 · Qiushan Guo, Shalini De Mello, Hongxu Yin, Wonmin Byeon, Ka Chun Cheung, Yizhou Yu, Ping Luo, Sifei Liu

Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs, yet they struggle with detailed regional visual understanding due to limited spatial awareness of the vision encoder, and the use of coarse-grained training data that lacks detailed, region-specific captions. To address this, we introduce RegionGPT (short as RGPT), a novel framework designed for complex region-level captioning and understanding. RGPT enhances the spatial awareness of regional representation with simple yet effective modifications to existing visual encoders in VLMs. We further improve performance on tasks requiring a specific output scope by integrating task-guided instruction prompts during both training and inference phases, while maintaining the model's versatility for general-purpose tasks. Additionally, we develop an automated region caption data generation pipeline, enriching the training set with detailed region-level captions. We demonstrate that a universal RGPT model can be effectively applied and significantly enhancing performance across a range of region-level tasks, including but not limited to complex region descriptions, reasoning, object classification, and referring expressions comprehension.

📄 PDF Abstract BibTeX arXiv:2403.02330

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingmodel

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

SLAN: Self-Locator Aided Network for Vision-Language Understanding

2023-01-01 · ICCV 2023 1 · Jiang-Tian Zhai, Qi Zhang, Tong Wu, Xing-Yu Chen 외

Learning fine-grained interplay between vision and language contributes to a more accurate understanding for Vision-Language tasks. However, it remains challenging to extract key image regions according to the texts …

Image RetrievalImage to textRetrieval

Toward Interactive Regional Understanding in Vision-Large Language Models

2024-03-27 · Jungbeom Lee, Sanghyuk Chun, Sangdoo Yun

Recent Vision-Language Pre-training (VLP) models have demonstrated significant advancements. Nevertheless, these models heavily rely on image-text pairs that capture only coarse and global information of an image, leadin…

RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding

2023-04-03 · CVPR 2024 1 · Jihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang 외

We propose a lightweight and scalable Regional Point-Language Contrastive learning framework, namely \textbf{RegionPLC}, for open-world 3D scene understanding, aiming to identify and recognize open-set objects and catego…

Contrastive LearningInstance SegmentationScene UnderstandingSemantic Segmentation

TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

2026-02-23 · Fan Yang, Shurong Zheng, Hongyin Zhao, Yufei Zhan 외 arxiv

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, strug…

Trajectory PredictionScene UnderstandingLogical Reasoning

SLAN: Self-Locator Aided Network for Cross-Modal Understanding

2022-11-28 · Jiang-Tian Zhai, Qi Zhang, Tong Wu, Xing-Yu Chen 외

Learning fine-grained interplay between vision and language allows to a more accurate understanding for VisionLanguage tasks. However, it remains challenging to extract key image regions according to the texts for semant…

Image RetrievalImage to textRetrieval