paper-with-me

홈 › Papers

Toward Interactive Regional Understanding in Vision-Large Language Models

2024-03-27 · Jungbeom Lee, Sanghyuk Chun, Sangdoo Yun

Recent Vision-Language Pre-training (VLP) models have demonstrated significant advancements. Nevertheless, these models heavily rely on image-text pairs that capture only coarse and global information of an image, leading to a limitation in their regional understanding ability. In this work, we introduce \textbf{RegionVLM}, equipped with explicit regional modeling capabilities, allowing them to understand user-indicated image regions. To achieve this, we design a simple yet innovative architecture, requiring no modifications to the model architecture or objective function. Additionally, we leverage a dataset that contains a novel source of information, namely Localized Narratives, which has been overlooked in previous VLP research. Our experiments demonstrate that our single generalist model not only achieves an interactive dialogue system but also exhibits superior performance on various zero-shot region understanding tasks, without compromising its ability for global image understanding.

📄 PDF Abstract BibTeX arXiv:2403.18260

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MedP-CLIP: Medical CLIP with Region-Aware Prompt Integration

2026-04-13 · Jiahui Peng, He Yao, Jingwen Li, Yanzhou Su 외 arxiv

Contrastive Language-Image Pre-training (CLIP) has demonstrated outstanding performance in global image understanding and zero-shot transfer through large-scale text-image alignment. However, the core of medical image an…

Interactive Segmentation

RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding

2023-04-03 · CVPR 2024 1 · Jihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang 외

We propose a lightweight and scalable Regional Point-Language Contrastive learning framework, namely \textbf{RegionPLC}, for open-world 3D scene understanding, aiming to identify and recognize open-set objects and catego…

Contrastive LearningInstance SegmentationScene UnderstandingSemantic Segmentation

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

2026-05-18 · Qianhao Yuan, Jie Lou, Xing Yu, Hongyu Lin 외 arxiv

Multimodal Large Language Models (MLLMs) still struggle with fine-grained visual understanding, where answers often depend on small but decisive evidence in the full image. We observe a regional-to-global perception gap:…

RegionGPT: Towards Region Understanding Vision Language Model

2024-03-04 · CVPR 2024 1 · Qiushan Guo, Shalini De Mello, Hongxu Yin, Wonmin Byeon 외

Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs, yet they struggle with detailed regional visual understanding due to limite…

Language ModelingLanguage Modellingmodel

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

2026-04-13 · Samuel Cahyawijaya, Peerat Limkonchotiwat, Tack Hwa Wong, Hitesh Laxmichand Patel 외 arxiv

While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and domains, there is still no dedicated framework for assessing human-centr…