paper-with-me

홈 › Papers

How Well Can Vision Language Models See Image Details?

2024-08-07 · Chenhui Gou, Abdulwahab Felemban, Faizan Farooq Khan, Deyao Zhu, Jianfei Cai, Hamid Rezatofighi, Mohamed Elhoseiny

Large Language Model-based Vision-Language Models (LLM-based VLMs) have demonstrated impressive results in various vision-language understanding tasks. However, how well these VLMs can see image detail beyond the semantic level remains unclear. In our study, we introduce a pixel value prediction task (PVP) to explore "How Well Can Vision Language Models See Image Details?" and to assist VLMs in perceiving more details. Typically, these models comprise a frozen CLIP visual encoder, a large language model, and a connecting module. After fine-tuning VLMs on the PVP task, we find: 1) existing VLMs struggle to predict precise pixel values by only fine-tuning the connection module and LLM; and 2) prediction precision is significantly improved when the vision encoder is also adapted. Additionally, our research reveals that incorporating pixel value prediction as one of the VLM pre-training tasks and vision encoder adaptation markedly boosts VLM performance on downstream image-language understanding tasks requiring detailed image perception, such as referring image segmentation (with an average +10.19 cIoU improvement) and video game decision making (with average score improvements of +80.34 and +70.54 on two games, respectively).

📄 PDF Abstract BibTeX arXiv:2408.03940

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingImage SegmentationLanguage ModelingLanguage ModellingLarge Language ModelPredictionSemantic SegmentationValue prediction

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models

2024-11-22 · CVPR 2025 1 · Junzhe Chen, Tianshu Zhang, Shiyu Huang, Yuwei Niu 외

Despite the recent breakthroughs achieved by Large Vision Language Models (LVLMs) in understanding and responding to complex visual-textual contexts, their inherent hallucination tendencies limit their practical applicat…

HallucinationObjectObject Hallucination

xT: Nested Tokenization for Larger Context in Large Images

2024-03-04 · Ritwik Gupta, Shufan Li, Tyler Zhu, Jitendra Malik 외

Modern computer vision pipelines handle large images in one of two sub-optimal ways: down-sampling or cropping. These two methods incur significant losses in the amount of information and context present in an image. The…

FuseCap: Leveraging Large Language Models for Enriched Fused Image Captions

2023-05-28 · Noam Rotstein, David Bensaid, Shaked Brody, Roy Ganz 외

The advent of vision-language pre-training techniques enhanced substantial progress in the development of models for image captioning. However, these models frequently produce generic captions and may omit semantically i…

AttributeImage CaptioningLanguage ModellingLarge Language Model+3

Pixel Perfect MegaMed: A Megapixel-Scale Vision-Language Foundation Model for Generating High Resolution Medical Images

2025-07-17 · Zahra Tehraninasab, Amar Kumar, Tal Arbel

Medical image synthesis presents unique challenges due to the inherent complexity and high-resolution details required in clinical contexts. Traditional generative architectures such as Generative Adversarial Networks (G…

Data AugmentationImage GenerationMedical Image Generation

Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?

2024-06-18 · Mingqian Feng, Yunlong Tang, Zeliang Zhang, Chenliang Xu

Large Vision-Language Models (LVLMs) excel in integrating visual and linguistic contexts to produce detailed content, facilitating applications such as image captioning. However, using LVLMs to generate descriptions ofte…

AttributeHallucinationImage CaptioningObject Hallucination