paper-with-me

홈 › Papers

A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment

2024-03-16 · Tianhe Wu, Kede Ma, Jie Liang, Yujiu Yang, Lei Zhang

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality Assessment (IQA) remains largely unexplored. In this paper, we conduct a comprehensive and systematic study of prompting MLLMs for IQA. We first investigate nine prompting systems for MLLMs as the combinations of three standardized testing procedures in psychophysics (i.e., the single-stimulus, double-stimulus, and multiple-stimulus methods) and three popular prompting strategies in natural language processing (i.e., the standard, in-context, and chain-of-thought prompting). We then present a difficult sample selection procedure, taking into account sample diversity and uncertainty, to further challenge MLLMs equipped with the respective optimal prompting systems. We assess three open-source and one closed-source MLLMs on several visual attributes of image quality (e.g., structural and textural distortions, geometric transformations, and color differences) in both full-reference and no-reference scenarios. Experimental results show that only the closed-source GPT-4V provides a reasonable account for human perception of image quality, but is weak at discriminating fine-grained quality variations (e.g., color differences) and at comparing visual quality of multiple images, tasks humans can perform effortlessly.

📄 PDF Abstract BibTeX arXiv:2403.10854

Code (2)

tianhewu/mllms-for-iqa 공식 구현 pytorch
tianhewu/visualquality-r1 pytorch

Tasks

Image Quality Assessment

Similar Papers 제목 키워드 기반

Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency

2024-06-19 · Sakib Shahriar, Brady Lund, Nishith Reddy Mannuru, Muhammad Arbab Arshad 외

As large language models (LLMs) continue to advance, evaluating their comprehensive capabilities becomes significant for their application in various fields. This research study comprehensively evaluates the language, vi…

Few-Shot Learningimage-classificationImage ClassificationObject Recognition

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

2024-03-14 · Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge 외

In this work, we discuss building performant Multimodal Large Language Models (MLLMs). In particular, we study the importance of various architecture components and data choices. Through careful and comprehensive ablatio…

In-Context LearningMixture-of-ExpertsVisual Question Answering

GRAPHGPT-O: Synergistic Multimodal Comprehension and Generation on Graphs

2025-02-17 · CVPR 2025 1 · Yi Fang, Bowen Jin, Jiacheng Shen, Sirui Ding 외

The rapid development of Multimodal Large Language Models (MLLMs) has enabled the integration of multiple modalities, including texts and images, within the large language model (LLM) framework. However, texts and images…

Image GenerationLanguage ModelingLanguage ModellingLarge Language Model

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

2023-05-13 · Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang 외

Large models have recently played a dominant role in natural language processing and multimodal vision-language learning. However, their effectiveness in text-related visual tasks remains relatively unexplored. In this p…

Key Information ExtractionNutritionOptical Character RecognitionOptical Character Recognition (OCR)+3

What Makes Multimodal In-Context Learning Work?

2024-04-24 · Folco Bertini Baldassini, Mustafa Shukor, Matthieu Cord, Laure Soulier 외

Large Language Models have demonstrated remarkable performance across various tasks, exhibiting the capacity to swiftly acquire new skills, such as through In-Context Learning (ICL) with minimal demonstration examples. I…

In-Context Learning