paper-with-me

Papers

Revisiting MLLM Based Image Quality Assessment: Errors and Remedy

2025-11-11 · Zhenchen Tang, Songlin Yang, Bo Peng, Zichuan Wang, Jing Dong arxiv

The rapid progress of multi-modal large language models (MLLMs) has boosted the task of image quality assessment (IQA). However, a key challenge arises from the inherent mismatch between the discrete token outputs of MLLMs and the continuous nature of quality scores required by IQA tasks. This discrepancy significantly hinders the performance of MLLM-based IQA methods. Previous approaches that convert discrete token predictions into continuous scores often suffer from conversion errors. Moreover, the semantic confusion introduced by level tokens (e.g., ``good'') further constrains the performance of MLLMs on IQA tasks and degrades their original capabilities for related tasks. To tackle these problems, we provide a theoretical analysis of the errors inherent in previous approaches and, motivated by this analysis, propose a simple yet effective framework, Q-Scorer. This framework incorporates a lightweight regression module and IQA-specific score tokens into the MLLM pipeline. Extensive experiments demonstrate that Q-Scorer achieves state-of-the-art performance across multiple IQA benchmarks, generalizes well to mixed datasets, and further improves when combined with other methods.

📄 PDF Abstract BibTeX arXiv:2511.07812

Code (0)

등록된 구현이 없습니다.

Tasks

Image Quality Assessment

Similar Papers 제목 키워드 기반

Ultrasound-QBench: Can LLMs Aid in Quality Assessment of Ultrasound Imaging?

2025-01-06 · Hongyi Miao, Jun Jia, Yankun Cao, Yingjie Zhou 외

With the dramatic upsurge in the volume of ultrasound examinations, low-quality ultrasound imaging has gradually increased due to variations in operator proficiency and imaging circumstances, imposing a severe burden on …

DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment

2025-07-17 · Junjie Gao, Runze Liu, Yingzhe Peng, Shujian Yang 외

Document quality assessment is critical for a wide range of applications including document digitization, OCR, and archival. However, existing approaches often struggle to provide accurate and robust quality scores, limi…

Document Image Quality AssessmentImage Quality AssessmentOptical Character Recognition (OCR)

Revisiting Multi-Modal LLM Evaluation

2024-08-09 · Jian Lu, Shikhar Srivastava, Junyu Chen, Robik Shrestha 외

With the advent of multi-modal large language models (MLLMs), datasets used for visual question answering (VQA) and referring expression comprehension have seen a resurgence. However, the most popular datasets used to ev…

Chart UnderstandingOptical Character RecognitionQuestion AnsweringReferring Expression+3

Q-Boost: On Visual Quality Assessment Ability of Low-level Multi-Modality Foundation Models

2023-12-23 · ZiCheng Zhang, HaoNing Wu, Zhongpeng Ji, Chunyi Li 외

Recent advancements in Multi-modality Large Language Models (MLLMs) have demonstrated remarkable capabilities in complex high-level vision tasks. However, the exploration of MLLM potential in visual quality assessment, a…

Image Quality AssessmentVideo Quality AssessmentVisual Question Answering (VQA)

Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models

2025-11-14 · Jiaxi Huang, Dongxu Wu, Hanwei Zhu, Lingyu Zhu 외 arxiv

The rapid advancement of Multi-modal Large Language Models (MLLMs) has expanded their capabilities beyond high-level vision tasks. Nevertheless, their potential for Document Image Quality Assessment (DIQA) remains undere…

Image Quality Assessment