paper-with-me

Papers

2AFC Prompting of Large Multimodal Models for Image Quality Assessment

2024-02-02 · Hanwei Zhu, Xiangjie Sui, Baoliang Chen, Xuelin Liu, Peilin Chen, Yuming Fang, Shiqi Wang

While abundant research has been conducted on improving high-level visual understanding and reasoning capabilities of large multimodal models~(LMMs), their visual quality assessment~(IQA) ability has been relatively under-explored. Here we take initial steps towards this goal by employing the two-alternative forced choice~(2AFC) prompting, as 2AFC is widely regarded as the most reliable way of collecting human opinions of visual quality. Subsequently, the global quality score of each image estimated by a particular LMM can be efficiently aggregated using the maximum a posterior estimation. Meanwhile, we introduce three evaluation criteria: consistency, accuracy, and correlation, to provide comprehensive quantifications and deeper insights into the IQA capability of five LMMs. Extensive experiments show that existing LMMs exhibit remarkable IQA ability on coarse-grained quality comparison, but there is room for improvement on fine-grained quality discrimination. The proposed dataset sheds light on the future development of IQA models based on LMMs. The codes will be made publicly available at https://github.com/h4nwei/2AFC-LMMs.

📄 PDF Abstract BibTeX arXiv:2402.01162

Code (1)

h4nwei/2afc-lmms 공식 구현

Tasks

Image Quality Assessment

Similar Papers 제목 키워드 기반

A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment

2024-03-16 · Tianhe Wu, Kede Ma, Jie Liang, Yujiu Yang 외

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Ima…

Image Quality Assessment

Redefining Quality Criteria and Distance-Aware Score Modeling for Image Editing Assessment

2026-04-14 · Xinjie Zhang, Qiang Li, Xiaowen Ma, Axi Niu 외 arxiv

Recent advances in image editing have heightened the need for reliable Image Editing Quality Assessment (IEQA). Unlike traditional methods, IEQA requires complex reasoning over multimodal inputs and multi-dimensional ass…

Distance regressionImage Editing

Towards Synthesizing Normative Data for Cognitive Assessments Using Generative Multimodal Large Language Models

2025-08-25 · Victoria Yan, Honor Chotkowski, Fengran Wang, Xinhui Li 외 arxiv

Cognitive assessments require normative data as essential benchmarks for evaluating individual performance. Hence, developing new cognitive tests based on novel image stimuli is challenging due to the lack of readily ava…

Exploring Diagnostic Prompting Approach for Multimodal LLM-based Visual Complexity Assessment: A Case Study of Amazon Search Result Pages

2025-11-26 · Divendar Murtadak, Yoon Kim, Trilokya Akula arxiv

This study investigates whether diagnostic prompting can improve Multimodal Large Language Model (MLLM) reliability for visual complexity assessment of Amazon Search Results Pages (SRP). We compare diagnostic prompting w…

From Drone Imagery to Livability Mapping: AI-powered Environment Perception in Rural China

2025-08-29 · Weihuan Deng, Yaofu Huang, Luan Chen, Xun Li 외 arxiv

The high cost of acquiring rural street view images has constrained comprehensive environmental perception in rural areas. Drone photographs, with their advantages of easy acquisition, broad coverage, and high spatial re…

Computational Efficiency