paper-with-me

Papers

Grounding-IQA: Multimodal Language Grounding Model for Image Quality Assessment

2024-11-26 · Zheng Chen, Xun Zhang, Wenbo Li, Renjing Pei, Fenglong Song, Xiongkuo Min, Xiaohong Liu, Xin Yuan, Yong Guo, Yulun Zhang

The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based IQA methods primarily rely on general contextual descriptions, sometimes limiting fine-grained quality assessment. To address this limitation, we introduce a new image quality assessment (IQA) task paradigm, grounding-IQA. This paradigm integrates multimodal referring and grounding with IQA to realize more fine-grained quality perception. Specifically, grounding-IQA comprises two subtasks: grounding-IQA-description (GIQA-DES) and visual question answering (GIQA-VQA). GIQA-DES involves detailed descriptions with precise locations (e.g., bounding boxes), while GIQA-VQA focuses on quality QA for local regions. To realize grounding-IQA, we construct a corresponding dataset, GIQA-160K, through our proposed automated annotation pipeline. Furthermore, we develop a well-designed benchmark, GIQA-Bench. The benchmark comprehensively evaluates the model grounding-IQA performance from three perspectives: description quality, VQA accuracy, and grounding precision. Experiments demonstrate that our proposed task paradigm, dataset, and benchmark facilitate the more fine-grained IQA application. Code: https://github.com/zhengchen1999/Grounding-IQA.

📄 PDF Abstract BibTeX arXiv:2411.17237

Code (1)

zhengchen1999/grounding-iqa 공식 구현 pytorch

Tasks

Image Quality AssessmentQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Learning to Ground VLMs without Forgetting

2024-10-14 · Aritra Bhowmik, Mohammad Mahdi Derakhshani, Dennis Koelma, Martin R. Oswald 외

Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Visual Language Models (VLMs) struggle at this task. In this paper, we introduce LynX, a framew…

DecoderLanguage ModellingMixture-of-ExpertsMultimodal Reasoning+3

Visual Grounding Strategies for Text-Only Natural Language Processing

2021-03-25 · EACL (LANTERN) 2021 4 · Damien Sileo

Visual grounding is a promising path toward more robust and accurate Natural Language Processing (NLP) models. Many multimodal extensions of BERT (e.g., VideoBERT, LXMERT, VL-BERT) allow a joint modeling of texts and ima…

Image RetrievalLanguage ModelingLanguage ModellingQuestion Answering+4

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

2025-01-10 · You Li, Heyu Huang, Chi Chen, Kaiyu Huang 외

The recent advancement of Multimodal Large Language Models (MLLMs) has significantly improved their fine-grained perception of single images and general comprehension across multiple images. However, existing MLLMs still…

FormImage ComprehensionInstruction Following

Multimodal Reference Visual Grounding

2025-04-02 · Yangxiao Lu, Ruosen Li, Liqiang Jing, Jikai Wang 외

Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding performance by training large models with …

Few-Shot Object DetectionVisual Grounding

Parameter-Efficient Fine-Tuning Medical Multimodal Large Language Models for Medical Visual Grounding

2024-10-31 · Jinlong He, Pengfei Li, Gang Liu, Shenjun Zhong

Multimodal Large Language Models (MLLMs) inherit the superior text understanding capabilities of LLMs and extend these capabilities to multimodal scenarios. These models achieve excellent results in the general domain of…

parameter-efficient fine-tuningVisual Grounding