paper-with-me

Papers

iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA

2025-10-20 · Zhaoran Zhao, Xinli Yue, Jianhui Sun, Yuhao Xie, Tao Shao, Liangchao Yao, Fan Xia, Yuetang Deng arxiv

Image Quality Assessment (IQA) has progressed from scalar quality prediction to more interpretable, human-aligned evaluation paradigms. In this work, we address the emerging challenge of detailed and explainable IQA by proposing iDETEX-a unified multimodal large language model (MLLM) capable of simultaneously performing three key tasks: quality grounding, perception, and description. To facilitate efficient and generalizable training across these heterogeneous subtasks, we design a suite of task-specific offline augmentation modules and a data mixing strategy. These are further complemented by online enhancement strategies to fully exploit multi-sourced supervision. We validate our approach on the large-scale ViDA-UGC benchmark, where iDETEX achieves state-of-the-art performance across all subtasks. Our model ranks first in the ICCV MIPI 2025 Detailed Image Quality Assessment Challenge, demonstrating its effectiveness and robustness in delivering accurate and interpretable quality assessments.

📄 PDF Abstract BibTeX arXiv:2510.17332

Code (0)

등록된 구현이 없습니다.

Tasks

Image Quality Assessment

Similar Papers 제목 키워드 기반

Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator

2023-12-11 · Henry Hengyuan Zhao, Pan Zhou, Mike Zheng Shou

Multimodal Large Language Models (MLLMs) demonstrate exceptional problem-solving capabilities, but few research studies aim to gauge the ability to generate visual instruction tuning data. This paper proposes to explore …

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images

2025-08-18 · Wenjie Liao, Jieyu Yuan, Yifang Xu, Chunle Guo 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have introduced a paradigm shift for Image Quality Assessment (IQA) from unexplainable image quality scoring to explainable IQA, demonstrating practical applica…

Image Quality AssessmentImage Restoration

Empowering Segmentation Ability to Multi-modal Large Language Models

2024-03-21 · YuQi Yang, Peng-Tao Jiang, Jing Wang, Hao Zhang 외

Multi-modal large language models (MLLMs) can understand image-language prompts and demonstrate impressive reasoning ability. In this paper, we extend MLLMs' output by empowering MLLMs with the segmentation ability. The …

Dialogue GenerationReasoning SegmentationSegmentationWord Embeddings

GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning

2026-03-11 · Ruiheng Liu, Haihong Hao, Mingfei Han, Xin Gu 외 arxiv

Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understanding of Multimodal Large Language Model…

Multimodal ReasoningSpatial ReasoningVisual Reasoning

SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing

2026-04-25 · Zi Meng, Wanli Song, Yi Hu, Jiayuan Rao 외 arxiv

Refereeing is vital in sports, where fair, accurate, and explainable decisions are fundamental. While intelligent assistant technologies are being widely adopted in soccer refereeing, current AI-assisted approaches remai…