paper-with-me

홈 › Papers

IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting

2025-12-10 · Tao Zhang, Yuyang Hong, Yang Xia, Kun Ding, Zeyu Zhang, Ying Wang, Shiming Xiang, Chunhong Pan arxiv

Recent advances in multimodal large language models (MLLMs) have led to impressive progress across various benchmarks. However, their capability in understanding infrared images remains unexplored. To address this gap, we introduce IF-Bench, the first high-quality benchmark designed for evaluating multimodal understanding of infrared images. IF-Bench consists of 499 images sourced from 23 infrared datasets and 680 carefully curated visual question-answer pairs, covering 10 essential dimensions of image understanding. Based on this benchmark, we systematically evaluate over 40 open-source and closed-source MLLMs, employing cyclic evaluation, bilingual assessment, and hybrid judgment strategies to enhance the reliability of the results. Our analysis reveals how model scale, architecture, and inference paradigms affect infrared image comprehension, providing valuable insights for this area. Furthermore, we propose a training-free generative visual prompting (GenViP) method, which leverages advanced image editing models to translate infrared images into semantically and spatially aligned RGB counterparts, thereby mitigating domain distribution shifts. Extensive experiments demonstrate that our method consistently yields significant performance improvements across a wide range of MLLMs. The benchmark and code are available at https://github.com/casiatao/IF-Bench.

📄 PDF Abstract BibTeX arXiv:2512.09663

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

UAVBench and UAVIT-1M: Benchmarking and Enhancing MLLMs for Low-Altitude UAV Vision-Language Understanding

2026-03-15 · Yang Zhan, Yuan Yuan arxiv

Multimodal Large Language Models (MLLMs) have made significant strides in natural images and satellite remote sensing images. However, understanding low-altitude drone scenarios remains a challenge. Existing datasets pri…

MileBench: Benchmarking MLLMs in Long Context

2024-04-29 · Dingjie Song, Shunian Chen, Guiming Hardy Chen, Fei Yu 외

Despite the advancements and impressive performance of Multimodal Large Language Models (MLLMs) on benchmarks, their effectiveness in real-world, long-context, and multi-image tasks is unclear due to the benchmarks' limi…

BenchmarkingDiagnostic

Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition

2026-01-21 · Hatef Otroshi Shahreza, Anjith George, Sébastien Marcel arxiv

Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance on a wide range of vision-language tasks, raising interest in their potential use for biometric applications. In this paper, we condu…

Heterogeneous Face Recognition

From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning

2025-09-21 · Hang Du, Jiayang Zhang, Guoshun Nan, Wendi Deng 외 arxiv

Multi-image Interleaved Reasoning aims to improve Multi-modal Large Language Models (MLLMs) ability to jointly comprehend and reason across multiple images and their associated textual contexts, introducing unique challe…

Q-Bench-Portrait: Benchmarking Multimodal Large Language Models on Portrait Image Quality Perception

2026-01-26 · Sijing Wu, Yunhao Li, Zicheng Zhang, Qi Jia 외 arxiv

Recent advances in multimodal large language models (MLLMs) have demonstrated impressive performance on existing low-level vision benchmarks, which primarily focus on generic images. However, their capabilities to percei…