paper-with-me

홈 › Papers

From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning

2025-09-21 · Hang Du, Jiayang Zhang, Guoshun Nan, Wendi Deng, Zhenyan Chen, Chenyang Zhang, Wang Xiao, Shan Huang, Yuqi Pan, Tao Qi, Sicong Leng arxiv

Multi-image Interleaved Reasoning aims to improve Multi-modal Large Language Models (MLLMs) ability to jointly comprehend and reason across multiple images and their associated textual contexts, introducing unique challenges beyond single-image or non-interleaved multi-image tasks. While current multi-image benchmarks overlook interleaved textual contexts and neglect distinct relationships between individual images and their associated texts, enabling models to reason over multi-image interleaved data may significantly enhance their comprehension of complex scenes and better capture cross-modal correlations. To bridge this gap, we introduce a novel benchmark MIR, requiring joint reasoning over multiple images accompanied by interleaved textual contexts to accurately associate image regions with corresponding texts and logically connect information across images. To enhance MLLMs ability to comprehend multi-image interleaved data, we introduce reasoning steps for each instance within the benchmark and propose a stage-wise curriculum learning strategy. This strategy follows an "easy to hard" approach, progressively guiding models from simple to complex scenarios, thereby enhancing their ability to handle challenging tasks. Extensive experiments benchmarking multiple MLLMs demonstrate that our method significantly enhances models reasoning performance on MIR and other established benchmarks. We believe that MIR will encourage further research into multi-image interleaved reasoning, facilitating advancements in MLLMs capability to handle complex inter-modal tasks.

📄 PDF Abstract BibTeX arXiv:2509.17040

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IRIS: Interleaved Reinforcement with Incremental Staged Curriculum for Cross-Lingual Mathematical Reasoning

2026-04-27 · Navya Gupta, Rishitej Reddy Vyalla, Avinash Anand, Chhavi Kirtani 외 arxiv

Curriculum learning helps language models tackle complex reasoning by gradually increasing task difficulty. However, it often fails to generate consistent step-by-step reasoning, especially in multilingual and low-resour…

Reinforcement LearningCross-Lingual TransferMathematical Reasoning

M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?

2025-03-27 · Haolong Yan, Kaijun Tan, Yeqing Shen, Xin Huang 외

We investigate a critical yet under-explored question in Large Vision-Language Models (LVLMs): Do LVLMs genuinely comprehend interleaved image-text in the document? Existing document understanding benchmarks often assess…

Document Summarizationdocument understanding

From Easy to Hard: Two-stage Selector and Reader for Multi-hop Question Answering

2022-05-24 · Xin-Yi Li, Wei-Jun Lei, Yu-Bin Yang

Multi-hop question answering (QA) is a challenging task requiring QA systems to perform complex reasoning over multiple documents and provide supporting facts together with the exact answer. Existing works tend to utiliz…

Multi-hop Question AnsweringQuestion Answering

Texture-guided Saliency Distilling for Unsupervised Salient Object Detection

2022-07-13 · CVPR 2023 1 · Huajun Zhou, Bo Qiao, Lingxiao Yang, JianHuang Lai 외

Deep Learning-based Unsupervised Salient Object Detection (USOD) mainly relies on the noisy saliency pseudo labels that have been generated from traditional handcraft methods or pre-trained networks. To cope with the noi…

Objectobject-detectionObject DetectionOptical Flow Estimation+2

A New All-Digital Background Calibration Technique for Time-Interleaved ADC Using First Order Approximation FIR Filters

2018-06-24

This paper describes a new all-digital technique for calibration of the mismatches in time-interleaved analog-to-digital converters (TIADCs) to reduce the circuit area. The proposed technique gives the first order approx…

All