paper-with-me

Papers

MMCR: Advancing Visual Language Model in Multimodal Multi-Turn Contextual Reasoning

2025-03-24 · Dawei Yan, Yang Li, Qing-Guo Chen, Weihua Luo, Peng Wang, Haokui Zhang, Chunhua Shen

Compared to single-turn dialogue, multi-turn dialogue involving multiple images better aligns with the needs of real-world human-AI interactions. Additionally, as training data, it provides richer contextual reasoning information, thereby guiding the model to achieve better performance. However, existing vision-language models (VLMs) primarily rely on single-turn dialogue training and evaluation benchmarks. In this paper, following the characteristics of human dialogue, such as focused topics and concise, clear content, we present MMCR (Multimodal Multi-turn Contextual Reasoning), a novel dataset comprising: (1) MMCR-310k -- the largest multi-image multi-turn instruction tuning dataset with 310K contextual dialogues, each covering 1-4 images and 4 or 8 dialogue turns; and (2) MMCR-Bench -- a diagnostic benchmark featuring dialogues, spanning 8 domains (Humanities, Natural, Science, Education, etc.) and 40 sub-topics. Extensive evaluations demonstrate that models fine-tuned with MMCR-310k achieve 5.2\% higher contextual accuracy on MMCR-Bench, while showing consistent improvements on existing benchmarks (+1.1\% on AI2D, +1.2\% on MMMU and MMVet). MMCR and prompt engineering will be released publicly.

📄 PDF Abstract BibTeX arXiv:2503.18533

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticLanguage ModelingLanguage ModellingPrompt Engineering

Similar Papers 제목 키워드 기반

Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs

2025-08-24 · Somraj Gautam, Abhirama Subramanyam Penamakuri, Abhishek Bhandari, Gaurav Harit arxiv

We introduce MMCRICBENCH-3K, a benchmark for Visual Question Answering (VQA) on cricket scorecards, designed to evaluate large vision-language models (LVLMs) on complex numerical and cross-lingual reasoning over semi-str…

Visual Question Answering

Towards an Improved Understanding and Utilization of Maximum Manifold Capacity Representations

2024-06-13 · Rylan Schaeffer, Victor Lecomte, Dhruv Bhandarkar Pai, Andres Carranza 외

Maximum Manifold Capacity Representations (MMCR) is a recent multi-view self-supervised learning (MVSSL) method that matches or surpasses other leading MVSSL methods. MMCR is intriguing because it does not fit neatly int…

Self-Supervised Learning

ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area

2024-08-14 · Junxian Li, Di Zhang, Xunzhi Wang, Zeying Hao 외

Large Language Models (LLMs) have achieved remarkable success and have been applied across various scientific fields, including chemistry. However, many chemical tasks require the processing of visual information, which …

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3

Multimodal and Crossmodal AI for Smart Data Analysis

2022-09-03 · Minh-Son Dao

Recently, the multimodal and crossmodal AI techniques have attracted the attention of communities. The former aims to collect disjointed and heterogeneous data to compensate for complementary information to enhance robus…

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

2025-05-29 · CVPR 2025 1 · Yunze Man, De-An Huang, Guilin Liu, Shiwei Sheng 외

Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks, yet they often struggle with vision-centric scenarios where precise visual focus is needed f…

Multimodal Reasoning