paper-with-me

홈 › Papers

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

2024-12-06 · Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, YuBo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, Xiang Yue

Open-source multimodal large language models (MLLMs) have shown significant potential in a broad range of multimodal tasks. However, their reasoning capabilities remain constrained by existing instruction-tuning datasets, which were predominately repurposed from academic datasets such as VQA, AI2D, and ChartQA. These datasets target simplistic tasks, and only provide phrase-level answers without any intermediate rationales. To address these challenges, we introduce a scalable and cost-effective method to construct a large-scale multimodal instruction-tuning dataset with rich intermediate rationales designed to elicit CoT reasoning. Using only open models, we create a dataset containing 12M instruction-response pairs to cover diverse, reasoning-intensive tasks with detailed and faithful rationales. Experiments demonstrate that training MLLMs on this dataset significantly improves reasoning capabilities, achieving state-of-the-art performance on benchmarks such as MathVerse (+8.1%), MMMU-Pro (+7%), and MuirBench (+13.3%). Additionally, the model demonstrates notable improvements of up to 4% on non-reasoning-based benchmarks. Ablation studies further highlight the importance of key components, such as rewriting and self-filtering, in the dataset construction process.

📄 PDF Abstract BibTeX arXiv:2412.05237

Code (1)

mammoth-vl/mammoth-vl pytorch

Tasks

Multimodal ReasoningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

MAmmoTH2: Scaling Instructions from the Web

2024-05-06 · Xiang Yue, Tuney Zheng, Ge Zhang, Wenhu Chen

Instruction tuning improves the reasoning abilities of large language models (LLMs), with data quality and scalability being the crucial factors. Most instruction tuning data come from human crowd-sourcing or GPT-4 disti…

ChatbotGSM8KMath

MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation

2025-11-23 · Tao Shen, Xin Wan, Taicai Chen, Rui Zhang 외 arxiv

Unified multimodal models aim to integrate understanding and generation within a single framework, yet bridging the gap between discrete semantic reasoning and high-fidelity visual synthesis remains challenging. We prese…

Reinforcement Learning

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

2023-09-11 · Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu 외

We introduce MAmmoTH, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. The MAmmoTH models are trained on MathInstruct, our meticulously curated instruction tuni…

MathMathematical Reasoning

Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models

2025-02-17 · Zikang Liu, Kun Zhou, Wayne Xin Zhao, Dawei Gao 외

Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language models (LVLMs). Despite the success, as visual instructions require images as…

Instruction Followingvisual instruction followingVisual Reasoning

VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search

2025-03-13 · Yiming Jia, Jiachen Li, Xiang Yue, Bo Li 외

Vision-Language Models have made significant progress on many perception-focused tasks, however, their progress on reasoning-focused tasks seem to be limited due to the lack of high-quality and diverse training data. In …

Image RetrievalMath