paper-with-me

Papers

Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

2024-06-25 · Wenhao Shi, Zhiqiang Hu, Yi Bin, Junhua Liu, Yang Yang, See-Kiong Ng, Lidong Bing, Roy Ka-Wei Lee

Large language models (LLMs) have demonstrated impressive reasoning capabilities, particularly in textual mathematical problem-solving. However, existing open-source image instruction fine-tuning datasets, containing limited question-answer pairs per image, do not fully exploit visual information to enhance the multimodal mathematical reasoning capabilities of Multimodal LLMs (MLLMs). To bridge this gap, we address the lack of high-quality, diverse multimodal mathematical datasets by collecting 40K high-quality images with question-answer pairs from 24 existing datasets and synthesizing 320K new pairs, creating the MathV360K dataset, which enhances both the breadth and depth of multimodal mathematical questions. We introduce Math-LLaVA, a LLaVA-1.5-based model fine-tuned with MathV360K. This novel approach significantly improves the multimodal mathematical reasoning capabilities of LLaVA-1.5, achieving a 19-point increase and comparable performance to GPT-4V on MathVista's minitest split, and yielding leading performance on Math-V and MathVerse. Furthermore, Math-LLaVA demonstrates enhanced generalizability, showing substantial improvements on the MMMU benchmark. Our research highlights the importance of dataset diversity and synthesis in advancing MLLMs' mathematical reasoning abilities. The code and data are available at: \url{https://github.com/HZQ950419/Math-LLaVA}.

📄 PDF Abstract BibTeX arXiv:2406.17294

Code (1)

hzq950419/math-llava 공식 구현 pytorch

Tasks

DiversityMathMathematical Problem-SolvingMathematical Reasoning

Similar Papers 제목 키워드 기반

G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

2023-12-18 · Jiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye 외

Large language models (LLMs) have shown remarkable proficiency in human-level reasoning and generation capabilities, which encourages extensive research on their application in mathematical problem solving. However, curr…

Language ModelingLanguage ModellingLarge Language ModelMathematical Problem-Solving

LEMMA: Bootstrapping High-Level Mathematical Reasoning with Learned Symbolic Abstractions

2022-11-16 · Zhening Li, Gabriel Poesia, Omar Costilla-Reyes, Noah Goodman 외

Humans tame the complexity of mathematical reasoning by developing hierarchies of abstractions. With proper abstractions, solutions to hard problems can be expressed concisely, thus making them more likely to be found. I…

LEMMAMathematical ReasoningVocal Bursts Intensity Prediction

Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning

2024-12-04 · Neale Ratzlaff, Man Luo, Xin Su, Vasudev Lal 외

Multimodal models typically combine a powerful large language model (LLM) with a vision encoder and are then trained on multimodal data via instruction tuning. While this process adapts LLMs to multimodal settings, it re…

GSM8KLanguage ModelingLanguage ModellingLarge Language Model+1

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

2023-09-21 · Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu 외

Large language models (LLMs) have pushed the limits of natural language understanding and exhibited excellent problem-solving ability. Despite the great success, most existing open-source LLMs (e.g., LLaMA-2) are still f…

Arithmetic ReasoningGSM8KLanguage ModelingLanguage Modelling+4

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

2024-12-16 · Yibo Yan, Jiamin Su, Jianxiang He, Fangteng Fu 외

Mathematical reasoning, a core aspect of human cognition, is vital across many domains, from educational problem-solving to scientific advancements. As artificial general intelligence (AGI) progresses, integrating large …

Language ModelingLanguage ModellingLarge Language ModelMath+4