paper-with-me

Papers

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

2025-03-10 · Yingzhe Peng, Gongrui Zhang, Miaosen Zhang, Zhiyuan You, Jie Liu, Qipeng Zhu, Kai Yang, Xingzhong Xu, Xin Geng, Xu Yang

Enhancing reasoning in Large Multimodal Models (LMMs) faces unique challenges from the complex interplay between visual perception and logical reasoning, particularly in compact 3B-parameter architectures where architectural constraints limit reasoning capacity and modality alignment. While rule-based reinforcement learning (RL) excels in text-only domains, its multimodal extension confronts two critical barriers: (1) data limitations due to ambiguous answers and scarce complex reasoning examples, and (2) degraded foundational reasoning induced by multimodal pretraining. To address these challenges, we propose \textbf{\method}, a two-stage framework adapting rule-based RL for multimodal reasoning through \textbf{Foundational Reasoning Enhancement (FRE)} followed by \textbf{Multimodal Generalization Training (MGT)}. The FRE stage first strengthens reasoning abilities using text-only data with rule-based RL, then the MGT stage generalizes these reasoning capabilities to multimodal domains. Experiments on Qwen2.5-VL-Instruct-3B demonstrate that \method achieves 4.83\% and 4.5\% average improvements over baselines in multimodal and text-only benchmarks, respectively, with a 3.63\% gain in complex Football Game tasks. These results validate that text-based reasoning enhancement enables effective multimodal generalization, offering a data-efficient paradigm that bypasses costly high-quality multimodal training data.

📄 PDF Abstract BibTeX arXiv:2503.07536

Code (1)

tidedra/lmm-r1 pytorch

Tasks

Logical ReasoningMultimodal ReasoningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

V-Thinker: Interactive Thinking with Images

2025-11-06 · Runqi Qiao, Qiuna Tan, Minghan Yang, Guanting Dong 외 arxiv

Empowering Large Multimodal Models (LMMs) to deeply integrate image interaction with long-horizon reasoning capabilities remains a long-standing challenge in this field. Recent advances in vision-centric reasoning explor…

Reinforcement LearningMultimodal Reasoning

An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

2024-11-09 · Fatemeh Shiri, Xiao-Yu Guo, Mona Golestan Far, Xin Yu 외

Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA d…

object-detectionObject DetectionSpatial ReasoningVisual Question Answering (VQA)

LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

2024-09-26 · Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang 외

Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the developm…

3D Question Answering (3D-QA)PositionScene Understanding

VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

2024-06-17 · Yunxin Li, Xinyu Chen, Baotian Hu, Longyue Wang 외

Despite significant breakthroughs in video analysis driven by the rapid development of large multimodal models (LMMs), there remains a lack of a versatile evaluation benchmark to comprehensively assess these models' perf…

Anomaly DetectionLogical ReasoningObject TrackingVideo Understanding

AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation

2025-05-05 · Qingqiu Li, Zihang Cui, Seongsu Bae, Jilan Xu 외

Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Large Multimodal Models (LMMs) have enabled automated CXR interpretation, enhancing diagnostic accur…

AnatomyDiagnosticVisual Question Answering (VQA)