paper-with-me

홈 › Papers

Empowering Lightweight MLLMs with Reasoning via Long CoT SFT

2025-09-03 · Linyu Ou, YuYang Yin arxiv

While Reinforcement Learning with Verifiable Rewards has enhanced the reasoning of large-scale language models (LLMs), its efficacy for lightweight multimodal language models (MLLMs) with fewer than seven billion parameters remains underexplored. This paper investigates the role of long Chain-of-Thought (long CoT) data in enhancing the reasoning abilities of such MLLMs. Our findings demonstrate that Supervised Fine-Tuning (SFT) with long CoT data significantly improves MLLM reasoning. Furthermore, we observe that after this initial SFT phase, MLLMs can achieve additional performance gains through a subsequent RL stage. We conclude that a SFT stage with long CoT data is a critical prerequisite for developing the reasoning capabilities of lightweight MLLMs.

📄 PDF Abstract BibTeX arXiv:2509.03321

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Empowering Segmentation Ability to Multi-modal Large Language Models

2024-03-21 · YuQi Yang, Peng-Tao Jiang, Jing Wang, Hao Zhang 외

Multi-modal large language models (MLLMs) can understand image-language prompts and demonstrate impressive reasoning ability. In this paper, we extend MLLMs' output by empowering MLLMs with the segmentation ability. The …

Dialogue GenerationReasoning SegmentationSegmentationWord Embeddings

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

2025-01-21 · Yi Wang, Xinhao Li, Ziang Yan, Yinan He 외

This paper aims to improve the performance of video multimodal large language models (MLLM) via long and rich context (LRC) modeling. As a result, we develop a new version of InternVideo2.5 with a focus on enhancing the …

Object TrackingReferring Expression SegmentationReferring Video Object SegmentationVideo Understanding

Position: Empowering Time Series Reasoning with Multimodal LLMs

2025-02-03 · Yaxuan Kong, Yiyuan Yang, Shiyu Wang, Chenghao Liu 외

Understanding time series data is crucial for multiple real-world applications. While large language models (LLMs) show promise in time series tasks, current approaches often rely on numerical data alone, overlooking the…

Decision MakingMultimodal ReasoningPositionTime Series+1

Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

2024-08-01 · CVPR 2025 1 · Benlin Liu, Yuhao Dong, Yiqin Wang, Zixian Ma 외

Multimodal language models (MLLMs) are increasingly being applied in real-world environments, necessitating their ability to interpret 3D spaces and comprehend temporal dynamics. Current methods often rely on specialized…

EgoSchemaLanguage ModelingLanguage ModellingSpatial Reasoning+1

M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning

2025-07-11 · Inclusion AI, :, Fudong Wang, Jiajia Liu 외

Recent advancements in Multimodal Large Language Models (MLLMs), particularly through Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced their reasoning abilities. However, a critical gap …

Spatial Reasoning