paper-with-me

홈 › Papers

R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

2025-03-07 · Hengguang Zhou, Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou, Cho-Jui Hsieh

Recently DeepSeek R1 demonstrated how reinforcement learning with simple rule-based incentives can enable autonomous development of complex reasoning in large language models, characterized by the "aha moment", in which the model manifest self-reflection and increased response length during training. However, attempts to extend this success to multimodal reasoning often failed to reproduce these key characteristics. In this report, we present the first successful replication of these emergent characteristics for multimodal reasoning on only a non-SFT 2B model. Starting with Qwen2-VL-2B and applying reinforcement learning directly on the SAT dataset, our model achieves 59.47% accuracy on CVBench, outperforming the base model by approximately ~30% and exceeding both SFT setting by ~2%. In addition, we share our failed attempts and insights in attempting to achieve R1-like reasoning using RL with instruct models. aiming to shed light on the challenges involved. Our key observations include: (1) applying RL on instruct model often results in trivial reasoning trajectories, and (2) naive length reward are ineffective in eliciting reasoning capabilities. The project code is available at https://github.com/turningpoint-ai/VisualThinker-R1-Zero

📄 PDF Abstract BibTeX arXiv:2503.05132

Code (1)

turningpoint-ai/visualthinker-r1-zero 공식 구현 pytorch

Tasks

Multimodal Reasoningreinforcement-learningReinforcement LearningVisual Reasoning

Methods 이 논문이 사용한 방법론

SFT Shrink and Fine-Tune, or SFT, is a type of distillation that avoids explicit distillation by copying parameters to a student student model and then fine-tuning.…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

2024-02-18 · Long Qian, Juncheng Li, Yu Wu, Yaobo Ye 외

Large Language Models (LLMs) demonstrate remarkable proficiency in comprehending and handling text-based tasks. Many efforts are being made to transfer these attributes to video modality, which are termed Video-LLMs. How…

Language ModelingLanguage ModellingLarge Language Model

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

2026-06-01 · Xiaolin Liu, Yilun Zhu, Xiangyu Zhao, Xuehui Wang 외 arxiv

Video multimodal large language models (MLLMs) have made rapid progress on general and long-form video understanding, yet their ability to preserve brief answer-critical visual evidence remains underexplored. Many practi…

SketchQL Demonstration: Zero-shot Video Moment Querying with Sketches

2024-05-28 · Renzhi Wu, Pramod Chunduri, Dristi J Shah, Ashmitha Julius Aravind 외

In this paper, we will present SketchQL, a video database management system (VDBMS) for retrieving video moments with a sketch-based query interface. This novel interface allows users to specify object trajectory events …

ManagementRetrieval

Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis

2024-08-27 · Aishik Nagar, Shantanu Jaiswal, Cheston Tan

Vision-language models (VLMs) have shown impressive zero- and few-shot performance on real-world visual question answering (VQA) benchmarks, alluding to their capabilities as visual reasoning engines. However, the benchm…

BenchmarkingLarge Language ModelQuestion AnsweringVisual Question Answering+3

TACIT: Transformation-Aware Capturing of Implicit Thought

2026-02-05 · Daniel Nobrega arxiv

We present TACIT (Transformation-Aware Capturing of Implicit Thought), a diffusion-based transformer for interpretable visual reasoning. Unlike language-based reasoning systems, TACIT operates entirely in pixel space usi…

Visual Reasoning