paper-with-me

홈 › Papers

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

2025-03-13 · Yi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang, Yan Deng, Xingtao Yang, Haoyu Lu, Dacheng Yin, Fengyun Rao, Minfeng Zhu, Bo Zhang, Wei Chen

Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge. Existing visual-language models often struggle to effectively analyze and reason visual content, resulting in suboptimal performance on complex reasoning tasks. Moreover, the absence of comprehensive benchmarks hinders the accurate assessment of multimodal reasoning capabilities. In this paper, we introduce R1-Onevision, a multimodal reasoning model designed to bridge the gap between visual perception and deep reasoning. To achieve this, we propose a cross-modal reasoning pipeline that transforms images into formal textural representations, enabling precise language-based reasoning. Leveraging this pipeline, we construct the R1-Onevision dataset which provides detailed, step-by-step multimodal reasoning annotations across diverse domains. We further develop the R1-Onevision model through supervised fine-tuning and reinforcement learning to cultivate advanced reasoning and robust generalization abilities. To comprehensively evaluate multimodal reasoning performance across different grades, we introduce R1-Onevision-Bench, a benchmark aligned with human educational stages, covering exams from junior high school to university and beyond. Experimental results show that R1-Onevision achieves state-of-the-art performance, outperforming models such as GPT-4o and Qwen2.5-VL on multiple challenging multimodal reasoning benchmarks.

📄 PDF Abstract BibTeX arXiv:2503.10615

Code (1)

Fancy-MLLM/R1-onevision 공식 구현 pytorch

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

R1-Onevision:An Open-Source Multimodal Large Language Model Capable of Deep Reasoning

2025-02-24 · ongoing 2025 2 · Yi Yang*, Xiaoxuan He*, Hongkun Pan*, Xiyan Jiang 외

R1-OneVision is a versatile multimodal reasoning large model, designed to tackle complex visual reasoning tasks. It seamlessly integrates visual and textual data to offer precise interpretations of multimodal information…

Language ModelingLanguage ModellingLarge Language ModelLogical Reasoning+3

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

2025-09-28 · Xiang An, Yin Xie, Kaicheng Yang, Wenkang Zhang 외 arxiv

We present LLaVA-OneVision-1.5, a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. Different from the existing works, …

Multimodal Reasoning

LLaVA-OneVision: Easy Visual Task Transfer

2024-08-06 · Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang 외

We present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT blog series. Our experimental results de…

3D Question Answering (3D-QA)Multiple-choiceTemporal Relation Extraction+6

Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks

2026-06-09 · Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou 외 arxiv

RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery. However, existing models support only a narrow range of sensor types and tasks, yielding a fragmented view of the earth…

Spatial ReasoningVisual Grounding

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

2026-05-25 · Xiang An, Yin Xie, Feilong Tang, Yunyao Yan 외 arxiv

We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of multimodal benchmarks. The model builds on…