paper-with-me

홈 › Papers

R1-Onevision:An Open-Source Multimodal Large Language Model Capable of Deep Reasoning

2025-02-24 · ongoing 2025 2 · Yi Yang*, Xiaoxuan He*, Hongkun Pan*, Xiyan Jiang, Yan Deng, Xingtao Yang, Haoyu Lu, Minfeng Zhu†, Bo Zhang†, Wei Chen†

R1-OneVision is a versatile multimodal reasoning large model, designed to tackle complex visual reasoning tasks. It seamlessly integrates visual and textual data to offer precise interpretations of multimodal information, excelling in areas such as mathematics, science, deep image understanding, and logical reasoning. With its robust ability to perform multimodal reasoning, R1-OneVision emerges as a powerful AI assistant capable of addressing a wide range of problem-solving challenges across different domains.

📄 PDF Abstract BibTeX

Code (1)

Fancy-MLLM/R1-onevision pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelLogical ReasoningMultimodal Large Language ModelMultimodal ReasoningVisual Reasoning

Similar Papers 제목 키워드 기반

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

2025-09-28 · Xiang An, Yin Xie, Kaicheng Yang, Wenkang Zhang 외 arxiv

We present LLaVA-OneVision-1.5, a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. Different from the existing works, …

Multimodal Reasoning

LLaVA-OneVision: Easy Visual Task Transfer

2024-08-06 · Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang 외

We present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT blog series. Our experimental results de…

3D Question Answering (3D-QA)Multiple-choiceTemporal Relation Extraction+6

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

2025-03-13 · Yi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang 외

Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge.…

Multimodal Reasoning

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

2026-05-25 · Xiang An, Yin Xie, Feilong Tang, Yunyao Yan 외 arxiv

We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of multimodal benchmarks. The model builds on…

Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks

2026-06-09 · Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou 외 arxiv

RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery. However, existing models support only a narrow range of sensor types and tasks, yielding a fragmented view of the earth…

Spatial ReasoningVisual Grounding