paper-with-me

Papers

Ovis-U1 Technical Report

2025-06-29 · Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang, Liangfu Cao, Pengxin Zhan, Lunhao Duan, Shiyin Lu, Minghao Fu, Xiaohao Chen, Jianshan Zhao, Yang Li, Qing-Guo Chen

In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Building on the foundation of the Ovis series, Ovis-U1 incorporates a diffusion-based visual decoder paired with a bidirectional token refiner, enabling image generation tasks comparable to leading models like GPT-4o. Unlike some previous models that use a frozen MLLM for generation tasks, Ovis-U1 utilizes a new unified training approach starting from a language model. Compared to training solely on understanding or generation tasks, unified training yields better performance, demonstrating the enhancement achieved by integrating these two tasks. Ovis-U1 achieves a score of 69.6 on the OpenCompass Multi-modal Academic Benchmark, surpassing recent state-of-the-art models such as Ristretto-3B and SAIL-VL-1.5-2B. In text-to-image generation, it excels with scores of 83.72 and 0.89 on the DPG-Bench and GenEval benchmarks, respectively. For image editing, it achieves 4.00 and 6.42 on the ImgEdit-Bench and GEdit-Bench-EN, respectively. As the initial version of the Ovis unified model series, Ovis-U1 pushes the boundaries of multimodal understanding, generation, and editing.

📄 PDF Abstract BibTeX arXiv:2506.23044

Code (1)

aidc-ai/ovis-u1 공식 구현 pytorch

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Ovis-Image Technical Report

2025-11-28 · Guo-Hua Wang, Liangfu Cao, Tianyu Cui, Minghao Fu 외 arxiv

We introduce $\textbf{Ovis-Image}$, a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational constraints. Built upon our previous Ovis…

OvisOCR2 Technical Report

2026-07-15 · Shiyin Lu, Yinglun Li, Yu Xia, Yuhui Chen 외 hf

We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, t…

Reinforcement Learning

Multi-Source Transformer Architectures for Audiovisual Scene Classification

2022-10-18 · Wim Boes, Hugo Van hamme

In this technical report, the systems we submitted for subtask 1B of the DCASE 2021 challenge, regarding audiovisual scene classification, are described in detail. They are essentially multi-source transformers employing…

ClassificationScene Classification

Ovis2.5 Technical Report

2025-08-15 · Shiyin Lu, Yang Li, Yu Xia, Yuwei Hu 외 arxiv

We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer that processes images at their native, …

Multimodal Reasoning

CASTLE2026 Team WDL Technical Report

2026-05-30 · Zhengyang Li, Zhenglin Du, Yi Wen, Fang Liu 외 arxiv

The CASTLE Challenge @ EgoVis 2026 evaluates long-form egocentric video question answering over 600+ hours of multi-perspective recordings. Each four-choice question requires evidence from videos, transcripts, auxiliary …

Video Question AnsweringMultimodal Reasoning