paper-with-me

Papers

Ovis-Image Technical Report

2025-11-28 · Guo-Hua Wang, Liangfu Cao, Tianyu Cui, Minghao Fu, Xiaohao Chen, Pengxin Zhan, Jianshan Zhao, Lan Li, Bowen Fu, Jiaqi Liu, Qing-Guo Chen arxiv

We introduce $\textbf{Ovis-Image}$, a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational constraints. Built upon our previous Ovis-U1 framework, Ovis-Image integrates a diffusion-based visual decoder with the stronger Ovis 2.5 multimodal backbone, leveraging a text-centric training pipeline that combines large-scale pre-training with carefully tailored post-training refinements. Despite its compact architecture, Ovis-Image achieves text rendering performance on par with significantly larger open models such as Qwen-Image and approaches closed-source systems like Seedream and GPT4o. Crucially, the model remains deployable on a single high-end GPU with moderate memory, narrowing the gap between frontier-level text rendering and practical deployment. Our results indicate that combining a strong multimodal backbone with a carefully designed, text-focused training recipe is sufficient to achieve reliable bilingual text rendering without resorting to oversized or proprietary models.

📄 PDF Abstract BibTeX arXiv:2511.22982

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Ovis-U1 Technical Report

2025-06-29 · Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang, Liangfu Cao 외

In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Building on the foundation of the Ovis series,…

Image GenerationText to Image GenerationText-to-Image Generation

OvisOCR2 Technical Report

2026-07-15 · Shiyin Lu, Yinglun Li, Yu Xia, Yuhui Chen 외 hf

We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, t…

Reinforcement Learning

CASTLE2026 Team WDL Technical Report

2026-05-30 · Zhengyang Li, Zhenglin Du, Yi Wen, Fang Liu 외 arxiv

The CASTLE Challenge @ EgoVis 2026 evaluates long-form egocentric video question answering over 600+ hours of multi-perspective recordings. Each four-choice question requires evidence from videos, transcripts, auxiliary …

Video Question AnsweringMultimodal Reasoning

Ovis2.5 Technical Report

2025-08-15 · Shiyin Lu, Yang Li, Yu Xia, Yuwei Hu 외 arxiv

We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer that processes images at their native, …

Multimodal Reasoning

Multi-Source Transformer Architectures for Audiovisual Scene Classification

2022-10-18 · Wim Boes, Hugo Van hamme

In this technical report, the systems we submitted for subtask 1B of the DCASE 2021 challenge, regarding audiovisual scene classification, are described in detail. They are essentially multi-source transformers employing…

ClassificationScene Classification