paper-with-me

홈 › Papers

Scalable Visual Pretraining for Language Intelligence

2026-07-10 · Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen arxiv

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.

📄 PDF Abstract BibTeX arXiv:2607.09657

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vision Pretraining for Dense Spatial Perception

2026-07-06 · Zelin Fu, Bin Tan, Changjiang Sun, Shaohui Liu 외 arxiv

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models te…

Depth CompletionDepth Estimation

Explore the Limits of Omni-modal Pretraining at Scale

2024-06-13 · Yiyuan Zhang, Handong Li, Jing Liu, Xiangyu Yue

We propose to build omni-modal intelligence, which is capable of understanding any modality and learning universal representations. In specific, we propose a scalable pretraining paradigm, named Multimodal Context (MiCo)…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+4

Rethinking Visual Intelligence: Insights from Video Pretraining

2025-10-28 · Pablo Acuaviva, Aram Davtyan, Mariam Hassan, Sebastian Stapf 외 arxiv

Large language models (LLMs) have demonstrated that large-scale pretraining enables systems to adapt rapidly to new problems with little supervision in the language domain. This success, however, has not translated as ef…

Scalable Vision Language Model Training via High Quality Data Curation

2025-01-10 · Hongyuan Dong, Zijian Kang, Weijie Yin, Xiao Liang 외

In this paper, we introduce SAIL-VL (ScAlable Vision Language Model TraIning via High QuaLity Data Curation), an open-source vision language model (VLM) of state-of-the-art (SOTA) performance with 2B parameters. We intro…

Instruction FollowingLanguage ModelingLanguage Modelling

Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos

2025-10-24 · Qixiu Li, Yu Deng, Yaobo Liang, Lin Luo 외 arxiv

This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as…