paper-with-me

홈 › Papers

InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing

2026-03-10 · Changyao Tian, Danni Yang, Guanzhou Chen, Erfei Cui, Zhaokai Wang, Yuchen Duan, Penghao Yin, Sitao Chen, Ganlin Yang, Mingxin Liu, Zirun Zhu, Ziqian Fan, Leyao Gu, Haomin Wang, Qi Wei, Jinhui Yin, Xue Yang, Zhihang Zhong, Qi Qin, Yi Xin, Bin Fu, Yihao Liu, Jiaye Ge, Qipeng Guo, Gen Luo, Hongsheng Li, Yu Qiao, Kai Chen, Hongjie Zhang arxiv

Unified multimodal models (UMMs) that integrate understanding, reasoning, generation, and editing face inherent trade-offs between maintaining strong semantic comprehension and acquiring powerful generation capabilities. In this report, we present InternVL-U, a lightweight 4B-parameter UMM that democratizes these capabilities within a unified framework. Guided by the principles of unified contextual modeling and modality-specific modular design with decoupled visual representations, InternVL-U integrates a state-of-the-art Multimodal Large Language Model (MLLM) with a specialized MMDiT-based visual generation head. To further bridge the gap between aesthetic generation and high-level intelligence, we construct a comprehensive data synthesis pipeline targeting high-semantic-density tasks, such as text rendering and scientific reasoning, under a reasoning-centric paradigm that leverages Chain-of-Thought (CoT) to better align abstract user intent with fine-grained visual generation details. Extensive experiments demonstrate that InternVL-U achieves a superior performance - efficiency balance. Despite using only 4B parameters, it consistently outperforms unified baseline models with over 3x larger scales such as BAGEL (14B) on various generation and editing tasks, while retaining strong multimodal understanding and reasoning capabilities.

📄 PDF Abstract BibTeX arXiv:2603.09877

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

2024-12-06 · Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu 외

We introduce InternVL 2.5, an advanced multimodal large language model (MLLM) series that builds upon InternVL 2.0, maintaining its core model architecture while introducing significant enhancements in training and testi…

document understandingHallucinationLanguage ModelingLanguage Modelling+7

InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation

2026-01-05 · Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen 외 arxiv

Prevalent Vision-Language-Action (VLA) models are typically built upon Multimodal Large Language Models (MLLMs) and demonstrate exceptional proficiency in semantic understanding, but they inherently lack the capability t…

Scene UnderstandingVideo Prediction

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

2025-08-25 · Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu 외 arxiv

We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the Cascade …

Reinforcement LearningOffline RL

FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion

2024-05-08 · Zehan Wang, Ziang Zhang, Xize Cheng, Rongjie Huang 외

Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhan…

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

2024-11-15 · Weiyun Wang, Zhe Chen, Wenhai Wang, Yue Cao 외

Existing open-source multimodal large language models (MLLMs) generally follow a training process involving pre-training and supervised fine-tuning. However, these models suffer from distribution shifts, which limit thei…

Multimodal Reasoning