paper-with-me

Papers

Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective

2024-10-29 · Shenghao Xie, Wenqiang Zu, Mingyang Zhao, Duo Su, Shilong Liu, Ruohua Shi, Guoqi Li, Shanghang Zhang, Lei Ma

Autoregression in large language models (LLMs) has shown impressive scalability by unifying all language tasks into the next token prediction paradigm. Recently, there is a growing interest in extending this success to vision foundation models. In this survey, we review the recent advances and discuss future directions for autoregressive vision foundation models. First, we present the trend for next generation of vision foundation models, i.e., unifying both understanding and generation in vision tasks. We then analyze the limitations of existing vision foundation models, and present a formal definition of autoregression with its advantages. Later, we categorize autoregressive vision foundation models from their vision tokenizers and autoregression backbones. Finally, we discuss several promising research challenges and directions. To the best of our knowledge, this is the first survey to comprehensively summarize autoregressive vision foundation models under the trend of unifying understanding and generation. A collection of related resources is available at https://github.com/EmmaSRH/ARVFM.

📄 PDF Abstract BibTeX arXiv:2410.22217

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Similar Papers 제목 키워드 기반

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

2026-06-24 · Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao 외 arxiv

Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepS…

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling

2025-02-12 · Shixiang Tang, Yizhou Wang, Lu Chen, YuAn Wang 외

Human understanding and generation are critical for modeling digital humans and humanoid embodiments. Recently, Human-centric Foundation Models (HcFMs) inspired by the success of generalist models, such as large language…

Survey

Multimodal Foundation Models: From Specialists to General-Purpose Assistants

2023-09-18 · Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang 외

This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to gene…

Image GenerationSurveyText to Image GenerationText-to-Image Generation

A Survey for Foundation Models in Autonomous Driving

2024-02-02 · Haoxiang Gao, Zhongruo Wang, Yaqian Li, Kaiwen Long 외

The advent of foundation models has revolutionized the fields of natural language processing and computer vision, paving the way for their application in autonomous driving (AD). This survey presents a comprehensive revi…

3D Object DetectionAutonomous DrivingCode Generationobject-detection+3

Igniting VLMs toward the Embodied Space

2025-09-15 · Andy Zhai, Brae Liu, Bruno Fang, Chalse Cai 외 arxiv

While foundation models show remarkable progress in language and vision, existing vision-language models (VLMs) still have limited spatial and embodiment understanding. Transferring VLMs to embodied domains reveals funda…