paper-with-me

홈 › Papers

Vision Generalist Model: A Survey

2025-06-11 · Ziyi Wang, Yongming Rao, Shuofeng Sun, Xinrun Liu, Yi Wei, Xumin Yu, Zuyan Liu, Yanbo Wang, Hongmin Liu, Jie zhou, Jiwen Lu

Recently, we have witnessed the great success of the generalist model in natural language processing. The generalist model is a general framework trained with massive data and is able to process various downstream tasks simultaneously. Encouraged by their impressive performance, an increasing number of researchers are venturing into the realm of applying these models to computer vision tasks. However, the inputs and outputs of vision tasks are more diverse, and it is difficult to summarize them as a unified representation. In this paper, we provide a comprehensive overview of the vision generalist models, delving into their characteristics and capabilities within the field. First, we review the background, including the datasets, tasks, and benchmarks. Then, we dig into the design of frameworks that have been proposed in existing research, while also introducing the techniques employed to enhance their performance. To better help the researchers comprehend the area, we take a brief excursion into related domains, shedding light on their interconnections and potential synergies. To conclude, we provide some real-world application scenarios, undertake a thorough examination of the persistent challenges, and offer insights into possible directions for future research endeavors.

📄 PDF Abstract BibTeX arXiv:2506.09954

Code (0)

등록된 구현이 없습니다.

Tasks

modelSurvey

Similar Papers 제목 키워드 기반

Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches

2025-06-12 · Andrea Moglia, Matteo Leccardi, Matteo Cavicchioli, Alice Maccarini 외

Following the successful paradigm shift of large language models, leveraging pre-training on a massive corpus of data and fine-tuning on different downstream tasks, generalist models have made their foray into computer v…

Image SegmentationMedical Image SegmentationSemantic Segmentation

Towards Generalist Robot Learning from Internet Video: A Survey

2024-04-30 · Robert McCarthy, Daniel C. H. Tan, Dominik Schmidt, Fernando Acero 외

Scaling deep learning to massive, diverse internet data has yielded remarkably general capabilities in visual and natural language understanding and generation. However, data has remained scarce and challenging to collec…

Natural Language UnderstandingReinforcement Learning (RL)Survey

Hidden-Shot: Towards One-Shot Task Generalization for Low-Level Vision Generalist Models

2026-07-01 · Shao-Jun Xia, Xianzheng Ma, Zichong Meng arxiv

Despite the intense engagement surrounding low-level vision generalist models, their effectiveness in zero/few-shot scenarios beyond learned tasks remains unverified. The primary challenge of developing an ideal generali…

Prompt Engineering

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM

2026-05-04 · Hyobin Park, Minseok Seo, Dong-Geol Choi arxiv

Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particularly important in remote sensing, where data acquisition is costly …

Image Retrieval

Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks

2022-11-17 · CVPR 2023 1 · Hao Li, Jinguo Zhu, Xiaohu Jiang, Xizhou Zhu 외

Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to eliminating this inconsistency is to use gene…

DecoderLanguage ModellingMulti-Task Learning