paper-with-me

홈 › Papers

Towards the Unification of Generative and Discriminative Visual Foundation Model: A Survey

2023-12-15 · Xu Liu, Tong Zhou, Yuanxin Wang, Yuping Wang, Qinjingwen Cao, Weizhi Du, Yonghuan Yang, Junjun He, Yu Qiao, Yiqing Shen

The advent of foundation models, which are pre-trained on vast datasets, has ushered in a new era of computer vision, characterized by their robustness and remarkable zero-shot generalization capabilities. Mirroring the transformative impact of foundation models like large language models (LLMs) in natural language processing, visual foundation models (VFMs) have become a catalyst for groundbreaking developments in computer vision. This review paper delineates the pivotal trajectories of VFMs, emphasizing their scalability and proficiency in generative tasks such as text-to-image synthesis, as well as their adeptness in discriminative tasks including image segmentation. While generative and discriminative models have historically charted distinct paths, we undertake a comprehensive examination of the recent strides made by VFMs in both domains, elucidating their origins, seminal breakthroughs, and pivotal methodologies. Additionally, we collate and discuss the extensive resources that facilitate the development of VFMs and address the challenges that pave the way for future research endeavors. A crucial direction for forthcoming innovation is the amalgamation of generative and discriminative paradigms. The nascent application of generative models within discriminative contexts signifies the early stages of this confluence. This survey aspires to be a contemporary compendium for scholars and practitioners alike, charting the course of VFMs and illuminating their multifaceted landscape.

📄 PDF Abstract BibTeX arXiv:2312.10163

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationImage SegmentationSemantic SegmentationZero-shot Generalization

Similar Papers 제목 키워드 기반

Anti-unification and Generalization: A Survey

2023-02-01 · David M. Cerna, Temur Kutsia

Anti-unification (AU) is a fundamental operation for generalization computation used for inductive inference. It is the dual operation to unification, an operation at the foundation of automated theorem proving. Interest…

Automated Theorem ProvingSurvey

A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems

2026-08-12 · Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang arxiv

As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.…

A Survey of Spatio-Temporal EEG data Analysis: from Models to Applications

2024-09-26 · Pengfei Wang, Huanran Zheng, Silong Dai, Yiqiao Wang 외

In recent years, the field of electroencephalography (EEG) analysis has witnessed remarkable advancements, driven by the integration of machine learning and artificial intelligence. This survey aims to encapsulate the la…

EEGSelf-Supervised LearningSurvey

Representation Potentials of Foundation Models for Multimodal Alignment: A Survey

2025-10-05 · Jianglin Lu, Hailing Wang, Yi Xu, Yizhou Wang 외 arxiv

Foundation models learn highly transferable representations through large-scale pretraining on diverse data. An increasing body of research indicates that these representations exhibit a remarkable degree of similarity a…

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

2026-08-05 · Junlin Han, Shengbang Tong, David Fan, Minghao Chen 외 hf

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities int…