paper-with-me

Papers

Multimodal Foundation Models: From Specialists to General-Purpose Assistants

2023-09-18 · Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, Jianfeng Gao

This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to general-purpose assistants. The research landscape encompasses five core topics, categorized into two classes. (i) We start with a survey of well-established research areas: multimodal foundation models pre-trained for specific purposes, including two topics -- methods of learning vision backbones for visual understanding and text-to-image generation. (ii) Then, we present recent advances in exploratory, open research areas: multimodal foundation models that aim to play the role of general-purpose assistants, including three topics -- unified vision models inspired by large language models (LLMs), end-to-end training of multimodal LLMs, and chaining multimodal tools with LLMs. The target audiences of the paper are researchers, graduate students, and professionals in computer vision and vision-language multimodal communities who are eager to learn the basics and recent advances in multimodal foundation models.

📄 PDF Abstract BibTeX arXiv:2309.10020

Code (1)

computer-vision-in-the-wild/cvinw_readings 공식 구현

Tasks

Image GenerationSurveyText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

A Foundational Multimodal Vision Language AI Assistant for Human Pathology

2023-12-13 · Ming Y. Lu, Bowen Chen, Drew F. K. Williamson, Richard J. Chen 외

The field of computational pathology has witnessed remarkable progress in the development of both task-specific predictive models and task-agnostic self-supervised vision encoders. However, despite the explosive growth o…

Decision MakingDiagnosticLanguage ModellingLarge Language Model+1

Vesta: A Generalist Embodied Reasoning Model

2026-06-18 · Johan Bjorck, Zhiqi Li, Yunze Man, Jing Wang 외 arxiv

Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model sta…

Spatial Reasoning

The Backfiring Effect of Weak AI Safety Regulation

2025-03-26 · Benjamin Laufer, Jon Kleinberg, Hoda Heidari

Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches to AI safety. We present a strategic model that explores the i…

MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine

2026-03-01 · Kai Zhang, Zhengqing Yuan, Cheng Peng, Songlin Zhao 외 arxiv

Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remains: top-performing systems are either closed-source or computationally pr…

Multimodal Reasoning

A General-Purpose Device for Interaction with LLMs

2024-08-02 · Jiajun Xu, Qun Wang, Yuhang Cao, Baitao Zeng 외

This paper investigates integrating large language models (LLMs) with advanced hardware, focusing on developing a general-purpose device designed for enhanced interaction with LLMs. Initially, we analyze the current land…