Large Multimodal Models: Notes on CVPR 2023 Tutorial
This tutorial note summarizes the presentation on `Large Multimodal Models: Towards Building and Surpassing Multimodal GPT-4'', a part of CVPR 2023 tutorial on `Recent Advances in Vision Foundation Models''. The tutorial consists of three parts. We first introduce the background on recent GPT-like large models for vision-and-language modeling to motivate the research in instruction-tuned large multimodal models (LMMs). As a pre-requisite, we describe the basics of instruction-tuning in large language models, which is further extended to the multimodal space. Lastly, we illustrate how to build the minimum prototype of multimodal GPT-4 like models with the open-source resource, and review the recently emerged topics.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Multimodal Machine Learning: Integrating Language, Vision and Speech
Multimodal machine learning is a vibrant multi-disciplinary research field which addresses some of the original goals of artificial intelligence by integrating and modeling multiple communicative modalities, including li…
Audio-Visual Speech RecognitionBIG-bench Machine LearningImage CaptioningQuestion Answering+8Lecture Notes: Optimization for Machine Learning
Lecture notes on optimization for machine learning, derived from a course at Princeton University and tutorials given in MLSS, Buenos Aires, as well as Simons Foundation, Berkeley.
BIG-bench Machine LearningTopology in Sound Synthesis and Digital Signal Processing -- DAFx2022 Lecture Notes
Lecture notes of a tutorial on topology in sound synthesis and digital signal processing held at international conference for digital audio effects (DAFx-22) in Vienna, Austria.
Combinatorial Hodge Theory in Simplicial Signal Processing -- DAFx2023 Lecture Notes
Lecture notes of a tutorial on Combinatorial Hodge Theory in Simplicial Signal Processing held at international conference for digital audio effects (DAFx-23) in Copenhagen, Denmark.
Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understandi…
Question AnsweringVisual Question AnsweringVisual Storytelling