paper-with-me

Papers

MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens

2023-10-03 · Kaizhi Zheng, Xuehai He, Xin Eric Wang

The effectiveness of Multimodal Large Language Models (MLLMs) demonstrates a profound capability in multimodal understanding. However, the simultaneous generation of images with coherent texts is still underdeveloped. Addressing this, we introduce a novel interleaved vision-and-language generation method, centered around the concept of ``generative vokens". These vokens serve as pivotal elements contributing to coherent image-text outputs. Our method is marked by a unique two-stage training strategy for description-free multimodal generation, which does not necessitate extensive descriptions of images. We integrate classifier-free guidance to enhance the alignment of generated images and texts, ensuring more seamless and contextually relevant multimodal interactions. Our model, MiniGPT-5, exhibits substantial improvement over the baseline models on multimodal generation datasets, including MMDialog and VIST. The human evaluation shows MiniGPT-5 is better than the baseline model on more than 56\% cases for multimodal generation, highlighting its efficacy across diverse benchmarks.

📄 PDF Abstract BibTeX arXiv:2310.02239

Code (1)

eric-ai-lab/minigpt-5 공식 구현 pytorch

Tasks

Image Generationmultimodal generationReading ComprehensionText Generation

Similar Papers 제목 키워드 기반

MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

2024-04-04 · Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman 외

This paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding. The model is capable of processing both temporal visual and textual data, making it adept at un…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+10

Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models

2024-04-26 · Yuhang Huang, Zihan Wu, Chongyang Gao, Jiawei Peng 외

Large Vision-Language Models (LVLMs) are gaining traction for their remarkable ability to process and integrate visual and textual data. Despite their popularity, the capacity of LVLMs to generate precise, fine-grained t…

Retrieval

MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis

2024-07-04 · Asma Alkhaldi, Raneem Alnajim, Layan Alabdullatef, Rawan Alyahya 외

Recent advancements in artificial intelligence (AI) have precipitated significant breakthroughs in healthcare, particularly in refining diagnostic procedures. However, previous studies have often been constrained to limi…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model+4

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

2023-04-20 · Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 외

The recent GPT-4 has demonstrated extraordinary multi-modal abilities, such as directly generating websites from handwritten text and identifying humorous elements within images. These features are rarely observed in pre…

Image DescriptionLanguage ModellingLarge Language ModelSpatial Reasoning+4

MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4

2024-06-03 · Vahid Azizi, Fatemeh Koochaki

Vision-Language Models (VLMs) have recently seen significant advancements through integrating with Large Language Models (LLMs). The VLMs, which process image and text modalities simultaneously, have demonstrated the abi…