Self-Improvement in Multimodal Large Language Models: A Survey
Recent advancements in self-improvement for Large Language Models (LLMs) have efficiently enhanced model capabilities without significantly increasing costs, particularly in terms of human effort. While this area is still relatively young, its extension to the multimodal domain holds immense potential for leveraging diverse data sources and developing more general self-improving models. This survey is the first to provide a comprehensive overview of self-improvement in Multimodal LLMs (MLLMs). We provide a structured overview of the current literature and discuss methods from three perspectives: 1) data collection, 2) data organization, and 3) model optimization, to facilitate the further development of self-improvement in MLLMs. We also include commonly used evaluations and downstream applications. Finally, we conclude by outlining open challenges and future research directions.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation
Large language models such as BERT and the GPT series started a paradigm shift that calls for building general-purpose models via pre-training on large datasets, followed by fine-tuning on task-specific datasets. There i…
Image CaptioningMachine TranslationMultimodal Machine TranslationQuestion Answering+2A Survey of Generative Categories and Techniques in Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) have rapidly evolved beyond text generation, now spanning diverse output modalities including images, music, video, human motion, and 3D objects, by integrating language with othe…
Mixture-of-ExpertsSelf-Supervised LearningSurveyText GenerationA Survey on Image-text Multimodal Models
With the significant advancements of Large Language Models (LLMs) in the field of Natural Language Processing (NLP), the development of image-text multimodal models has garnered widespread attention. Current surveys on i…
SurveyA Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models
Visual impairment represents a major global health challenge, with multimodal imaging providing complementary information that is essential for accurate ophthalmic diagnosis. This comprehensive survey systematically revi…
Multimodal Deep LearningSelf-Supervised LearningReinforcement LearningSurvey on Self-Supervised Multimodal Representation Learning and Foundation Models
Deep learning has been the subject of growing interest in recent years. Specifically, a specific type called Multimodal learning has shown great promise for solving a wide range of problems in domains such as language, v…
Representation LearningSelf-Supervised LearningSurvey