paper-with-me

Papers

Self-Improvement in Multimodal Large Language Models: A Survey

2025-10-03 · Shijian Deng, Kai Wang, Tianyu Yang, Harsh Singh, Yapeng Tian arxiv

Recent advancements in self-improvement for Large Language Models (LLMs) have efficiently enhanced model capabilities without significantly increasing costs, particularly in terms of human effort. While this area is still relatively young, its extension to the multimodal domain holds immense potential for leveraging diverse data sources and developing more general self-improving models. This survey is the first to provide a comprehensive overview of self-improvement in Multimodal LLMs (MLLMs). We provide a structured overview of the current literature and discuss methods from three perspectives: 1) data collection, 2) data organization, and 3) model optimization, to facilitate the further development of self-improvement in MLLMs. We also include commonly used evaluations and downstream applications. Finally, we conclude by outlining open challenges and future research directions.

📄 PDF Abstract BibTeX arXiv:2510.02665

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation

2023-06-12 · Jeremy Gwinnup, Kevin Duh

Large language models such as BERT and the GPT series started a paradigm shift that calls for building general-purpose models via pre-training on large datasets, followed by fine-tuning on task-specific datasets. There i…

Image CaptioningMachine TranslationMultimodal Machine TranslationQuestion Answering+2

A Survey of Generative Categories and Techniques in Multimodal Large Language Models

2025-05-29 · Longzhen Han, Awes Mubarak, Almas Baimagambetov, Nikolaos Polatidis 외

Multimodal Large Language Models (MLLMs) have rapidly evolved beyond text generation, now spanning diverse output modalities including images, music, video, human motion, and 3D objects, by integrating language with othe…

Mixture-of-ExpertsSelf-Supervised LearningSurveyText Generation

A Survey on Image-text Multimodal Models

2023-09-23 · Ruifeng Guo, Jingxuan Wei, Linzhuang Sun, Bihui Yu 외

With the significant advancements of Large Language Models (LLMs) in the field of Natural Language Processing (NLP), the development of image-text multimodal models has garnered widespread attention. Current surveys on i…

Survey

A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models

2025-07-31 · Xiaoling Luo, Ruli Zheng, Qiaojian Zheng, Zibo Du 외 arxiv

Visual impairment represents a major global health challenge, with multimodal imaging providing complementary information that is essential for accurate ophthalmic diagnosis. This comprehensive survey systematically revi…

Multimodal Deep LearningSelf-Supervised LearningReinforcement Learning

Survey on Self-Supervised Multimodal Representation Learning and Foundation Models

2022-11-29 · Sushil Thapa

Deep learning has been the subject of growing interest in recent years. Specifically, a specific type called Multimodal learning has shown great promise for solving a wide range of problems in domains such as language, v…

Representation LearningSelf-Supervised LearningSurvey