paper-with-me

Papers

A Survey of Multimodal Large Language Model from A Data-centric Perspective

2024-05-26 · Tianyi Bai, Hao Liang, Binwang Wan, Yanran Xu, Xi Li, Shiyu Li, Ling Yang, Bozhou Li, Yifan Wang, Bin Cui, Ping Huang, Jiulong Shan, Conghui He, Binhang Yuan, Wentao Zhang

Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal role in the development and refinement of these models. In this survey, we comprehensively review the literature on MLLMs from a data-centric perspective. Specifically, we explore methods for preparing multimodal data during the pretraining and adaptation phases of MLLMs. Additionally, we analyze the evaluation methods for the datasets and review the benchmarks for evaluating MLLMs. Our survey also outlines potential future research directions. This work aims to provide researchers with a detailed understanding of the data-driven aspects of MLLMs, fostering further exploration and innovation in this field.

📄 PDF Abstract BibTeX arXiv:2405.16640

Code (1)

beccabai/Data-centric_multimodal_LLM 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelSurvey

Similar Papers 제목 키워드 기반

Large Language Models Meet Text-Centric Multimodal Sentiment Analysis: A Survey

2024-06-12 · Hao Yang, Yanyan Zhao, Yang Wu, Shilong Wang 외

Compared to traditional sentiment analysis, which only considers text, multimodal sentiment analysis needs to consider emotional signals from multimodal sources simultaneously and is therefore more consistent with the wa…

Multimodal Sentiment AnalysisSentiment Analysis

Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks

2025-10-29 · Xu Zheng, Zihao Dongfang, Lutao Jiang, Boyuan Zheng 외 arxiv

Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend these abilities by learning to perceive…

Vision-Language NavigationVisual Question AnsweringMultimodal ReasoningSpatial Reasoning

A Survey of Token Compression for Efficient Multimodal Large Language Models

2025-07-27 · Kele Shao, Keda Tao, Kejia Zhang, Sicheng Feng 외 arxiv

Multimodal large language models (MLLMs) have made remarkable strides, largely driven by their ability to process increasingly long and complex contexts, such as high-resolution images, extended video sequences, and leng…

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

2026-06-24 · Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao 외 arxiv

Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepS…

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

2025-06-05 · Jisu An, Junseok Lee, Jeoungeun Lee, Yongseok Son

The rapid progress of Multimodal Large Language Models(MLLMs) has transformed the AI landscape. These models combine pre-trained LLMs with various modality encoders. This integration requires a systematic understanding o…

cross-modal alignmentLarge Language ModelMultimodal Large Language ModelRepresentation Learning