paper-with-me

홈 › Papers

A Comprehensive Review of Multimodal Large Language Models: Performance and Challenges Across Different Tasks

2024-08-02

In an era defined by the explosive growth of data and rapid technological advancements, Multimodal Large Language Models (MLLMs) stand at the forefront of artificial intelligence (AI) systems. Designed to seamlessly integrate diverse data types-including text, images, videos, audio, and physiological sequences-MLLMs address the complexities of real-world applications far beyond the capabilities of single-modality systems. In this paper, we systematically sort out the applications of MLLM in multimodal tasks such as natural language, vision, and audio. We also provide a comparative analysis of the focus of different MLLMs in the tasks, and provide insights into the shortcomings of current MLLMs, and suggest potential directions for future research. Through these discussions, this paper hopes to provide valuable insights for the further development and application of MLLM.

📄 PDF Abstract BibTeX arXiv:2408.01319

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation

2025-08-19 · Xian Gao, Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao 외 arxiv

With the rapid growth of academic publications, peer review has become an essential yet time-consuming responsibility within the research community. Large Language Models (LLMs) have increasingly been adopted to assist i…

Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review

2025-02-23 · Pei Fu, Tongkun Guan, Zining Wang, Zhentao Guo 외

The recent emergence of Multi-modal Large Language Models (MLLMs) has introduced a new dimension to the Text-rich Image Understanding (TIU) field, with models demonstrating impressive and inspiring performance. However, …

Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing

2024-09-26 · Huthaifa I. Ashqar, Ahmed Jaber, Taqwa I. Alhadidi, Mohammed Elhenawy

This study aims to comprehensively review and empirically evaluate the application of multimodal large language models (MLLMs) and Large Vision Models (VLMs) in object detection for transportation systems. In the first f…

Event DetectionObjectobject-detectionObject Detection+1

A Survey of Multimodal Large Language Model from A Data-centric Perspective

2024-05-26 · Tianyi Bai, Hao Liang, Binwang Wan, Yanran Xu 외

Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. …

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks

2025-10-29 · Xu Zheng, Zihao Dongfang, Lutao Jiang, Boyuan Zheng 외 arxiv

Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend these abilities by learning to perceive…

Vision-Language NavigationVisual Question AnsweringMultimodal ReasoningSpatial Reasoning