paper-with-me

홈 › Papers

Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities

2024-06-08 · Sai Munikoti, Ian Stewart, Sameera Horawalavithana, Henry Kvinge, Tegan Emerson, Sandra E Thompson, Karl Pazdernik

Multimodal models are expected to be a critical component to future advances in artificial intelligence. This field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural language processing (NLP) and vision. It is widely hoped that further extending the foundation models to multiple modalities (e.g., text, image, video, sensor, time series, graph, etc.) will ultimately lead to generalist multimodal models, i.e. one model across different data modalities and tasks. However, there is little research that systematically analyzes recent multimodal models (particularly the ones that work beyond text and vision) with respect to the underling architecture proposed. Therefore, this work provides a fresh perspective on generalist multimodal models (GMMs) via a novel architecture and training configuration specific taxonomy. This includes factors such as Unifiability, Modularity, and Adaptability that are pertinent and essential to the wide adoption and application of GMMs. The review further highlights key challenges and prospects for the field and guide the researchers into the new advancements.

📄 PDF Abstract BibTeX arXiv:2406.05496

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Toward Generalist Neural Motion Planners for Robotic Manipulators: Challenges and Opportunities

2026-03-25 · Davood Soleymanzadeh, Ivan Lopez-Sanchez, Hao Su, Yunzhu Li 외 arxiv

State-of-the-art generalist manipulation policies have enabled the deployment of robotic manipulators in unstructured human environments. However, these frameworks struggle in cluttered environments primarily because the…

Motion Planning

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

2025-07-02 · Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

Crash detection from video feeds is a critical problem in intelligent transportation systems. Recent developments in large language models (LLMs) and vision-language models (VLMs) have transformed how we process, reason …

Video Understanding

Medical Multimodal Foundation Models in Clinical Diagnosis and Treatment: Applications, Challenges, and Future Directions

2024-12-03 · Kai Sun, Siyan Xue, Fuchun Sun, Haoran Sun 외

Recent advancements in deep learning have significantly revolutionized the field of clinical diagnosis and treatment, offering novel approaches to improve diagnostic precision and treatment efficacy across diverse clinic…

Diagnostic

Deep Learning Advances in Vision-Based Traffic Accident Anticipation: A Comprehensive Review of Methods,Datasets,and Future Directions

2025-05-12 · Yi Zhang, Wenye Zhou, Ruonan Lin, Xin Yang 외

Traffic accident prediction and detection are critical for enhancing road safety,and vision-based traffic accident anticipation (Vision-TAA) has emerged as a promising approach in the era of deep learning.This paper revi…

Accident AnticipationPredictionScene UnderstandingSelf-Supervised Learning

Towards Generalist Robot Learning from Internet Video: A Survey

2024-04-30 · Robert McCarthy, Daniel C. H. Tan, Dominik Schmidt, Fernando Acero 외

Scaling deep learning to massive, diverse internet data has yielded remarkably general capabilities in visual and natural language understanding and generation. However, data has remained scarce and challenging to collec…

Natural Language UnderstandingReinforcement Learning (RL)Survey