paper-with-me

홈 › Papers

Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions

2022-09-07 · Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency

Multimodal machine learning is a vibrant multi-disciplinary research field that aims to design computer agents with intelligent capabilities such as understanding, reasoning, and learning through integrating multiple communicative modalities, including linguistic, acoustic, visual, tactile, and physiological messages. With the recent interest in video understanding, embodied autonomous agents, text-to-image generation, and multisensor fusion in application domains such as healthcare and robotics, multimodal machine learning has brought unique computational and theoretical challenges to the machine learning community given the heterogeneity of data sources and the interconnections often found between modalities. However, the breadth of progress in multimodal research has made it difficult to identify the common themes and open questions in the field. By synthesizing a broad range of application domains and theoretical frameworks from both historical and recent perspectives, this paper is designed to provide an overview of the computational and theoretical foundations of multimodal machine learning. We start by defining three key principles of modality heterogeneity, connections, and interactions that have driven subsequent innovations, and propose a taxonomy of six core technical challenges: representation, alignment, reasoning, generation, transference, and quantification covering historical and recent trends. Recent technical achievements will be presented through the lens of this taxonomy, allowing researchers to understand the similarities and differences across new approaches. We end by motivating several open problems for future research as identified by our taxonomy.

📄 PDF Abstract BibTeX arXiv:2209.03430

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image GenerationVideo Understanding

Similar Papers 제목 키워드 기반

Evolution of Video Generative Foundations

2026-04-07 · Teng Hu, Jiangning Zhang, Hongrui Huang, Ran Yi 외 arxiv

The rapid advancement of Artificial Intelligence Generated Content (AIGC) has revolutionized video generation, enabling systems ranging from proprietary pioneers like OpenAI's Sora, Google's Veo3, and Bytedance's Seedanc…

Autonomous DrivingVideo Generation

Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

2024-11-04 · Biao Wu, Yanda Li, Meng Fang, Zirui Song 외

Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and process multimodal data have grown. This su…

multimodal interactionSurvey

Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions

2025-04-20 · Luyang Fang, Xiaowei Yu, Jiazhang Cai, Yongkai Chen 외

The exponential growth of Large Language Models (LLMs) continues to highlight the need for efficient strategies to meet ever-expanding computational and data demands. This survey provides a comprehensive analysis of two …

Dataset DistillationDiversityKnowledge DistillationSurvey

Efficient Generalization via Multimodal Co-Training under Data Scarcity and Distribution Shift

2025-10-08 · Tianyu Bell Pan, Damon L. Woodard arxiv

This paper explores a multimodal co-training framework designed to enhance model generalization in situations where labeled data is limited and distribution shifts occur. We thoroughly examine the theoretical foundations…

Exploring the landscape of large language models: Foundations, techniques, and challenges

2024-04-18 · Milad Moradi, Ke Yan, David Colwell, Matthias Samwald 외

In this review paper, we delve into the realm of Large Language Models (LLMs), covering their foundational principles, diverse applications, and nuanced training processes. The article sheds light on the mechanics of in-…

In-Context LearningRetrievalRetrieval-augmented Generation