paper-with-me

Papers

Multimodal Intelligence: Representation Learning, Information Fusion, and Applications

2019-11-10 · Chao Zhang, Zichao Yang, Xiaodong He, Li Deng

Deep learning methods have revolutionized speech recognition, image recognition, and natural language processing since 2010. Each of these tasks involves a single modality in their input signals. However, many applications in the artificial intelligence field involve multiple modalities. Therefore, it is of broad interest to study the more difficult and complex problem of modeling and learning across multiple modalities. In this paper, we provide a technical review of available models and learning methods for multimodal intelligence. The main focus of this review is the combination of vision and natural language modalities, which has become an important topic in both the computer vision and natural language processing research communities. This review provides a comprehensive analysis of recent works on multimodal deep learning from three perspectives: learning multimodal representations, fusing multimodal signals at various levels, and multimodal applications. Regarding multimodal representation learning, we review the key concepts of embedding, which unify multimodal signals into a single vector space and thereby enable cross-modality signal processing. We also review the properties of many types of embeddings that are constructed and learned for general downstream tasks. Regarding multimodal fusion, this review focuses on special architectures for the integration of representations of unimodal signals for a particular task. Regarding applications, selected areas of a broad interest in the current literature are covered, including image-to-text caption generation, text-to-image generation, and visual question answering. We believe that this review will facilitate future studies in the emerging field of multimodal intelligence for related communities.

📄 PDF Abstract BibTeX arXiv:1911.03977

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationImage GenerationImage to textMultimodal Deep LearningQuestion AnsweringRepresentation Learningspeech-recognitionSpeech RecognitionText to Image GenerationText-to-Image GenerationVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Multimodal Machine Learning: A Survey and Taxonomy

2017-05-26 · Tadas Baltrušaitis, Chaitanya Ahuja, Louis-Philippe Morency

Our experience of the world is multimodal - we see objects, hear sounds, feel texture, smell odors, and taste flavors. Modality refers to the way in which something happens or is experienced and a research problem is cha…

BIG-bench Machine LearningSurveyTranslation

MolFusion: Multimodal Fusion Learning for Molecular Representations via Multi-granularity Views

2024-06-26 · MuZhen Cai, Sendong Zhao, Haochun Wang, Yanrui Du 외

Artificial Intelligence predicts drug properties by encoding drug molecules, aiding in the rapid screening of candidates. Different molecular representations, such as SMILES and molecule graphs, contain complementary inf…

DyMRL: Dynamic Multispace Representation Learning for Multimodal Event Forecasting in Knowledge Graph

2026-03-25 · Feng Zhao, Kangzheng Liu, Teng Peng, Yu Yang 외 arxiv

Accurate representation of multimodal knowledge is crucial for event forecasting in real-world scenarios. However, existing studies have largely focused on static settings, overlooking the dynamic acquisition and fusion …

Representation LearningLogical Reasoning

Brainish: Formalizing A Multimodal Language for Intelligence and Consciousness

2022-04-14 · Paul Pu Liang

Having a rich multimodal inner language is an important component of human intelligence that enables several necessary core cognitive functions such as multimodal prediction, translation, and generation. Building upon th…

RetrievalTranslation

Artificial Intelligence-Based Methods for Fusion of Electronic Health Records and Imaging Data

2022-10-23 · Farida Mohsen, Hazrat Ali, Nady El Hajj, Zubair Shah

Healthcare data are inherently multimodal, including electronic health records (EHR), medical images, and multi-omics data. Combining these multimodal data sources contributes to a better understanding of human health an…