Multimodal Representation Learning With Text and Images
In recent years, multimodal AI has seen an upward trend as researchers are integrating data of different types such as text, images, speech into modelling to get the best results. This project leverages multimodal AI and matrix factorization techniques for representation learning, on text and image data simultaneously, thereby employing the widely used techniques of Natural Language Processing (NLP) and Computer Vision. The learnt representations are evaluated using downstream classification and regression tasks. The methodology adopted can be extended beyond the scope of this project as it uses Auto-Encoders for unsupervised representation learning.
Code (1)
Tasks
Representation LearningSimilar Papers 제목 키워드 기반
Mr. Right: Multimodal Retrieval on Representation of ImaGe witH Text
Multimodal learning is a recent challenge that extends unimodal learning by generalizing its domain to diverse modalities, such as texts, images, or speech. This extension requires models to process and relate informatio…
Image CaptioningImage RetrievalImage-text RetrievalInformation Retrieval+2MM-Rec: Multimodal News Recommendation
Accurate news representation is critical for news recommendation. Most of existing news representation methods learn news representations only from news texts while ignore the visual information in news like images. In f…
News Recommendationobject-detectionObject DetectionMultimodal Logical Inference System for Visual-Textual Entailment
A large amount of research about multimodal inference across text and vision has been recently developed to obtain visually grounded word and sentence representations. In this paper, we use logic-based representations as…
Automated Theorem ProvingNatural Language InferenceSemantic ParsingSentenceMOTIF: Contextualized Images for Complex Words to Improve Human Reading
MOTIF (MultimOdal ConTextualized Images For Language Learners) is a multimodal dataset that consists of 1125 comprehension texts retrieved from Wikipedia Simple Corpus. Allowing multimodal processing or enriching the con…
Reading ComprehensionAccurate Word Representations with Universal Visual Guidance
Word representation is a fundamental component in neural language understanding models. Recently, pre-trained language models (PrLMs) offer a new performant method of contextualized word representations by leveraging the…
Machine TranslationNatural Language UnderstandingTranslation