paper-with-me

Papers

Multimodal Representation Learning With Text and Images

2022-04-30 · Aishwarya Jayagopal, Ankireddy Monica Aiswarya, Ankita Garg, Srinivasan Kolumam Nandakumar

In recent years, multimodal AI has seen an upward trend as researchers are integrating data of different types such as text, images, speech into modelling to get the best results. This project leverages multimodal AI and matrix factorization techniques for representation learning, on text and image data simultaneously, thereby employing the widely used techniques of Natural Language Processing (NLP) and Computer Vision. The learnt representations are evaluated using downstream classification and regression tasks. The methodology adopted can be extended beyond the scope of this project as it uses Auto-Encoders for unsupervised representation learning.

📄 PDF Abstract BibTeX arXiv:2205.00142

Code (1)

srini-98/cs5260-neural-networks-2 공식 구현

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Mr. Right: Multimodal Retrieval on Representation of ImaGe witH Text

2022-09-28 · Cheng-An Hsieh, Cheng-Ping Hsieh, Pu-Jen Cheng

Multimodal learning is a recent challenge that extends unimodal learning by generalizing its domain to diverse modalities, such as texts, images, or speech. This extension requires models to process and relate informatio…

Image CaptioningImage RetrievalImage-text RetrievalInformation Retrieval+2

MM-Rec: Multimodal News Recommendation

2021-04-15 · Chuhan Wu, Fangzhao Wu, Tao Qi, Yongfeng Huang

Accurate news representation is critical for news recommendation. Most of existing news representation methods learn news representations only from news texts while ignore the visual information in news like images. In f…

News Recommendationobject-detectionObject Detection

Multimodal Logical Inference System for Visual-Textual Entailment

2019-06-10 · ACL 2019 7 · Riko Suzuki, Hitomi Yanaka, Masashi Yoshikawa, Koji Mineshima 외

A large amount of research about multimodal inference across text and vision has been recently developed to obtain visually grounded word and sentence representations. In this paper, we use logic-based representations as…

Automated Theorem ProvingNatural Language InferenceSemantic ParsingSentence

MOTIF: Contextualized Images for Complex Words to Improve Human Reading

2022-06-01 · LREC 2022 6 · Xintong Wang, Florian Schneider, Özge Alacam, Prateek Chaudhury 외

MOTIF (MultimOdal ConTextualized Images For Language Learners) is a multimodal dataset that consists of 1125 comprehension texts retrieved from Wikipedia Simple Corpus. Allowing multimodal processing or enriching the con…

Reading Comprehension

Accurate Word Representations with Universal Visual Guidance

2020-12-30 · Zhuosheng Zhang, Haojie Yu, Hai Zhao, Rui Wang 외

Word representation is a fundamental component in neural language understanding models. Recently, pre-trained language models (PrLMs) offer a new performant method of contextualized word representations by leveraging the…

Machine TranslationNatural Language UnderstandingTranslation