Doubly-Attentive Decoder for Multi-modal Neural Machine Translation
We introduce a Multi-modal Neural Machine Translation model in which a doubly-attentive decoder naturally incorporates spatial visual features obtained using pre-trained convolutional neural networks, bridging the gap between image description and translation. Our decoder learns to attend to source-language words and parts of an image independently by means of two separate attention mechanisms as it generates words in the target language. We find that our model can efficiently exploit not just back-translated in-domain multi-modal data but also large general-domain text-only MT corpora. We also report state-of-the-art results on the Multi30k data set.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderImage DescriptionMachine TranslationMultimodal Machine TranslationTranslationSimilar Papers 제목 키워드 기반
Doubly Attentive Transformer Machine Translation
In this paper a doubly attentive transformer machine translation model (DATNMT) is presented in which a doubly-attentive transformer decoder normally joins spatial visual features obtained via pretrained convolutional ne…
DecoderImage CaptioningMachine TranslationMultimodal Machine Translation+1GLEAM: A Multimodal Imaging Dataset and HAMM for Glaucoma Classification
We propose glaucoma lesion evaluation and analysis with multimodal imaging (GLEAM), the first publicly available tri-modal glaucoma dataset comprising scanning laser ophthalmoscopy fundus images, circumpapillary OCT imag…
Representation LearningNeural network architectures for disentangling the multimodal structure of data ensembles
We introduce neural network architectures that model the mechanism that generates data and address the difficult problem of disentangling the multimodal structure of data ensembles. We provide (i) an autoencoder-decoder …
DecoderDoubly Residual Neural Decoder: Towards Low-Complexity High-Performance Channel Decoding
Recently deep neural networks have been successfully applied in channel coding to improve the decoding performance. However, the state-of-the-art neural channel decoders cannot achieve high decoding performance and low c…
DecoderVocal Bursts Intensity PredictionAttentive Cross-modal Connections for Deep Multimodal Wearable-based Emotion Recognition
Classification of human emotions can play an essential role in the design and improvement of human-machine systems. While individual biological signals such as Electrocardiogram (ECG) and Electrodermal Activity (EDA) hav…
ClassificationEmotion ClassificationEmotion Recognition