Imagination improves Multimodal Translation
We decompose multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. In a multitask learning framework, translations are learned in an attention-based encoder-decoder, and grounded representations are learned through image representation prediction. Our approach improves translation performance compared to the state of the art on the Multi30K dataset. Furthermore, it is equally effective if we train the image prediction task on the external MS COCO dataset, and we find improvements if we train the translation model on the external News Commentary parallel text.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderPredictionTranslationSimilar Papers 제목 키워드 기반
Generative Imagination Elevates Machine Translation
There are common semantics shared across text and images. Given a sentence in a source language, whether depicting the visual scene helps translation into a target language? Existing multimodal neural machine translation…
Machine TranslationMultimodal Machine TranslationSentenceTransfer Learning+1Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
Visual information has been introduced for enhancing machine translation (MT), and its effectiveness heavily relies on the availability of large amounts of bilingual parallel sentence pairs with manual image annotations.…
Language ModelingLanguage ModellingLarge Language ModelMachine Translation+3SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, r…
multimodal generationSpatial ReasoningMissing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities
Multimodal fusion has been proved to improve emotion recognition performance in previous works. However, in real-world applications, we often encounter the problem of missing modality, and which modalities will be missin…
Emotion RecognitionExploiting modality-invariant feature for robust multimodal emotion recognition with missing modalities
Multimodal emotion recognition leverages complementary information across modalities to gain performance. However, we cannot guarantee that the data of all modalities are always present in practice. In the studies to pre…
Emotion RecognitionMultimodal Emotion Recognition