Adversarial reconstruction for Multi-modal Machine Translation
Even with the growing interest in problems at the intersection of Computer Vision and Natural Language, grounding (i.e. identifying) the components of a structured description in an image still remains a challenging task. This contribution aims to propose a model which learns grounding by reconstructing the visual features for the Multi-modal translation task. Previous works have partially investigated standard approaches such as regression methods to approximate the reconstruction of a visual input. In this paper, we propose a different and novel approach which learns grounding by adversarial feedback. To do so, we modulate our network following the recent promising adversarial architectures and evaluate how the adversarial response from a visual reconstruction as an auxiliary task helps the model in its learning. We report the highest scores in term of BLEU and METEOR metrics on the different datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationSimilar Papers 제목 키워드 기반
Adversarial Evaluation of Multimodal Machine Translation
The promise of combining language and vision in multimodal machine translation is that systems will produce better translations by leveraging the image data. However, the evidence surrounding whether the images are usefu…
Machine TranslationMultimodal Machine Translationtext similarityTranslationEntity-level Cross-modal Learning Improves Multi-modal Machine Translation
Multi-modal machine translation (MMT) aims at improving translation performance by incorporating visual information. Most of the studies leverage the visual information through integrating the global image features as au…
Machine TranslationRepresentation LearningTranslationDeterministic Medical Image Translation via High-fidelity Brownian Bridges
Recent studies have shown that diffusion models produce superior synthetic images when compared to Generative Adversarial Networks (GANs). However, their outputs are often non-deterministic and lack high fidelity to the …
Image Super-ResolutionSuper-ResolutionTranslationTowards Multimodal Simultaneous Neural Machine Translation
Simultaneous translation involves translating a sentence before the speaker's utterance is completed in order to realize real-time understanding in multiple languages. This task is significantly more challenging than the…
Machine TranslationSentenceTranslationSemi-Supervised Image-to-Image Translation
Image-to-image translation is a long-established and a difficult problem in computer vision. In this paper we propose an adversarial based model for image-to-image translation. The regular deep neural-network based metho…
Generative Adversarial NetworkImage SegmentationImage-to-Image TranslationMultimodal Unsupervised Image-To-Image Translation+4