On Vision Features in Multimodal Machine Translation
Previous work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is on the quality of vision models. In this work, we investigate the impact of vision models on MMT. Given the fact that Transformer is becoming popular in computer vision, we experiment with various strong models (such as Vision Transformer) and enhanced features (such as object-detection and image captioning). We develop a selective attention model to study the patch-level contribution of an image in MMT. On detailed probing tasks, we find that stronger vision models are helpful for learning translation from the vision modality. Our results also suggest the need of carefully examining MMT models, especially when current benchmarks are small-scale and biased.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningMachine TranslationMultimodal Machine Translationobject-detectionObject DetectionTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
On Vision Features in Multimodal Machine Translation
Previous work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is on the quality of vision models. In this work, we investigate the imp…
Image CaptioningMachine TranslationMultimodal Machine Translationobject-detection+2Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps Grounding
Visual context provides grounding information for multimodal machine translation (MMT). However, previous MMT models and probing studies on visual features suggest that visual information is less explored in MMT as it is…
Machine TranslationMultimodal Machine TranslationObjectTranslationCLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine Translation
There has been a growing interest in developing multimodal machine translation (MMT) systems that enhance neural machine translation (NMT) with visual knowledge. This problem setup involves using images as auxiliary info…
Image CaptioningMachine TranslationMultimodal Machine TranslationNMT+1Improved English to Hindi Multimodal Neural Machine Translation
Machine translation performs automatic translation from one natural language to another. Neural machine translation attains a state-of-the-art approach in machine translation, but it requires adequate training data, whic…
Data AugmentationMachine TranslationNMTTranslationDynamic Context-guided Capsule Network for Multimodal Machine Translation
Multimodal machine translation (MMT), which mainly focuses on enhancing text-only translation with visual features, has attracted considerable attention from both computer vision and natural language processing communiti…
DecoderMachine TranslationMultimodal Machine TranslationRepresentation Learning+1