Probing the Need for Visual Context in Multimodal Machine Translation
Current work on multimodal machine translation (MMT) has suggested that the visual modality is either unnecessary or only marginally beneficial. We posit that this is a consequence of the very simple, short and repetitive sentences used in the only available dataset for the task (Multi30K), rendering the source text sufficient as context. In the general case, however, we believe that it is possible to combine visual and textual information in order to ground translations. In this paper we probe the contribution of the visual modality to state-of-the-art MMT models by conducting a systematic analysis where we partially deprive the models from source-side textual context. Our results show that under limited textual context, models are capable of leveraging the visual input to generate better translations. This contradicts the current belief that MMT models disregard the visual modality because of either the quality of the image features or the way they are integrated into the model.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationMultimodal Machine TranslationTranslationSimilar Papers 제목 키워드 기반
Is BERT Blind? Exploring the Effect of Vision-and-Language Pretraining on Visual Language Understanding
Most humans use visual imagination to understand and reason about language, but models such as BERT reason about language using knowledge acquired during text-only pretraining. In this work, we investigate whether vision…
Knowledge ProbingLanguage ModellingNatural Language UnderstandingVisual ReasoningEfficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps Grounding
Visual context provides grounding information for multimodal machine translation (MMT). However, previous MMT models and probing studies on visual features suggest that visual information is less explored in MMT as it is…
Machine TranslationMultimodal Machine TranslationObjectTranslationProbing Multimodal Embeddings for Linguistic Properties: the Visual-Semantic Case
Semantic embeddings have advanced the state of the art for countless natural language processing tasks, and various extensions to multimodal domains, such as visual-semantic embeddings, have been proposed. While the powe…
Incorporating Probing Signals into Multimodal Machine Translation via Visual Question-Answering Pairs
This paper presents an in-depth study of multimodal machine translation (MMT), examining the prevailing understanding that MMT systems exhibit decreased sensitivity to visual information when text inputs are complete. In…
AttributeMachine TranslationMultimodal Machine TranslationQuestion Answering+2Probing Contextualized Sentence Representations with Visual Awareness
We present a universal framework to model contextualized sentence representations with visual awareness that is motivated to overcome the shortcomings of the multimodal parallel data with manual annotations. For each sen…
DiversityMachine TranslationNatural Language InferenceSentence+1