paper-with-me

홈 › Papers

Visual Grounding in Video for Unsupervised Word Translation

2020-03-11 · CVPR 2020 6 · Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh, Lucas Smaira, Mateusz Malinowski, João Carreira, Phil Blunsom, Andrew Zisserman

There are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to improve unsupervised word mapping between languages. The key idea is to establish a common visual representation between two languages by learning embeddings from unpaired instructional videos narrated in the native language. Given this shared embedding we demonstrate that (i) we can map words between the languages, particularly the 'visual' words; (ii) that the shared embedding provides a good initialization for existing unsupervised text-based word translation techniques, forming the basis for our proposed hybrid visual-text mapping algorithm, MUVE; and (iii) our approach achieves superior performance by addressing the shortcomings of text-based methods -- it is more robust, handles datasets with less commonality, and is applicable to low-resource languages. We apply these methods to translate words from English to French, Korean, and Japanese -- all without any parallel corpora and simply by watching many videos of people speaking while doing things.

📄 PDF Abstract BibTeX arXiv:2003.05078

Code (1)

gsig/visual-grounding 공식 구현

Tasks

TranslationVisual GroundingWord Translation

Similar Papers 제목 키워드 기반

Grounded Word Sense Translation

2019-06-01 · WS 2019 6 · Chiraag Lala, Pranava Madhyastha, Lucia Specia

Recent work on visually grounded language learning has focused on broader applications of grounded representations, such as visual question answering and multimodal machine translation. In this paper we consider grounded…

Grounded language learningMachine TranslationMultimodal Machine TranslationQuestion Answering+3

Video-to-Video Translation for Visual Speech Synthesis

2019-05-28 · Michail C. Doukas, Viktoriia Sharmanska, Stefanos Zafeiriou

Despite remarkable success in image-to-image translation that celebrates the advancements of generative adversarial networks (GANs), very limited attempts are known for video domain translation. We study the task of vide…

Image-to-Image TranslationSpeech SynthesisTranslation

Natural Language Grounding and Grammar Induction for Robotic Manipulation Commands

2017-08-01 · WS 2017 8 · Muhannad Alomari, Paul Duckworth, Majd Hawasly, David C. Hogg 외

We present a cognitively plausible system capable of acquiring knowledge in language and vision from pairs of short video clips and linguistic descriptions. The aim of this work is to teach a robot manipulator how to exe…

Weakly-Supervised Generation and Grounding of Visual Descriptions With Conditional Generative Models

2022-01-01 · CVPR 2022 1 · Effrosyni Mavroudi, René Vidal

Given weak supervision from image- or video-caption pairs, we address the problem of grounding (localizing) each object word of a ground-truth or generated sentence describing a visual input. Recent weakly-supervised…

Sentence

Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps Grounding

2020-12-18 · Dexin Wang, Deyi Xiong

Visual context provides grounding information for multimodal machine translation (MMT). However, previous MMT models and probing studies on visual features suggest that visual information is less explored in MMT as it is…

Machine TranslationMultimodal Machine TranslationObjectTranslation