Exploration on Grounded Word Embedding: Matching Words and Images with Image-Enhanced Skip-Gram Model
Word embedding is designed to represent the semantic meaning of a word with low dimensional vectors. The state-of-the-art methods of learning word embeddings (word2vec and GloVe) only use the word co-occurrence information. The learned embeddings are real number vectors, which are obscure to human. In this paper, we propose an Image-Enhanced Skip-Gram Model to learn grounded word embeddings by representing the word vectors in the same hyper-plane with image vectors. Experiments show that the image vectors and word embeddings learned by our model are highly correlated, which indicates that our model is able to provide a vivid image-based explanation to the word embeddings.
Code (0)
등록된 구현이 없습니다.
Tasks
Learning Word EmbeddingsWord EmbeddingsSimilar Papers 제목 키워드 기반
Learning to Recognise Words using Visually Grounded Speech
We investigated word recognition in a Visually Grounded Speech model. The model has been trained on pairs of images and spoken captions to create visually grounded embeddings which can be used for speech to image retriev…
Image RetrievalRetrievalWordsurf : un outil pour naviguer dans un espace de « Word Embeddings » (Wordsurf : a tool to surf in a ``word embeddings'' space)
Dans cet article, nous pr{\'e}sentons un outil appel{\'e} « Wordsurf » pour faciliter la phase d{'}exploration et de navigation dans un espace de « Word Embeddings » pr{\'e}alablement entrain{\'e} sur des corpus de texte…
Word EmbeddingsAcoustically Grounded Word Embeddings for Improved Acoustics-to-Word Speech Recognition
Direct acoustics-to-word (A2W) systems for end-to-end automatic speech recognition are simpler to train, and more efficient to decode with, than sub-word systems. However, A2W systems can have difficulties at training ti…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1Language with Vision: a Study on Grounded Word and Sentence Embeddings
Grounding language in vision is an active field of research seeking to construct cognitively plausible word and sentence representations by incorporating perceptual knowledge from vision into text-based representations. …
SentenceSentence EmbeddingsVisual GroundingWord Embeddings+1Robust Backed-off Estimation of Out-of-Vocabulary Embeddings
Out-of-vocabulary (oov) words cause serious troubles in solving natural language tasks with a neural network. Existing approaches to this problem resort to using subwords, which are shorter and more ambiguous units than …
Word EmbeddingsWord Similarity