paper-with-me

홈 › Papers

Word Recognition, Competition, and Activation in a Model of Visually Grounded Speech

2019-09-18 · CONLL 2019 11 · William N. Havard, Jean-Pierre Chevrot, Laurent Besacier

In this paper, we study how word-like units are represented and activated in a recurrent neural model of visually grounded speech. The model used in our experiments is trained to project an image and its spoken description in a common representation space. We show that a recurrent model trained on spoken sentences implicitly segments its input into word-like units and reliably maps them to their correct visual referents. We introduce a methodology originating from linguistics to analyse the representation learned by neural networks -- the gating paradigm -- and show that the correct representation of a word is only activated if the network has access to first phoneme of the target word, suggesting that the network does not rely on a global acoustic pattern. Furthermore, we find out that not all speech frames (MFCC vectors in our case) play an equal role in the final encoded representation of a given word, but that some frames have a crucial effect on it. Finally, we suggest that word representation could be activated through a process of lexical competition.

📄 PDF Abstract BibTeX arXiv:1909.08491

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning to Recognise Words using Visually Grounded Speech

2020-05-31 · Sebastiaan Scholten, Danny Merkx, Odette Scharenborg

We investigated word recognition in a Visually Grounded Speech model. The model has been trained on pairs of images and spoken captions to create visually grounded embeddings which can be used for speech to image retriev…

Image RetrievalRetrieval

Modelling word learning and recognition using visually grounded speech

2022-03-14 · Danny Merkx, Sebastiaan Scholten, Stefan L. Frank, Mirjam Ernestus 외

Background: Computational models of speech recognition often assume that the set of target words is already given. This implies that these models do not learn to recognise speech from scratch without prior knowledge and …

Representation Learningspeech-recognitionSpeech Recognition

Towards Visually Grounded Sub-Word Speech Unit Discovery

2019-02-21 · David Harwath, James Glass

In this paper, we investigate the manner in which interpretable sub-word speech units emerge within a convolutional neural network model trained to associate raw speech waveforms with semantically related natural image s…

Language learning using Speech to Image retrieval

2019-09-09 · Danny Merkx, Stefan L. Frank, Mirjam Ernestus

Humans learn language by interaction with their environment and listening to other humans. It should also be possible for computational models to learn language directly from speech but so far most approaches require tex…

Grounded language learningImage RetrievalRetrievalSentence+3

Seeing the advantage: visually grounding word embeddings to better capture human semantic knowledge

2022-02-21 · CMCL (ACL) 2022 5 · Danny Merkx, Stefan L. Frank, Mirjam Ernestus

Distributional semantic models capture word-level meaning that is useful in many natural language processing tasks and have even been shown to capture cognitive aspects of word meaning. The majority of these models are p…

Grounded language learningImage RetrievalLearning Semantic RepresentationsVisual Grounding+2