Towards Visually Grounded Sub-Word Speech Unit Discovery
In this paper, we investigate the manner in which interpretable sub-word speech units emerge within a convolutional neural network model trained to associate raw speech waveforms with semantically related natural image scenes. We show how diphone boundaries can be superficially extracted from the activation patterns of intermediate layers of the model, suggesting that the model may be leveraging these events for the purpose of word recognition. We present a series of experiments investigating the information encoded by these events.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Word Discovery in Visually Grounded, Self-Supervised Speech Models
We present a method for visually-grounded spoken term discovery. After training either a HuBERT or wav2vec2.0 model to associate spoken captions with natural images, we show that powerful word segmentation and clustering…
ClusteringSegmentationVisual GroundingModelling word learning and recognition using visually grounded speech
Background: Computational models of speech recognition often assume that the set of target words is already given. This implies that these models do not learn to recognise speech from scratch without prior knowledge and …
Representation Learningspeech-recognitionSpeech RecognitionSyllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model
In this paper, we show that representations capturing syllabic units emerge when training a self-supervised speech model with a visually-grounded training objective. We demonstrate that a nearly identical model architect…
Language ModelingLanguage ModellingMasked Language ModelingSegmentation+2Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
In this paper, we present a method for learning discrete linguistic units by incorporating vector quantization layers into neural models of visually grounded speech. We show that our method is capable of capturing both w…
Image RetrievalQuantizationRetrievalModels of Visually Grounded Speech Signal Pay Attention To Nouns: a Bilingual Experiment on English and Japanese
We investigate the behaviour of attention in neural models of visually grounded speech trained on two languages: English and Japanese. Experimental results show that attention focuses on nouns and this behaviour holds tr…
Retrieval