paper-with-me

홈 › Papers

Towards Visually Grounded Sub-Word Speech Unit Discovery

2019-02-21 · David Harwath, James Glass

In this paper, we investigate the manner in which interpretable sub-word speech units emerge within a convolutional neural network model trained to associate raw speech waveforms with semantically related natural image scenes. We show how diphone boundaries can be superficially extracted from the activation patterns of intermediate layers of the model, suggesting that the model may be leveraging these events for the purpose of word recognition. We present a series of experiments investigating the information encoded by these events.

📄 PDF Abstract BibTeX arXiv:1902.08213

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Word Discovery in Visually Grounded, Self-Supervised Speech Models

2022-03-28 · Puyuan Peng, David Harwath

We present a method for visually-grounded spoken term discovery. After training either a HuBERT or wav2vec2.0 model to associate spoken captions with natural images, we show that powerful word segmentation and clustering…

ClusteringSegmentationVisual Grounding

Modelling word learning and recognition using visually grounded speech

2022-03-14 · Danny Merkx, Sebastiaan Scholten, Stefan L. Frank, Mirjam Ernestus 외

Background: Computational models of speech recognition often assume that the set of target words is already given. This implies that these models do not learn to recognise speech from scratch without prior knowledge and …

Representation Learningspeech-recognitionSpeech Recognition

Syllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model

2023-05-19 · Puyuan Peng, Shang-Wen Li, Okko Räsänen, Abdelrahman Mohamed 외

In this paper, we show that representations capturing syllabic units emerge when training a self-supervised speech model with a visually-grounded training objective. We demonstrate that a nearly identical model architect…

Language ModelingLanguage ModellingMasked Language ModelingSegmentation+2

Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech

2019-11-21 · ICLR 2020 1 · David Harwath, Wei-Ning Hsu, James Glass

In this paper, we present a method for learning discrete linguistic units by incorporating vector quantization layers into neural models of visually grounded speech. We show that our method is capable of capturing both w…

Image RetrievalQuantizationRetrieval

Models of Visually Grounded Speech Signal Pay Attention To Nouns: a Bilingual Experiment on English and Japanese

2019-02-08 · William N. Havard, Jean-Pierre Chevrot, Laurent Besacier

We investigate the behaviour of attention in neural models of visually grounded speech trained on two languages: English and Japanese. Experimental results show that attention focuses on nouns and this behaviour holds tr…

Retrieval