paper-with-me

홈 › Papers

Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech

2019-11-21 · ICLR 2020 1 · David Harwath, Wei-Ning Hsu, James Glass

In this paper, we present a method for learning discrete linguistic units by incorporating vector quantization layers into neural models of visually grounded speech. We show that our method is capable of capturing both word-level and sub-word units, depending on how it is configured. What differentiates this paper from prior work on speech unit learning is the choice of training objective. Rather than using a reconstruction-based loss, we use a discriminative, multimodal grounding objective which forces the learned units to be useful for semantic image retrieval. We evaluate the sub-word units on the ZeroSpeech 2019 challenge, achieving a 27.3\% reduction in ABX error rate over the top-performing submission, while keeping the bitrate approximately the same. We also present experiments demonstrating the noise robustness of these units. Finally, we show that a model with multiple quantizers can simultaneously learn phone-like detectors at a lower layer and word-like detectors at a higher layer. We show that these detectors are highly accurate, discovering 279 words with an F1 score of greater than 0.5.

📄 PDF Abstract BibTeX arXiv:1911.09602

Code (1)

DannyMerkx/speech2image pytorch

Tasks

Image RetrievalQuantizationRetrieval

Similar Papers 제목 키워드 기반

Word Recognition, Competition, and Activation in a Model of Visually Grounded Speech

2019-09-18 · CONLL 2019 11 · William N. Havard, Jean-Pierre Chevrot, Laurent Besacier

In this paper, we study how word-like units are represented and activated in a recurrent neural model of visually grounded speech. The model used in our experiments is trained to project an image and its spoken descripti…

Catplayinginthesnow: Impact of Prior Segmentation on a Model of Visually Grounded Speech

2020-06-15 · CONLL 2020 · William N. Havard, Jean-Pierre Chevrot, Laurent Besacier

The language acquisition literature shows that children do not build their lexicon by segmenting the spoken input into phonemes and then building up words from them, but rather adopt a top-down approach and start by segm…

Image RetrievalLanguage AcquisitionRetrieval

GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion

2026-04-23 · Qizhuo Xie, Yunhui Liu, Yu Xing, Qianzi Hou 외 arxiv

Large Language Models (LLMs) have shown immense potential in Knowledge Graph Completion (KGC), yet bridging the modality gap between continuous graph embeddings and discrete LLM tokens remains a critical challenge. While…

Knowledge Graph Completion

A Linguistic Analysis of Visually Grounded Dialogues Based on Spatial Expressions

2020-10-07 · Findings of the Association for Computational Linguistics 2020 · Takuma Udagawa, Takato Yamazaki, Akiko Aizawa

Recent models achieve promising results in visually grounded dialogues. However, existing datasets often contain undesirable biases and lack sophisticated linguistic analyses, which make it difficult to understand how we…

Coreference ResolutionNatural Language Visual GroundingSpatial Relation Recognition

Semantic Composition in Visually Grounded Language Models

2023-05-15 · Rohan Pandey

What is sentence meaning and its ideal representation? Much of the expressive power of human language derives from semantic composition, the mind's ability to represent meaning hierarchically & relationally over constitu…

Image CaptioningInductive BiasPhilosophyQuestion Answering+6