paper-with-me

홈 › Papers

Self-Supervised Speech Representations are More Phonetic than Semantic

2024-06-12 · Kwanghee Choi, Ankita Pasad, Tomohiko Nakamura, Satoru Fukayama, Karen Livescu, Shinji Watanabe

Self-supervised speech models (S3Ms) have become an effective backbone for speech applications. Various analyses suggest that S3Ms encode linguistic properties. In this work, we seek a more fine-grained analysis of the word-level linguistic properties encoded in S3Ms. Specifically, we curate a novel dataset of near homophone (phonetically similar) and synonym (semantically similar) word pairs and measure the similarities between S3M word representation pairs. Our study reveals that S3M representations consistently and significantly exhibit more phonetic than semantic similarity. Further, we question whether widely used intent classification datasets such as Fluent Speech Commands and Snips Smartlights are adequate for measuring semantic abilities. Our simple baseline, using only the word identity, surpasses S3M-based models. This corroborates our findings and suggests that high scores on these datasets do not necessarily guarantee the presence of semantic content.

📄 PDF Abstract BibTeX arXiv:2406.08619

Code (1)

juice500ml/phonetic_semantic_probing 공식 구현

Tasks

intent-classificationIntent ClassificationSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

2025-12-22 · Angelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun, Thomas Hueber 외 arxiv

This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on …

Self-Supervised LearningRepresentation Learning

UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data

2021-01-19 · Chengyi Wang, Yu Wu, Yao Qian, Kenichi Kumatani 외

In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both unlabeled and labeled data, in which supervised phonetic CTC learning and phonetically-aware contrastiv…

Multi-Task LearningRepresentation LearningSelf-Supervised Learningspeech-recognition+3

An Information-Theoretic Analysis of Self-supervised Discrete Representations of Speech

2023-06-04 · Badr M. Abdullah, Mohammed Maqsood Shaik, Bernd Möbius, Dietrich Klakow

Self-supervised representation learning for speech often involves a quantization step that transforms the acoustic input into discrete units. However, it remains unclear how to characterize the relationship between these…

QuantizationRepresentation Learning

Probing self-supervised speech models for phonetic and phonemic information: a case study in aspiration

2023-06-09 · Kinan Martin, Jon Gauthier, Canaan Breiss, Roger Levy

Textless self-supervised speech models have grown in capabilities in recent years, but the nature of the linguistic information they encode has not yet been thoroughly examined. We evaluate the extent to which these mode…

Learning Dependencies of Discrete Speech Representations with Neural Hidden Markov Models

2022-10-29 · Sung-Lin Yeh, Hao Tang

While discrete latent variable models have had great success in self-supervised learning, most models assume that frames are independent. Due to the segmental nature of phonemes in speech perception, modeling dependencie…

Self-Supervised Learning