paper-with-me

홈 › Papers

Analysis of Predictive Coding Models for Phonemic Representation Learning in Small Datasets

2020-07-08 · María Andrea Cruz Blandón, Okko Räsänen

Neural network models using predictive coding are interesting from the viewpoint of computational modelling of human language acquisition, where the objective is to understand how linguistic units could be learned from speech without any labels. Even though several promising predictive coding -based learning algorithms have been proposed in the literature, it is currently unclear how well they generalise to different languages and training dataset sizes. In addition, despite that such models have shown to be effective phonemic feature learners, it is unclear whether minimisation of the predictive loss functions of these models also leads to optimal phoneme-like representations. The present study investigates the behaviour of two predictive coding models, Autoregressive Predictive Coding and Contrastive Predictive Coding, in a phoneme discrimination task (ABX task) for two languages with different dataset sizes. Our experiments show a strong correlation between the autoregressive loss and the phoneme discrimination scores with the two datasets. However, to our surprise, the CPC model shows rapid convergence already after one pass over the training data, and, on average, its representations outperform those of APC on both languages.

📄 PDF Abstract BibTeX arXiv:2007.04205

Code (0)

등록된 구현이 없습니다.

Tasks

Language AcquisitionRepresentation Learning

Methods 이 논문이 사용한 방법론

InfoNCE 설명 없음
Contrastive Predictive Coding Contrastive Predictive Coding (CPC) learns self-supervised representations by predicting the future in latent space by using powerful autoregressive models. The model uses a…

Similar Papers 제목 키워드 기반

Mitigating the Linguistic Gap with Phonemic Representations for Robust Cross-lingual Transfer

2024-02-22 · Haeji Jung, Changdae Oh, Jooeon Kang, Jimin Sohn 외

Approaches to improving multilingual language understanding often struggle with significant performance gaps between high-resource and low-resource languages. While there are efforts to align the languages in a single la…

Cross-Lingual TransferLanguage Modelling

SCaLa: Supervised Contrastive Learning for End-to-End Speech Recognition

2021-10-08 · Li Fu, Xiaoxiao Li, Runyu Wang, Lu Fan 외

End-to-end Automatic Speech Recognition (ASR) models are usually trained to optimize the loss of the whole token sequence, while neglecting explicit phonemic-granularity supervision. This could result in recognition erro…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive LearningRepresentation Learning+2

IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling

2025-04-03 · Zébulon Goriely, Paula Buttery

In this paper, we introduce two resources: (i) G2P+, a tool for converting orthographic datasets to a consistent phonemic representation; and (ii) IPA CHILDES, a phonemic dataset of child-centered speech across 31 langua…

Grapheme-to-Phoneme ConversionLanguage ModelingLanguage Modelling

A Phonemic Corpus of Polish Child-Directed Speech

2012-05-01 · LREC 2012 5 · Luc Boruta, Justyna Jastrzebska

Recent advances in modeling early language acquisition are due not only to the development of machine-learning techniques, but also to the increasing availability of data on child language and child-adult interaction. In…

Language Acquisition

Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation

2025-06-09 · Rui Hu, Xiaolong Lin, Jiawang Liu, Shixi Huang 외

In this paper, we propose a method for annotating phonemic and prosodic labels on a given audio-transcript pair, aimed at constructing Japanese text-to-speech (TTS) datasets. Our approach involves fine-tuning a large-sca…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2