paper-with-me

Papers

How transferable are features in convolutional neural network acoustic models across languages?

2018-10-22 · NIPS Workshop IRASL 2018 · Anonymous

Characterization of the representations learned in intermediate layers of deep networks can provide valuable insight into the nature of a task and can guide the development of well-tailored learning strategies. Here we study convolutional neural network-based acoustic models in the context of automatic speech recognition. Adapting a method proposed by Yosinski et al. [2014], we measure the transferability of each layer between German and English to assess the their language-specifity. We observe three distinct regions of transferability: (1) the first two layers are entirely transferable between languages, (2) layers 2–8 are also highly transferable but we find evidence of some language specificity, (3) the subsequent fully connected layers are more language specific but can be successfully finetuned to the target language. To further probe the effect of weight freezing, we performed follow-up experiments using freeze-training [Raghu et al., 2017]. Our results are consistent with the observation that CCNs converge 'bottom up' during training and demonstrate the benefit of freeze training, especially for transfer learning.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Specificityspeech-recognitionSpeech RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models

2026-05-24 · Muhammad Ashad Kabir, Sirajam Munira arxiv

Large audio and language models have recently demonstrated zero-shot reasoning capabilities across various domains. However, it remains unclear how the form of audio input, whether handcrafted acoustic features extracted…

Probing Acoustic Representations for Phonetic Properties

2020-10-25 · Danni Ma, Neville Ryant, Mark Liberman

Pre-trained acoustic representations such as wav2vec and DeCoAR have attained impressive word error rates (WER) for speech recognition benchmarks, particularly when labeled data is limited. But little is known about what…

Benchmarkingspeech-recognitionSpeech Recognition

Landmark-based consonant voicing detection on multilingual corpora

2016-11-10 · Xiang Kong, Xuesong Yang, Mark Hasegawa-Johnson, Jeung-Yoon Choi 외

This paper tests the hypothesis that distinctive feature classifiers anchored at phonetic landmarks can be transferred cross-lingually without loss of accuracy. Three consonant voicing classifiers were developed: (1) man…

Fully Convolutional ASR for Less-Resourced Endangered Languages

2020-05-01 · LREC 2020 5 · Bao Thai, Robert Jimerson, Raymond Ptucha, Emily Prud{'}hommeaux

The application of deep learning to automatic speech recognition (ASR) has yielded dramatic accuracy increases for languages with abundant training data, but languages with limited training resources have yet to see accu…

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1

A multi-lingual and cross-domain analysis of features for text simplification

2020-05-01 · LREC 2020 5 · Regina Stodden, Laura Kallmeyer

In text simplification and readability research, several features have been proposed to estimate or simplify a complex text, e.g., readability scores, sentence length, or proportion of POS tags. These features are howeve…

POSSentenceText Simplification