paper-with-me

홈 › Papers

Analyzing the factors affecting usefulness of Self-Supervised Pre-trained Representations for Speech Recognition

2022-03-31 · Ashish Seth, Lodagala V S V Durga Prasad, Sreyan Ghosh, S. Umesh

Self-supervised learning (SSL) to learn high-level speech representations has been a popular approach to building Automatic Speech Recognition (ASR) systems in low-resource settings. However, the common assumption made in literature is that a considerable amount of unlabeled data is available for the same domain or language that can be leveraged for SSL pre-training, which we acknowledge is not feasible in a real-world setting. In this paper, as part of the Interspeech Gram Vaani ASR challenge, we try to study the effect of domain, language, dataset size, and other aspects of our upstream pre-training SSL data on the final performance low-resource downstream ASR task. We also build on the continued pre-training paradigm to study the effect of prior knowledge possessed by models trained using SSL. Extensive experiments and studies reveal that the performance of ASR systems is susceptible to the data used for SSL pre-training. Their performance improves with an increase in similarity and volume of pre-training data. We believe our work will be helpful to the speech community in building better ASR systems in low-resource settings and steer research towards improving generalization in SSL-based pre-training for speech systems.

📄 PDF Abstract BibTeX arXiv:2203.16973

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Recent Advancements in Self-Supervised Paradigms for Visual Feature Representation

2021-11-03 · Mrinal Anand, Aditya Garg

We witnessed a massive growth in the supervised learning paradigm in the past decade. Supervised learning requires a large amount of labeled data to reach state-of-the-art performance. However, labeling the samples requi…

NFTVis: Visual Analysis of NFT Performance

2023-06-05 · Fan Yan, Xumeng Wang, Ketian Mao, Wei zhang 외

A non-fungible token (NFT) is a data unit stored on the blockchain. Nowadays, more and more investors and collectors (NFT traders), who participate in transactions of NFTs, have an urgent need to assess the performance o…

Time Series

TweetCOVID: A System for Analyzing Public Sentiments and Discussions about COVID-19 via Twitter Activities

2021-03-02 · Jolin Shaynn-Ly Kwan, Kwan Hui Lim

The COVID-19 pandemic has created widespread health and economical impacts, affecting millions around the world. To better understand these impacts, we present the TweetCOVID system that offers the capability to understa…

Weakly Supervised Disentanglement with Guarantees

2019-10-22 · ICLR 2020 1 · Rui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon 외

Learning disentangled representations that correspond to factors of variation in real-world data is critical to interpretable and human-controllable machine learning. Recently, concerns about the viability of learning di…

Disentanglement

Driving through the Lens: Improving Generalization of Learning-based Steering using Simulated Adversarial Examples

2021-01-01 · Yu Shen, Laura Yu Zheng, Manli Shu, Weizi Li 외

To ensure the wide adoption and safety of autonomous driving, the vehicles need to be able to drive under various lighting, weather, and visibility conditions in different environments. These external and environmental f…

Autonomous DrivingData AugmentationDecision MakingSelf-Driving Cars+1