paper-with-me

홈 › Papers

An Empirical Analysis of Speech Self-Supervised Learning at Multiple Resolutions

2024-10-31 · Theo Clark, Benedetta Cevoli, Eloy de Jong, Timofey Abramski, Jamie Dougherty

Self-supervised learning (SSL) models have become crucial in speech processing, with recent advancements concentrating on developing architectures that capture representations across multiple timescales. The primary goal of these multi-scale architectures is to exploit the hierarchical nature of speech, where lower-resolution components aim to capture representations that align with increasingly abstract concepts (e.g., from phones to words to sentences). Although multi-scale approaches have demonstrated some improvements over single-scale models, the precise reasons for these enhancements have poor empirical support. In this study, we present an initial analysis of layer-wise representations in multi-scale architectures, with a focus on Canonical Correlation Analysis (CCA) and Mutual Information (MI). We apply this analysis to Multi-Resolution HuBERT (MR-HuBERT) and find that (1) the improved performance on SUPERB tasks is primarily due to the auxiliary low-resolution loss rather than the downsampling itself, and (2) downsampling to lower resolutions neither improves downstream performance nor correlates with higher-level information (e.g., words), though it does improve computational efficiency. These findings challenge assumptions about the multi-scale nature of MR-HuBERT and motivate the importance of disentangling computational efficiency from learning better representations.

📄 PDF Abstract BibTeX arXiv:2410.23955

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencySelf-Supervised Learning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Focus 설명 없음

Similar Papers 제목 키워드 기반

Towards a Common Speech Analysis Engine

2022-03-01 · Hagai Aronowitz, Itai Gat, Edmilson Morais, Weizhong Zhu 외

Recent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-supervised representation learning-based syst…

Emotion RecognitionLanguage IdentificationRepresentation Learning

RemixIT: Continual self-training of speech enhancement models via bootstrapped remixing

2022-02-17 · Efthymios Tzinis, Yossi Adi, Vamsi Krishna Ithapu, Buye Xu 외

We present RemixIT, a simple yet effective self-supervised method for training speech enhancement without the need of a single isolated in-domain speech nor a noise waveform. Our approach overcomes limitations of previou…

Domain AdaptationSpeech EnhancementUnsupervised Domain Adaptation

When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation

2025-02-01 · Anna Min, Chenxu Hu, Yi Ren, Hang Zhao

Though end-to-end speech-to-text translation has been a great success, we argue that the cascaded speech-to-text translation model still has its place, which is usually criticized for the error propagation between automa…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+4

Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis

2025-05-27 · Tianyi Xu, Hongjie Chen, Wang Qing, Lv Hang 외

Large-scale training corpora have significantly improved the performance of ASR models. Unfortunately, due to the relative scarcity of data, Chinese accents and dialects remain a challenge for most ASR models. Recent adv…

Accented Speech RecognitionSelf-Supervised Learningspeech-recognitionSpeech Recognition

An empirical study on speech restoration guided by self supervised speech representation

2023-05-30 · Jaeuk Byun, Youna Ji, Soo Whan Chung, Soyeon Choe 외

Enhancing speech quality is an indispensable yet difficult task as it is often complicated by a range of degradation factors. In addition to additive noise, reverberation, clipping, and speech attenuation can all adverse…

Representation LearningSpeech Representation Learning