paper-with-me

홈 › Papers

SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition

2023-09-29 · Hongfei Xue, Qijie Shao, Kaixun Huang, Peikun Chen, Jie Liu, Lei Xie

Multilingual automatic speech recognition (ASR) systems have garnered attention for their potential to extend language coverage globally. While self-supervised learning (SSL) models, like MMS, have demonstrated their effectiveness in multilingual ASR, it is worth noting that various layers' representations potentially contain distinct information that has not been fully leveraged. In this study, we propose a novel method that leverages self-supervised hierarchical representations (SSHR) to fine-tune the MMS model. We first analyze the different layers of MMS and show that the middle layers capture language-related information, and the high layers encode content-related information, which gradually decreases in the final layers. Then, we extract a language-related frame from correlated middle layers and guide specific language extraction through self-attention mechanisms. Additionally, we steer the model toward acquiring more content-related information in the final layers using our proposed Cross-CTC. We evaluate SSHR on two multilingual datasets, Common Voice and ML-SUPERB, and the experimental results demonstrate that our method achieves state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2309.16937

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Single-Stage Hierarchical Rectification for Weakly Supervised Histopathology Segmentation

2026-06-18 · Duc T. Nguyen, Hoang-Long Nguyen, Thanh-Ha DO, Huy-Hieu Pham arxiv

Existing weakly supervised semantic segmentation (WSSS) methods in computational pathology rely on a multi-stage paradigm: class activation map (CAM) generation, offline pseudo-mask refinement, and fully supervised retra…

Semantic Segmentation

Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning

2024-03-21 · Hasindri Watawana, Kanchana Ranasinghe, Tariq Mahmood, Muzammal Naseer 외

Self-supervised representation learning has been highly promising for histopathology image analysis with numerous approaches leveraging their patient-slide-patch hierarchy to learn better representations. In this paper, …

Representation LearningSelf-Supervised Learning

Use All The Labels: A Hierarchical Multi-Label Contrastive Learning Framework

2022-04-27 · CVPR 2022 1 · Shu Zhang, ran Xu, Caiming Xiong, Chetan Ramaiah

Current contrastive learning frameworks focus on leveraging a single supervisory signal to learn representations, which limits the efficacy on unseen data and downstream tasks. In this paper, we present a hierarchical mu…

AllContrastive LearningRepresentation Learning

Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency

2022-04-06 · CVPR 2022 1 · Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu 외

Natural videos provide rich visual contents for self-supervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos, leading to limited diversity in visual …

Contrastive LearningRepresentation LearningSelf-Supervised Learning

OPERA: Omni-Supervised Representation Learning with Hierarchical Supervisions

2022-10-11 · ICCV 2023 1 · Chengkun Wang, Wenzhao Zheng, Zheng Zhu, Jie zhou 외

The pretrain-finetune paradigm in modern computer vision facilitates the success of self-supervised learning, which tends to achieve better transferability than supervised learning. However, with the availability of mass…

image-classificationImage Classificationobject-detectionObject Detection+2