paper-with-me

홈 › Papers

Understanding Shared Speech-Text Representations

2023-04-27 · Gary Wang, Kyle Kastner, Ankur Bapna, Zhehuai Chen, Andrew Rosenberg, Bhuvana Ramabhadran, Yu Zhang

Recently, a number of approaches to train speech models by incorpo-rating text into end-to-end models have been developed, with Mae-stro advancing state-of-the-art automatic speech recognition (ASR)and Speech Translation (ST) performance. In this paper, we expandour understanding of the resulting shared speech-text representationswith two types of analyses. First we examine the limits of speech-free domain adaptation, finding that a corpus-specific duration modelfor speech-text alignment is the most important component for learn-ing a shared speech-text representation. Second, we inspect the sim-ilarities between activations of unimodal (speech or text) encodersas compared to the activations of a shared encoder. We find that theshared encoder learns a more compact and overlapping speech-textrepresentation than the uni-modal encoders. We hypothesize that thispartially explains the effectiveness of the Maestro shared speech-textrepresentations.

📄 PDF Abstract BibTeX arXiv:2304.14514

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain Adaptationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Few-Shot Spoken Language Understanding via Joint Speech-Text Models

2023-10-09 · Chung-Ming Chien, Mingjiamei Zhang, Ju-chieh Chou, Karen Livescu

Recent work on speech representation models jointly pre-trained with text has demonstrated the potential of improving speech representations by encoding speech and text in a shared space. In this paper, we leverage such …

named-entity-recognitionNamed Entity RecognitionSentiment AnalysisSpoken Language Understanding

SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding

2020-10-05 · NAACL 2021 4 · Yu-An Chung, Chenguang Zhu, Michael Zeng

Spoken language understanding (SLU) requires a model to analyze input acoustic signal to understand its linguistic content and make predictions. To boost the models' performance, various pre-training methods have been pr…

Language ModelingLanguage ModellingMasked Language ModelingSpoken Language Understanding

CLARA: Multilingual Contrastive Learning for Audio Representation Acquisition

2023-10-18 · Kari A Noriy, Xiaosong Yang, Marcin Budka, Jian Jun Zhang

Multilingual speech processing requires understanding emotions, a task made difficult by limited labelled data. CLARA, minimizes reliance on labelled data, enhancing generalization across languages. It excels at fosterin…

Audio ClassificationContrastive LearningCross-Lingual TransferData Augmentation+6

Toward Joint Language Modeling for Speech Units and Text

2023-10-12 · Ju-chieh Chou, Chung-Ming Chien, Wei-Ning Hsu, Karen Livescu 외

Speech and text are two major forms of human language. The research community has been focusing on mapping speech to text or vice versa for many years. However, in the field of language modeling, very little effort has b…

Language ModelingLanguage ModellingSpeech-to-TextSpoken Language Understanding

mSLAM: Massively multilingual joint pre-training for speech and text

2022-02-03 · Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia 외

We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unlabeled speech and text in multiple langua…

cross-modal alignmentintent-classificationIntent ClassificationLanguage Modeling+4