paper-with-me

홈 › Papers

On the Use of Semantically-Aligned Speech Representations for Spoken Language Understanding

2022-10-11 · Gaëlle Laperrière, Valentin Pelloin, Mickaël Rouvier, Themos Stafylakis, Yannick Estève

In this paper we examine the use of semantically-aligned speech representations for end-to-end spoken language understanding (SLU). We employ the recently-introduced SAMU-XLSR model, which is designed to generate a single embedding that captures the semantics at the utterance level, semantically aligned across different languages. This model combines the acoustic frame-level speech representation learning model (XLS-R) with the Language Agnostic BERT Sentence Embedding (LaBSE) model. We show that the use of the SAMU-XLSR model instead of the initial XLS-R model improves significantly the performance in the framework of end-to-end SLU. Finally, we present the benefits of using this model towards language portability in SLU.

📄 PDF Abstract BibTeX arXiv:2210.05291

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSentenceSentence EmbeddingSentence-EmbeddingSpeech Representation LearningSpoken Language Understanding

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Weight Decay 설명 없음
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.

Similar Papers 제목 키워드 기반

Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects

2026-01-12 · Kalvin Chang, Yiwen Shao, Jiahong Li, Dong Yu arxiv

Despite having hundreds of millions of speakers, Chinese dialects lag behind Mandarin in speech and language technologies. Most varieties are primarily spoken, making dialect-to-Mandarin speech-LLMs (large language model…

Speech Recognition

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue

2026-06-11 · Wen Zhang, Xiaocui Yang, Zhuoyue Gao, Shi Feng 외 arxiv

Empathetic spoken dialogue systems require not only semantically appropriate responses but also emotionally aligned prosodic expression. However, cascade pipelines often discard acoustic cues during speech-to-text conver…

Dialogue GenerationResponse GenerationSpeech Synthesis

CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving

2024-06-16 · Bhavani Shankar, Preethi Jyothi, Pushpak Bhattacharyya

Code-switching is a widely prevalent linguistic phenomenon in multilingual societies like India. Building speech-to-text models for code-switched speech is challenging due to limited availability of datasets. In this wor…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+3

Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text

2026-08-24 · Hyeonyu Kim, Hwayeon Kim, Youngwon Choi, Myeongkyun Cho 외 arxiv

Spoken Language Models (SLMs) generate textual responses directly from speech, offering an alternative to cascaded systems. Despite recent advances, existing SLMs still exhibit weaker instruction-following behavior and l…

TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling

2025-04-09 · Liang-Hsuan Tseng, Yi-Chang Chen, Kuan-Yi Lee, Da-Shan Shiu 외

Recent efforts target spoken language models (SLMs) that not only listen but also speak for more natural human-LLM interaction. Joint speech-text modeling is a promising direction to achieve this. However, the effectiven…

Language ModelingLanguage Modellingparameter-efficient fine-tuningSpeech Tokenization