paper-with-me

홈 › Papers

Learning Multiple Utterance-Level Attribute Representations with a Unified Speech Encoder

2026-03-09 · Maryem Bouziane, Salima Mdhaffar, Yannick Estève arxiv

Speech foundation models trained with self-supervised learning produce generic speech representations that support a wide range of speech processing tasks. When further adapted with supervised learning, these models can achieve strong performance on specific downstream tasks. Recent post-training approaches, such as SAMU-XSLR and SONAR, align speech representations with utterance-level semantic representations, enabling effective multimodal (speech-text) and multilingual applications. While speech foundation models typically learn contextual embeddings at the acoustic frame level, these methods learn representations at the utterance level. In this work, we extend this paradigm to arbitrary utterance-level attributes and propose a unified post-training framework that enables a single speech foundation model to generate multiple types of utterance-level representations. We demonstrate the effectiveness of this approach by jointly learning semantic and speaker representations and evaluating them on multilingual speech retrieval and speaker recognition tasks.

📄 PDF Abstract BibTeX arXiv:2603.08312

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSpeaker Recognition

Similar Papers 제목 키워드 기반

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

2026-06-30 · Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong, Shilei Zhang 외 arxiv

Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emot…

Learning utterance-level representations through token-level acoustic latents prediction for Expressive Speech Synthesis

2022-11-01 · Karolos Nikitaras, Konstantinos Klapsas, Nikolaos Ellinas, Georgia Maniati 외

This paper proposes an Expressive Speech Synthesis model that utilizes token-level latent prosodic variables in order to capture and control utterance-level attributes, such as character acting voice and speaking style. …

DisentanglementDiversityExpressive Speech SynthesisSpeech Synthesis

A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units

2022-11-12 · Li-Wei Chen, Shinji Watanabe, Alexander Rudnicky

We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody and language content, leading to the deg…

RhythmVoice Conversion

Hierarchical Recurrent Attention Network for Response Generation

2017-01-25 · Chen Xing, Wei Wu, Yu Wu, Ming Zhou 외

We study multi-turn response generation in chatbots where a response is generated according to a conversation context. Existing work has modeled the hierarchy of the context, but does not pay enough attention to the fact…

Response Generation

Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction

2023-11-08 · Cam-Van Thi Nguyen, Anh-Tuan Mai, The-Son Le, Hai-Dang Kieu 외

Emotion recognition is a crucial task for human conversation understanding. It becomes more challenging with the notion of multimodal data, e.g., language, voice, and facial expressions. As a typical solution, the global…

Emotion RecognitionGraph Neural NetworkMultimodal Emotion Recognition