paper-with-me

홈 › Papers

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining

2026-03-30 · Anuj Diwan, Eunsol Choi, David Harwath arxiv

We introduce ParaSpeechCLAP, a family of dual-encoder models that map speech and text style captions into a shared embedding space, supporting rich intrinsic (speaker-level) and situational (utterance-level) descriptors, such as pitch, texture, and emotion, beyond the narrow set handled by existing models. We train separate Intrinsic and Situational models alongside a unified Combined model, finding that specialized models are stronger on individual style dimensions while the unified model excels on compositional evaluation. We further show that ParaSpeechCLAP-Intrinsic benefits from an additional classification loss and class-balanced training. We demonstrate performance on style caption retrieval, speech attribute classification, and usability as inference-time reward models for style-prompted TTS. ParaSpeechCLAP models outperform baselines on most metrics across all three applications. Our models and code are released at https://github.com/ajd12342/paraspeechclap .

📄 PDF Abstract BibTeX arXiv:2603.28737

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text

2024-09-17 · Hongfei Xue, Wei Ren, Xuelong Geng, Kun Wei 외

Integrating audio encoders with LLMs through connectors has enabled these models to process and comprehend audio modalities, significantly enhancing speech-to-text tasks, including automatic speech recognition (ASR) and …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationspeech-recognition+2

A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding

2024-06-17 · Gaëlle Laperrière, Sahar Ghannay, Bassam Jabaian, Yannick Estève

Self-Supervised Learning is vastly used to efficiently represent speech for Spoken Language Understanding, gradually replacing conventional approaches. Meanwhile, textual SSL models are proposed to encode language-agnost…

Self-Supervised LearningSpoken Language Understanding

Residual Tokens Enhance Masked Autoencoders for Speech Modeling

2026-01-27 · Samir Sadok, Stéphane Lathuilière, Xavier Alameda-Pineda arxiv

Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce RT-MAE, a novel masked autoencoder frame…

Speech Enhancement

ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation

2026-01-16 · Zhuoyue Gao, Xiaohui Wang, Xiaocui Yang, Wen Zhang 외 arxiv

Empathetic speech dialogue requires not only understanding linguistic content but also perceiving rich paralinguistic information such as prosody, tone, and emotional intensity for affective understandings. Existing spee…

Empathetic Response GenerationSpeech Synthesis

Multi-reference Tacotron by Intercross Training for Style Disentangling,Transfer and Control in Speech Synthesis

2019-04-04 · Yanyao Bian, Changbin Chen, Yongguo Kang, Zhenglin Pan

Speech style control and transfer techniques aim to enrich the diversity and expressiveness of synthesized speech. Existing approaches model all speech styles into one representation, lacking the ability to control a spe…

DiversitySpeech Synthesis