paper-with-me

Papers

GEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio Pretraining for Accurate Speech Emotion Recognition

2023-06-13 · Yu Pan, Yanni Hu, Yuguang Yang, Wen Fei, Jixun Yao, Heng Lu, Lei Ma, Jianjun Zhao

Contrastive cross-modality pretraining has recently exhibited impressive success in diverse fields, whereas there is limited research on their merits in speech emotion recognition (SER). In this paper, we propose GEmo-CLAP, a kind of gender-attribute-enhanced contrastive language-audio pretraining (CLAP) method for SER. Specifically, we first construct an effective emotion CLAP (Emo-CLAP) for SER, using pre-trained text and audio encoders. Second, given the significance of gender information in SER, two novel multi-task learning based GEmo-CLAP (ML-GEmo-CLAP) and soft label based GEmo-CLAP (SL-GEmo-CLAP) models are further proposed to incorporate gender information of speech signals, forming more reasonable objectives. Experiments on IEMOCAP indicate that our proposed two GEmo-CLAPs consistently outperform Emo-CLAP with different pre-trained models. Remarkably, the proposed WavLM-based SL-GEmo-CLAP obtains the best WAR of 83.16\%, which performs better than state-of-the-art SER methods.

📄 PDF Abstract BibTeX arXiv:2306.07848

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeContrastive LearningEmotion RecognitionMulti-Task LearningSelf-Supervised LearningSpeech Emotion Recognition

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
WordPiece 설명 없음

Similar Papers 제목 키워드 기반

T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining

2024-04-27 · Yi Yuan, Zhuo Chen, Xubo Liu, Haohe Liu 외

Contrastive language-audio pretraining~(CLAP) has been developed to align the representations of audio and language, achieving remarkable performance in retrieval and classification tasks. However, current CLAP struggles…

Retrieval

EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast

2025-05-29 · Shreeram Suresh Chandra, Lucas Goncalves, Junchen Lu, Carlos Busso 외

Current emotion-based contrastive language-audio pretraining (CLAP) methods typically learn by na\"ively aligning audio samples with corresponding text prompts. Consequently, this approach fails to capture the ordinal na…

Contrastive Learningcross-modal alignmentCross-Modal Retrieval

CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models

2023-10-12 · Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi 외

A fundamental characteristic of audio is its compositional nature. Audio-language models (ALMs) trained using a contrastive approach (e.g., CLAP) that learns a shared representation between audio and language modalities …

AttributeAudio ClassificationZero-shot Audio Classification

FIGMA: Towards FIne-Grained Music retrievAl

2026-06-04 · Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha 외 arxiv

Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries. When descriptions specify fine-grained mus…

Contrastive Learning

Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows

2025-04-22 · Haohe Liu, Thomas Deacon, Wenwu Wang, Matt Paradis 외

Locating the right sound effect efficiently is an important yet challenging topic for audio production. Most current sound-searching systems rely on pre-annotated audio labels created by humans, which can be time-consumi…

valid