paper-with-me

Papers

Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations

2024-09-26 · Yujia Sun, Zeyu Zhao, Korin Richmond, Yuanchao Li

Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and music, particularly those encoded by Self-Supervised Learning (SSL) models, remain largely unexplored, given the fact that SSL models for speech and music have rarely been applied in cross-domain research. In this work, we revisit the acoustic similarity between emotion speech and music, starting with an analysis of the layerwise behavior of SSL models for Speech Emotion Recognition (SER) and Music Emotion Recognition (MER). Furthermore, we perform cross-domain adaptation by comparing several approaches in a two-stage fine-tuning process, examining effective ways to utilize music for SER and speech for MER. Lastly, we explore the acoustic similarities between emotional speech and music using Frechet audio distance for individual emotions, uncovering the issue of emotion bias in both speech and music SSL models. Our findings reveal that while speech and music SSL models do capture shared acoustic features, their behaviors can vary depending on different emotions due to their training strategies and domain-specificities. Additionally, parameter-efficient fine-tuning can enhance SER and MER performance by leveraging knowledge from each other. This study provides new insights into the acoustic similarity between emotional speech and music, and highlights the potential for cross-domain generalization to improve SER and MER systems.

📄 PDF Abstract BibTeX arXiv:2409.17899

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationDomain GeneralizationEmotion RecognitionMusic Emotion Recognitionparameter-efficient fine-tuningSelf-Supervised LearningSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

2026-04-29 · Yun-Shao Tsai, Yi-Cheng Lin, Huang-Cheng Chou, Tzu-Wen Hsu 외 arxiv

Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on…

Voice Conversion

Can Musical Emotion Be Quantified With Neural Jitter Or Shimmer? A Novel EEG Based Study With Hindustani Classical Music

2017-04-29

The term jitter and shimmer has long been used in the domain of speech and acoustic signal analysis as a parameter for speaker identification and other prosodic features. In this study, we look forward to use the same pa…

EEGElectroencephalogram (EEG)Speaker Identification

EMPHASIS: An Emotional Phoneme-based Acoustic Model for Speech Synthesis System

2018-06-26

We present EMPHASIS, an emotional phoneme-based acoustic model for speech synthesis system. EMPHASIS includes a phoneme duration prediction model and an acoustic parameter prediction model. It uses a CBHG-based regressio…

Emotional Speech SynthesisParameter PredictionregressionSpeech Synthesis

Shaping the Epochal Individuality and Generality: The Temporal Dynamics of Uncertainty and Prediction Error in Musical Improvisation

2023-10-04 · Tatsuya Daikoku

Musical improvisation, much like spontaneous speech, reveals intricate facets of the improviser's state of mind and emotional character. However, the specific musical components that reveal such individuality remain larg…

Rhythm

A Theory-Based Explainable Deep Learning Architecture for Music Emotion

2024-08-13 · Hortense Fong, Vineet Kumar, K. Sudhir

This paper paper develops a theory-based, explainable deep learning convolutional neural network (CNN) classifier to predict the time-varying emotional response to music. We design novel CNN filters that leverage the fre…

Deep Learning