paper-with-me

홈 › Papers

Textless Speech-to-Music Retrieval Using Emotion Similarity

2023-03-19 · Seungheon Doh, Minz Won, Keunwoo Choi, Juhan Nam

We introduce a framework that recommends music based on the emotions of speech. In content creation and daily life, speech contains information about human emotions, which can be enhanced by music. Our framework focuses on a cross-domain retrieval system to bridge the gap between speech and music via emotion labels. We explore different speech representations and report their impact on different speech types, including acting voice and wake-up words. We also propose an emotion similarity regularization term in cross-domain retrieval tasks. By incorporating the regularization term into training, similar speech-and-music pairs in the emotion space are closer in the joint embedding space. Our comprehensive experimental results show that the proposed model is effective in textless speech-to-music retrieval.

📄 PDF Abstract BibTeX arXiv:2303.10539

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Similar Papers 제목 키워드 기반

Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations

2024-09-26 · Yujia Sun, Zeyu Zhao, Korin Richmond, Yuanchao Li

Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and…

Domain AdaptationDomain GeneralizationEmotion RecognitionMusic Emotion Recognition+3

Enhancing expressivity transfer in textless speech-to-speech translation

2023-10-11 · Jarod Duret, Benjamin O'Brien, Yannick Estève, Titouan Parcollet

Textless speech-to-speech translation systems are rapidly advancing, thanks to the integration of self-supervised learning techniques. However, existing state-of-the-art systems fall short when it comes to capturing and …

Self-Supervised LearningSpeech-to-Speech TranslationTranslation

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

2025-09-09 · Peijin Xie, Shun Qian, Bingquan Liu, Dexin Wang 외 arxiv

Document images encapsulate a wealth of knowledge, while the portability of spoken queries enables broader and flexible application scenarios. Yet, no prior work has explored knowledge base question answering over visual…

Knowledge Base Question Answering

Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation

2024-07-08 · Jarod Duret, Yannick Estève, Titouan Parcollet

Recent advancements in textless speech-to-speech translation systems have been driven by the adoption of self-supervised learning techniques. Although most state-of-the-art systems adopt a similar architecture to transfo…

Automatic Speech RecognitionEmotion Recognitionfeature selectionResynthesis+7

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

2026-01-25 · Xilin Jiang, Qiaolin Wang, Junkai Wu, Xiaomin He 외 arxiv

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we…