paper-with-me

Papers

A General Framework for Learning Prosodic-Enhanced Representation of Rap Lyrics

2021-03-23 · Hongru Liang, Haozheng Wang, Qian Li, Jun Wang, Guandong Xu, Jiawei Chen, Jin-Mao Wei, Zhenglu Yang

Learning and analyzing rap lyrics is a significant basis for many web applications, such as music recommendation, automatic music categorization, and music information retrieval, due to the abundant source of digital music in the World Wide Web. Although numerous studies have explored the topic, knowledge in this field is far from satisfactory, because critical issues, such as prosodic information and its effective representation, as well as appropriate integration of various features, are usually ignored. In this paper, we propose a hierarchical attention variational autoencoder framework (HAVAE), which simultaneously consider semantic and prosodic features for rap lyrics representation learning. Specifically, the representation of the prosodic features is encoded by phonetic transcriptions with a novel and effective strategy~(i.e., rhyme2vec). Moreover, a feature aggregation strategy is proposed to appropriately integrate various features and generate prosodic-enhanced representation. A comprehensive empirical evaluation demonstrates that the proposed framework outperforms the state-of-the-art approaches under various metrics in different rap lyrics learning tasks.

📄 PDF Abstract BibTeX arXiv:2103.12615

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalMusic Information RetrievalMusic RecommendationRepresentation LearningRetrieval

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text

2024-05-30 · Jiaben Chen, Xin Yan, Yihang Chen, Siyuan Cen 외

In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these tw…

Motion Generation

It's Only Words And Words Are All I Have

2019-01-16 · Barman Manash Pratim, Dahekar Kavish, Anshuman Abhinav, Awekar Amit

The central idea of this paper is to demonstrate the strength of lyrics for music mining and natural language processing (NLP) tasks using the distributed representation paradigm. For music mining, we address two predict…

AllBinary ClassificationGeneral ClassificationMulti-class Classification

Unsupervised Generative Adversarial Alignment Representation for Sheet music, Audio and Lyrics

2020-07-29 · Donghuo Zeng, Yi Yu, Keizo Oyama

Sheet music, audio, and lyrics are three main modalities during writing a song. In this paper, we propose an unsupervised generative adversarial alignment representation (UGAAR) model to learn deep discriminative represe…

Representation Learning

Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations

2026-04-02 · Haitong Sun, Stephen McIntosh, Kwanghee Choi, Eunjung Yeo 외 arxiv

Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. The ABX discrimination task ha…

Non-verbal information in spontaneous speech -- towards a new framework of analysis

2024-03-06 · Tirza Biron, Moshe Barboy, Eran Ben-Artzy, Alona Golubchik 외

Non-verbal signals in speech are encoded by prosody and carry information that ranges from conversation action to attitude and emotion. Despite its importance, the principles that govern prosodic structure are not yet ad…

speech-recognitionSpeech Recognition