paper-with-me

Papers

Multi-Accent Mandarin Dry-Vocal Singing Dataset: Benchmark for Singing Accent Recognition

2025-12-07 · Zihao Wang, Ruibin Yuan, Ziqi Geng, Hengjia Li, Xingwei Qu, Xinyi Li, Songye Chen, Haoying Fu, Roger B. Dannenberg, Kejun Zhang arxiv

Singing accent research is underexplored compared to speech accent studies, primarily due to the scarcity of suitable datasets. Existing singing datasets often suffer from detail loss, frequently resulting from the vocal-instrumental separation process. Additionally, they often lack regional accent annotations. To address this, we introduce the Multi-Accent Mandarin Dry-Vocal Singing Dataset (MADVSD). MADVSD comprises over 670 hours of dry vocal recordings from 4,206 native Mandarin speakers across nine distinct Chinese regions. In addition to each participant recording audio of three popular songs in their native accent, they also recorded phonetic exercises covering all Mandarin vowels and a full octave range. We validated MADVSD through benchmark experiments in singing accent recognition, demonstrating its utility for evaluating state-of-the-art speech models in singing contexts. Furthermore, we explored dialectal influences on singing accent and analyzed the role of vowels in accentual variations, leveraging MADVSD's unique phonetic exercises.

📄 PDF Abstract BibTeX arXiv:2512.07005

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Metrical-accent Aware Vocal Onset Detection in Polyphonic Audio

2017-07-19 · Georgi Dzhambazov, Andre Holzapfel, Ajay Srinivasamurthy, Xavier Serra

The goal of this study is the automatic detection of onsets of the singing voice in polyphonic audio recordings. Starting with a hypothesis that the knowledge of the current position in a metrical cycle (i.e. metrical ac…

Onset DetectionPosition

VOCAL: Vowel and Consonant Layering for Expressive Animator-Centric Singing Animation

2022-11-30 · Siggraph Asia 2022 2022 11 · Yifang Pan, Chris Landreth, Eugene Fiume, Karan Singh Authors Info & Claims

Singing and speaking are two fundamental forms of human communication. From a modeling perspective however, speaking can be seen as a subset of singing. We present VOCAL, a system that automatically generates expressive,…

Sensitivity

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction

2024-12-11 · Sowmya Cheripally

This paper presents a new voice conversion model capable of transforming both speaking and singing voices. It addresses key challenges in current systems, such as conveying emotions, managing pronunciation and accent cha…

DecoderSelf-Supervised Learningtext-to-speechText to Speech+1

VocalSet: A Singing Voice Dataset

2018-09-25 · International Society for Music Information Retrieval Conference 2018 9 · Julia Wilkins, Prem Seetharaman, Alison Wahl, Bryan Pardo

We present VocalSet, a singing voice dataset of a capella singing. Existing singing voice datasets either do not capture a large range of vocal techniques, have very few singers, or are single-pitch and devoid of musical…

Singer IdentificationVocal technique classification

Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition

2022-01-24 · Xurong Xie, Xiang Sui, Xunying Liu, Lan Wang

The Mandarin Chinese language is known to be strongly influenced by a rich set of regional accents, while Mandarin speech with each accent is quite low resource. Hence, an important task in Mandarin speech recognition is…

Acoustic Modellingspeech-recognitionSpeech Recognition