paper-with-me

Papers

Metrical-accent Aware Vocal Onset Detection in Polyphonic Audio

2017-07-19 · Georgi Dzhambazov, Andre Holzapfel, Ajay Srinivasamurthy, Xavier Serra

The goal of this study is the automatic detection of onsets of the singing voice in polyphonic audio recordings. Starting with a hypothesis that the knowledge of the current position in a metrical cycle (i.e. metrical accent) can improve the accuracy of vocal note onset detection, we propose a novel probabilistic model to jointly track beats and vocal note onsets. The proposed model extends a state of the art model for beat and meter tracking, in which a-priori probability of a note at a specific metrical accent interacts with the probability of observing a vocal note onset. We carry out an evaluation on a varied collection of multi-instrument datasets from two music traditions (English popular music and Turkish makam) with different types of metrical cycles and singing styles. Results confirm that the proposed model reasonably improves vocal note onset detection accuracy compared to a baseline model that does not take metrical position into account.

📄 PDF Abstract BibTeX arXiv:1707.06163

Code (3)

georgid/lakh_vocal_segments_dataset 공식 구현
georgid/madmom 공식 구현
georgid/pypYIN 공식 구현

Tasks

Onset DetectionPosition

Similar Papers 제목 키워드 기반

Multi-Accent Mandarin Dry-Vocal Singing Dataset: Benchmark for Singing Accent Recognition

2025-12-07 · Zihao Wang, Ruibin Yuan, Ziqi Geng, Hengjia Li 외 arxiv

Singing accent research is underexplored compared to speech accent studies, primarily due to the scarcity of suitable datasets. Existing singing datasets often suffer from detail loss, frequently resulting from the vocal…

STRUM: A Spectral Transcription and Rhythm Understanding Model for End-to-End Generation of Playable Rhythm-Game Charts

2026-05-12 · Joshua Opria arxiv

We present STRUM (Spectral Transcription and Rhythm Understanding Model), an audio-to-chart pipeline that converts raw recordings into playable Clone Hero / YARG charts for drums, guitar, bass, vocals, and keys without a…

The Trajectory of Voice Onset Time with Vocal Aging

2018-10-15 · Xuanda Chen, Ziyu Xiong, Jian Hu

Vocal aging, a universal process of human aging, can largely affect one's language use, possibly including some subtle acoustic features of one's utterances like Voice Onset Time. To figure out the time effects, Queen El…

Human Aging

Robust detection of overlapping bioacoustic sound events

2025-03-04 · Louis Mahon, Benjamin Hoffman, Logan S James, Maddie Cusimano 외

We propose a method for accurately detecting bioacoustic sound events that is robust to overlapping events, a common issue in domains such as ethology, ecology and conservation. While standard methods employ a frame-base…

Event DetectionGraph Matchingobject-detectionObject Detection+1

TRIDENT: A Redundant Architecture for Caribbean-Accented Emergency Speech Triage

2025-12-11 · Elroy Galbraith, Chadwick Sutherland, Donahue Morgan arxiv

Emergency speech recognition systems exhibit systematic performance degradation on non-standard English varieties, creating a critical gap in services for Caribbean populations. We present TRIDENT (Transcription and Rout…

Speech Recognition