paper-with-me

Papers

DeepFry: Identifying Vocal Fry Using Deep Neural Networks

2022-03-31 · Bronya R. Chernyak, Talia Ben Simon, Yael Segal, Jeremy Steffman, Eleanor Chodroff, Jennifer S. Cole, Joseph Keshet

Vocal fry or creaky voice refers to a voice quality characterized by irregular glottal opening and low pitch. It occurs in diverse languages and is prevalent in American English, where it is used not only to mark phrase finality, but also sociolinguistic factors and affect. Due to its irregular periodicity, creaky voice challenges automatic speech processing and recognition systems, particularly for languages where creak is frequently used. This paper proposes a deep learning model to detect creaky voice in fluent speech. The model is composed of an encoder and a classifier trained together. The encoder takes the raw waveform and learns a representation using a convolutional neural network. The classifier is implemented as a multi-headed fully-connected network trained to detect creaky voice, voicing, and pitch, where the last two are used to refine creak prediction. The model is trained and tested on speech of American English speakers, annotated for creak by trained phoneticians. We evaluated the performance of our system using two encoders: one is tailored for the task, and the other is based on a state-of-the-art unsupervised representation. Results suggest our best-performing system has improved recall and F1 scores compared to previous methods on unseen data.

📄 PDF Abstract BibTeX arXiv:2203.17019

Code (1)

bronichern/deepfry 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Vocal Interactivity in Crowds, Flocks and Swarms: Implications for Voice User Interfaces

2019-07-26 · Roger K. Moore

Recent years have seen an explosion in the availability of Voice User Interfaces. However, user surveys suggest that there are issues with respect to usability, and it has been hypothesised that contemporary voice-enable…

Addressing the confounds of accompaniments in singer identification

2020-02-17 · Tsung-Han Hsieh, Kai-Hsiang Cheng, Zhe-Cheng Fan, Yu-Ching Yang 외

Identifying singers is an important task with many applications. However, the task remains challenging due to many issues. One major issue is related to the confounding factors from the background instrumental music that…

Data AugmentationSinger Identification

Survey on biomarkers in human vocalizations

2024-07-07 · Aki Härmä, Bert den Brinker, Ulf Grossekathofer, Okke Ouweltjes 외

Recent years has witnessed an increase in technologies that use speech for the sensing of the health of the talker. This survey paper proposes a general taxonomy of the technologies and a broad overview of current progre…

Survey

Towards Automated Animal Density Estimation with Acoustic Spatial Capture-Recapture

2023-08-24 · Yuheng Wang, Juan Ye, David L. Borchers

Passive acoustic monitoring can be an effective way of monitoring wildlife populations that are acoustically active but difficult to survey visually. Digital recorders allow surveyors to gather large volumes of data at l…

Density EstimationSurvey

Employing self-supervised learning models for cross-linguistic child speech maturity classification

2025-06-10 · Theo Zhang, Madurya Suresh, Anne S. Warlaumont, Kasia Hitczenko 외

Speech technology systems struggle with many downstream tasks for child speech due to small training corpora and the difficulties that child speech pose. We apply a novel dataset, SpeechMaturity, to state-of-the-art tran…

Self-Supervised Learningvalid