paper-with-me

홈 › Papers

Leveraging Laryngograph Data for Robust Voicing Detection in Speech

2023-12-05 · Yixuan Zhang, Heming Wang, DeLiang Wang

Accurately detecting voiced intervals in speech signals is a critical step in pitch tracking and has numerous applications. While conventional signal processing methods and deep learning algorithms have been proposed for this task, their need to fine-tune threshold parameters for different datasets and limited generalization restrict their utility in real-world applications. To address these challenges, this study proposes a supervised voicing detection model that leverages recorded laryngograph data. The model is based on a densely-connected convolutional recurrent neural network (DC-CRN), and trained on data with reference voicing decisions extracted from laryngograph data sets. Pretraining is also investigated to improve the generalization ability of the model. The proposed model produces robust voicing detection results, outperforming other strong baseline methods, and generalizes well to unseen datasets. The source code of the proposed model with pretraining is provided along with the list of used laryngograph datasets to facilitate further research in this area.

📄 PDF Abstract BibTeX arXiv:2312.03129

Code (1)

yixuanz/rvd 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Extracting Linguistic Knowledge from Speech: A Study of Stop Realization in 5 Romance Languages

2022-06-01 · LREC 2022 6 · Yaru Wu, Mathilde Hutin, Ioana Vasilescu, Lori Lamel 외

This paper builds upon recent work in leveraging the corpora and tools originally used to develop speech technologies for corpus-based linguistic studies. We address the non-canonical realization of consonants in connect…

speech-recognitionSpeech Recognition

Joint Robust Voicing Detection and Pitch Estimation Based on Residual Harmonics

2019-12-28 · Thomas Drugman, Abeer Alwan

This paper focuses on the problem of pitch tracking in noisy conditions. A method using harmonic information in the residual signal is presented. The proposed criterion is used both for pitch estimation, as well as for d…

Lenition and Fortition of Stop Codas in Romanian

2020-05-01 · LREC 2020 5 · Mathilde Hutin, Oana Niculescu, Ioana Vasilescu, Lori Lamel 외

The present paper aims at providing a first study of lenition- and fortition-type phenomena in coda position in Romanian, a language that can be considered as less-resourced. Our data show that there are two contexts for…

Voicing Personas: Rewriting Persona Descriptions into Style Prompts for Controllable Text-to-Speech

2025-05-21 · Yejin Lee, Jaehoon Kang, Kyuhong Shim

In this paper, we propose a novel framework to control voice style in prompt-based, controllable text-to-speech systems by leveraging textual personas as voice style prompts. We present two persona rewriting strategies t…

text-to-speechText to Speech

The Use of Voice Source Features for Sung Speech Recognition

2021-02-20 · Gerardo Roa Dabike, Jon Barker

In this paper, we ask whether vocal source features (pitch, shimmer, jitter, etc) can improve the performance of automatic sung speech recognition, arguing that conclusions previously drawn from spoken speech studies may…

speech-recognitionSpeech Recognitionvalid