paper-with-me

Papers

Landmark-based consonant voicing detection on multilingual corpora

2016-11-10 · Xiang Kong, Xuesong Yang, Mark Hasegawa-Johnson, Jeung-Yoon Choi, Stefanie Shattuck-Hufnagel

This paper tests the hypothesis that distinctive feature classifiers anchored at phonetic landmarks can be transferred cross-lingually without loss of accuracy. Three consonant voicing classifiers were developed: (1) manually selected acoustic features anchored at a phonetic landmark, (2) MFCCs (either averaged across the segment or anchored at the landmark), and(3) acoustic features computed using a convolutional neural network (CNN). All detectors are trained on English data (TIMIT),and tested on English, Turkish, and Spanish (performance measured using F1 and accuracy). Experiments demonstrate that manual features outperform all MFCC classifiers, while CNNfeatures outperform both. MFCC-based classifiers suffer an F1reduction of 16% absolute when generalized from English to other languages. Manual features suffer only a 5% F1 reduction,and CNN features actually perform better in Turkish and Span-ish than in the training language, demonstrating that features capable of representing long-term spectral dynamics (CNN and landmark-based features) are able to generalize cross-lingually with little or no loss of accuracy

📄 PDF Abstract BibTeX arXiv:1611.03533

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Extracting Linguistic Knowledge from Speech: A Study of Stop Realization in 5 Romance Languages

2022-06-01 · LREC 2022 6 · Yaru Wu, Mathilde Hutin, Ioana Vasilescu, Lori Lamel 외

This paper builds upon recent work in leveraging the corpora and tools originally used to develop speech technologies for corpus-based linguistic studies. We address the non-canonical realization of consonants in connect…

speech-recognitionSpeech Recognition

Lenition and Fortition of Stop Codas in Romanian

2020-05-01 · LREC 2020 5 · Mathilde Hutin, Oana Niculescu, Ioana Vasilescu, Lori Lamel 외

The present paper aims at providing a first study of lenition- and fortition-type phenomena in coda position in Romanian, a language that can be considered as less-resourced. Our data show that there are two contexts for…

Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music

2026-01-16 · Venkat Suprabath Bitra, Homayoon Beigi arxiv

Reliable fundamental frequency (F 0) and voicing estimation is essential for neural synthesis, yet many pitch extractors depend on large labeled corpora and degrade under realistic recording artifacts. We propose a light…

Sur le voisement des consonnes fricatives finales en fran\ccais du Qu\'ebec (On final fricative consonant voicing in Quebec French)

2020-06-01 · JEPTALNRECITAL 2020 6 · Josiane Riverin-Coutl{\'e}e

Cette {\'e}tude s{'}int{\'e}resse aux indices acoustiques qui concourent {\`a} distinguer les fricatives non vois{\'e}es /f s ʃ/ et vois{\'e}es /v z ʒ/ en position de finale absolue en fran{\c{c}}ais du Qu{\'e}bec. La du…

Detection of Consonant Errors in Disordered Speech Based on Consonant-vowel Segment Embedding

2021-06-16 · Si-Ioi Ng, Cymie Wing-Yee Ng, Jingyu Li, Tan Lee

Speech sound disorder (SSD) refers to a type of developmental disorder in young children who encounter persistent difficulties in producing certain speech sounds at the expected age. Consonant errors are the major indica…