paper-with-me

홈 › Papers

Automatic measurement of vowel duration via structured prediction

2016-10-26 · Yossi Adi, Joseph Keshet, Emily Cibelli, Erin Gustafson, Cynthia Clopper, Matthew Goldrick

A key barrier to making phonetic studies scalable and replicable is the need to rely on subjective, manual annotation. To help meet this challenge, a machine learning algorithm was developed for automatic measurement of a widely used phonetic measure: vowel duration. Manually-annotated data were used to train a model that takes as input an arbitrary length segment of the acoustic signal containing a single vowel that is preceded and followed by consonants and outputs the duration of the vowel. The model is based on the structured prediction framework. The input signal and a hypothesized set of a vowel's onset and offset are mapped to an abstract vector space by a set of acoustic feature functions. The learning algorithm is trained in this space to minimize the difference in expectations between predicted and manually-measured vowel durations. The trained model can then automatically estimate vowel durations without phonetic or orthographic transcription. Results comparing the model to three sets of manually annotated data suggest it out-performed the current gold standard for duration measurement, an HMM-based forced aligner (which requires orthographic or phonetic transcription as an input).

📄 PDF Abstract BibTeX arXiv:1610.08166

Code (1)

adiyoss/AutoVowelDuration 공식 구현

Tasks

PredictionStructured Prediction

Similar Papers 제목 키워드 기반

Multidimensional acoustic variation in vowels across English dialects

2022-07-01 · NAACL (SIGMORPHON) 2022 7 · James Tanner, Morgan Sonderegger, Jane Stuart-Smith

Vowels are typically characterized in terms of their static position in formant space, though vowels have also been long-known to undergo dynamic formant change over their timecourse. Recent studies have demonstrated tha…

Automatic Measurement of Pre-aspiration

2017-04-05 · Yaniv Sheena, Míša Hejná, Yossi Adi, Joseph Keshet

Pre-aspiration is defined as the period of glottal friction occurring in sequences of vocalic/consonantal sonorants and phonetically voiceless obstruents. We propose two machine learning methods for automatic measurement…

FrictionStructured Prediction

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

2026-02-06 · Yancheng Wang, Osama Hanna, Ruiming Xie, Xianfeng Rui 외 arxiv

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features such as fundamental frequency, intensity, a…

Reinforcement LearningEmotion Recognition

Duration Modeling by Multi-Models based on Vowel Production characteristics

2014-12-01 · WS 2014 12 · V Ramu Reddy, Parakrant Sarkar, K. Sreenivasa Rao
Speech SynthesisText-To-Speech Synthesis

Phonetic Segmentation of the UCLA Phonetics Lab Archive

2024-03-28 · Eleanor Chodroff, Blaž Pažon, Annie Baker, Steven Moran

Research in speech technologies and comparative linguistics depends on access to diverse and accessible speech data. The UCLA Phonetics Lab Archive is one of the earliest multilingual speech corpora, with long-form audio…