paper-with-me

홈 › Papers

An Automatic Prosody Tagger for Spontaneous Speech

2016-12-01 · COLING 2016 12 · M{\'o}nica Dom{\'\i}nguez, Mireia Farr{\'u}s, Leo Wanner

Speech prosody is known to be central in advanced communication technologies. However, despite the advances of theoretical studies in speech prosody, so far, no large scale prosody annotated resources that would facilitate empirical research and the development of empirical computational approaches are available. This is to a large extent due to the fact that current common prosody annotation conventions offer a descriptive framework of intonation contours and phrasing based on labels. This makes it difficult to reach a satisfactory inter-annotator agreement during the annotation of gold standard annotations and, subsequently, to create consistent large scale annotations. To address this problem, we present an annotation schema for prominence and boundary labeling of prosodic phrases based upon acoustic parameters and a tagger for prosody annotation at the prosodic phrase level. Evaluation proves that inter-annotator agreement reaches satisfactory values, from 0.60 to 0.80 Cohen{'}s kappa, while the prosody tagger achieves acceptable recall and f-measure figures for five spontaneous samples used in the evaluation of monologue and dialogue formats in English and Spanish. The work presented in this paper is a first step towards a semi-automatic acquisition of large corpora for empirical prosodic analysis.

📄 PDF Abstract BibTeX

Code (1)

monikaUPF/modularProsodyTagger 공식 구현

Tasks

Descriptive

Similar Papers 제목 키워드 기반

Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models

2024-07-18 · Weiqin Li, Peiji Yang, Yicheng Zhong, Yixuan Zhou 외

Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS sy…

Language ModelingLanguage ModellingSpeech Synthesistext-to-speech+2

Prosody-controllable spontaneous TTS with neural HMMs

2022-11-24 · Harm Lameris, Shivam Mehta, Gustav Eje Henter, Joakim Gustafson 외

Spontaneous speech has many affective and pragmatic functions that are interesting and challenging to model in TTS. However, the presence of reduced articulation, fillers, repetitions, and other disfluencies in spontaneo…

Diversityvalid

Integrating Disfluency-based and Prosodic Features with Acoustics in Automatic Fluency Evaluation of Spontaneous Speech

2020-05-01 · LREC 2020 5 · Huaijin Deng, Youchao Lin, Takehito Utsuro, Akio Kobayashi 외

This paper describes an automatic fluency evaluation of spontaneous speech. In the task of automatic fluency evaluation, we integrate diverse features of acoustics, prosody, and disfluency-based ones. Then, we attempt to…

Data augmentation using prosody and false starts to recognize non-native children's speech

2020-08-29 · Hemant Kathania, Mittul Singh, Tamás Grósz, Mikko Kurimo

This paper describes AaltoASR's speech recognition system for the INTERSPEECH 2020 shared task on Automatic Speech Recognition (ASR) for non-native children's speech. The task is to recognize non-native speech from child…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+3

On the Role of Style in Parsing Speech with Neural Models

2020-10-08 · Trang Tran, Jiahong Yuan, Yang Liu, Mari Ostendorf

The differences in written text and conversational speech are substantial; previous parsers trained on treebanked text have given very poor results on spontaneous speech. For spoken language, the mismatch in style also e…