paper-with-me

홈 › Papers

Using a Knowledge Base to Automatically Annotate Speech Corpora and to Identify Sociolinguistic Variation

2022-06-01 · LREC 2022 6 · Yaru Wu, Fabian Suchanek, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker

Speech characteristics vary from speaker to speaker. While some variation phenomena are due to the overall communication setting, others are due to diastratic factors such as gender, provenance, age, and social background. The analysis of these factors, although relevant for both linguistic and speech technology communities, is hampered by the need to annotate existing corpora or to recruit, categorise, and record volunteers as a function of targeted profiles. This paper presents a methodology that uses a knowledge base to provide speaker-specific information. This can facilitate the enrichment of existing corpora with new annotations extracted from the knowledge base. The method also helps the large scale analysis by automatically extracting instances of speech variation to correlate with diastratic features. We apply our method to an over 120-hour corpus of broadcast speech in French and investigate variation patterns linked to reduction phenomena and/or specific to connected speech such as disfluencies. We find significant differences in speech rate, the use of filler words, and the rate of non-canonical realisations of frequent segments as a function of different professional categories and age groups.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semi-automatically Alignment of Predicates between Speech and OntoNotes data

2016-05-01 · LREC 2016 5 · Niraj Shrestha, Marie-Francine Moens

Speech data currently receives a growing attention and is an important source of information. We still lack suitable corpora of transcribed speech annotated with semantic roles that can be used for semantic role labeling…

Semantic Role LabelingSentence

Dodging the Data Bottleneck: Automatic Subtitling with Automatically Segmented ST Corpora

2022-09-21 · Sara Papi, Alina Karakanta, Matteo Negri, Marco Turchi

Speech translation for subtitling (SubST) is the task of automatically translating speech data into well-formed subtitles by inserting subtitle breaks compliant to specific displaying guidelines. Similar to speech transl…

Translation

QTLeap WSD/NED Corpora: Semantic Annotation of Parallel Corpora in Six Languages

2016-05-01 · LREC 2016 5 · Arantxa Otegi, Nora Aranberri, Antonio Branco, Jan Haji{\v{c}} 외

This work presents parallel corpora automatically annotated with several NLP tools, including lemma and part-of-speech tagging, named-entity recognition and classification, named-entity disambiguation, word-sense disambi…

Cross-Lingual TransferEntity DisambiguationGeneral ClassificationLEMMA+7

Scaling Rich Style-Prompted Text-to-Speech Datasets

2025-03-06 · Anuj Diwan, Zhisheng Zheng, David Harwath, Eunsol Choi

We introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scale dataset that annotates speech utterances with rich style captions. While rich abstract tags (e.g. guttural, nasal, pained) have been explored in…

Language ModelingLanguage ModellingTAGtext-to-speech+1

Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models

2025-08-10 · Qiongqiong Wang, Hardik B. Sailor, Jeremy H. M. Wong, Tianchi Liu 외 arxiv

Current large speech language models (Speech-LLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In t…