paper-with-me

Papers

CoRuSS - a New Prosodically Annotated Corpus of Russian Spontaneous Speech

2016-05-01 · LREC 2016 5 · Tatiana Kachkovskaia, Daniil Kocharov, Pavel Skrelin, Nina Volskaya

This paper describes speech data recording, processing and annotation of a new speech corpus CoRuSS (Corpus of Russian Spontaneous Speech), which is based on connected communicative speech recorded from 60 native Russian male and female speakers of different age groups (from 16 to 77). Some Russian speech corpora available at the moment contain plain orthographic texts and provide some kind of limited annotation, but there are no corpora providing detailed prosodic annotation of spontaneous conversational speech. This corpus contains 30 hours of high quality recorded spontaneous Russian speech, half of it has been transcribed and prosodically labeled. The recordings consist of dialogues between two speakers, monologues (speakers{'} self-presentations) and reading of a short phonetically balanced text. Since the corpus is labeled for a wide range of linguistic - phonetic and prosodic - information, it provides basis for empirical studies of various spontaneous speech phenomena as well as for comparison with those we observe in prepared read speech. Since the corpus is designed as a open-access resource of speech data, it will also make possible to advance corpus-based analysis of spontaneous speech data across languages and speech technology development as well.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sense-Annotated Corpus for Russian

2022-09-01 · CLIB 2022 9 · Alexander Kirillovich, Natalia Loukachevitch, Maksim Kulaev, Angelina Bolshina 외

We present a sense-annotated corpus for Russian. The resource was obtained my manually annotating texts from the OpenCorpora corpus, an open corpus for the Russian language, by senses of Russian wordnet RuWordNet. The an…

Word Sense Disambiguation

Towards Large Vocabulary Kazakh-Russian Sign Language Dataset: KRSL-OnlineSchool

2022-06-01 · SignLang (LREC) 2022 6 · Medet Mukushev, Aigerim Kydyrbekova, Vadim Kimmelman, Anara Sandygulova

This paper presents a new dataset for Kazakh-Russian Sign Language (KRSL) created for the purposes of Sign Language Processing. In 2020, Kazakhstan’s schools were quickly switched to online mode due to the COVID-19 pande…

Sign Language TranslationSpeech-to-Text

RuCoCo: a new Russian corpus with coreference annotation

2022-06-10 · Vladimir Dobrovolskii, Mariia Michurina, Alexandra Ivoylova

We present a new corpus with coreference annotation, Russian Coreference Corpus (RuCoCo). The goal of RuCoCo is to obtain a large number of annotated texts while maintaining high inter-annotator agreement. RuCoCo contain…

Semi-automatically Annotated Learner Corpus for Russian

2022-06-01 · LREC 2022 6 · Anisia Katinskaia, Maria Lebedeva, Jue Hou, Roman Yangarber

We present ReLCo— the Revita Learner Corpus—a new semi-automatically annotated learner corpus for Russian. The corpus was collected while several thousand L2 learners were performing exercises using the Revita language-l…

Grammatical Error CorrectionGrammatical Error Detection

Designing a Russian Idiom-Annotated Corpus

2018-05-01 · LREC 2018 5 · Katsiaryna Aharodnik, Anna Feldman, Jing Peng
Machine TranslationWord Embeddings