paper-with-me

Papers

Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems

2024-10-03 · Olga Iakovenko, Ivan Bondarenko, Mariya Borovikova, Daniil Vodolazsky

This paper presents an overview of rule-based system for automatic accentuation and phonemic transcription of Russian texts for speech connected tasks, such as Automatic Speech Recognition (ASR). Two parts of the developed system, accentuation and transcription, use different approaches to achieve correct phonemic representations of input phrases. Accentuation is based on "Grammatical dictionary of the Russian language" of A.A. Zaliznyak and wiktionary corpus. To distinguish homographs, the accentuation system also utilises morphological information of the sentences based on Recurrent Neural Networks (RNN). Transcription algorithms apply the rules presented in the monograph of B.M. Lobanov and L.I. Tsirulnik "Computer Synthesis and Voice Cloning". The rules described in the present paper are implemented in an open-source module, which can be of use to any scientific study connected to ASR or Speech To Text (STT) tasks. Automatically marked up text annotations of the Russian Voxforge database were used as training data for an acoustic model in CMU Sphinx. The resulting acoustic model was evaluated on cross-validation, mean Word Accuracy being 71.2%. The developed toolkit is written in the Python language and is accessible on GitHub for any researcher interested.

📄 PDF Abstract BibTeX arXiv:2410.02538

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSpeech-to-TextVoice Cloning

Similar Papers 제목 키워드 기반

A Cross-language Corpus for Studying the Phonetics and Phonology of Prominence

2014-05-01 · LREC 2014 5 · Bistra Andreeva, William Barry, Jacques Koreman

The present article describes a corpus which was collected for the cross-language comparison of prominence. In the data analysis, the acoustic-phonetic properties of words spoken with two different levels of accentuation…

Creating an Aligned Russian Text Simplification Dataset from Language Learner Data

2021-04-01 · EACL (BSNLP) 2021 4 · Anna Dmitrieva, Jörg Tiedemann

Parallel language corpora where regular texts are aligned with their simplified versions can be used in both natural language processing and theoretical linguistic studies. They are essential for the task of automatic te…

Text Simplification

Methods for Detoxification of Texts for the Russian Language

2021-05-19 · Daryna Dementieva, Daniil Moskovskiy, Varvara Logacheva, David Dale 외

We introduce the first study of automatic detoxification of Russian texts to combat offensive language. Such a kind of textual style transfer can be used, for instance, for processing toxic content in social media. While…

Style Transfer

Translationese in Russian Literary Texts

2021-11-01 · EMNLP (LaTeCHCLfL, CLFL, LaTeCH) 2021 11 · Maria Kunilovskaya, Ekaterina Lapshinova-Koltunski, Ruslan Mitkov

The paper reports the results of a translationese study of literary texts based on translated and non-translated Russian. We aim to find out if translations deviate from non-translated literary texts, and if the establis…

Specificity

Automatic Aspect Extraction from Scientific Texts

2023-10-06 · Anna Marshalova, Elena Bruches, Tatiana Batura

Being able to extract from scientific papers their main points, key insights, and other important information, referred to here as aspects, might facilitate the process of conducting a scientific literature review. There…

Aspect Extraction