Normalization of Non-Standard Words in Croatian Texts
This paper presents text normalization which is an integral part of any text-to-speech synthesis system. Text normalization is a set of methods with a task to write non-standard words, like numbers, dates, times, abbreviations, acronyms and the most common symbols, in their full expanded form are presented. The whole taxonomy for classification of non-standard words in Croatian language together with rule-based normalization methods combined with a lookup dictionary are proposed. Achieved token rate for normalization of Croatian texts is 95%, where 80% of expanded words are in correct morphological form.
Code (0)
등록된 구현이 없습니다.
Tasks
FormGeneral ClassificationSpeech SynthesisText Normalizationtext-to-speechText to SpeechText-To-Speech SynthesisSimilar Papers 제목 키워드 기반
Non-Standard Words as Features for Text Categorization
This paper presents categorization of Croatian texts using Non-Standard Words (NSW) as features. Non-Standard Words are: numbers, dates, acronyms, abbreviations, currency, etc. NSWs in Croatian language are determined ac…
LemmatizationText CategorizationComplex Networks Measures for Differentiation between Normal and Shuffled Croatian Texts
This paper studies the properties of the Croatian texts via complex networks. We present network properties of normal and shuffled Croatian texts for different shuffling principles: on the sentence level and on the text …
SentenceInitial Comparison of Linguistic Networks Measures for Parallel Texts
This paper presents preliminary results of Croatian syllable networks analysis. Syllable network is a network in which nodes are syllables and links between them are constructed according to their connections within word…
ClusteringA preliminary study of Croatian Language Syllable Networks
This paper presents preliminary results of Croatian syllable networks analysis. Syllable network is a network in which nodes are syllables and links between them are constructed according to their connections within word…
ClusteringComparison of Short-Text Sentiment Analysis Methods for Croatian
We focus on the task of supervised sentiment classification of short and informal texts in Croatian, using two simple yet effective methods: word embeddings and string kernels. We investigate whether word embeddings offe…
General ClassificationSentiment AnalysisSentiment ClassificationStock Price Prediction+2