paper-with-me

Papers

Normalization of Non-Standard Words in Croatian Texts

2015-03-27 · Slobodan Beliga, Miran Pobar, Sanda Martinčić-Ipšić

This paper presents text normalization which is an integral part of any text-to-speech synthesis system. Text normalization is a set of methods with a task to write non-standard words, like numbers, dates, times, abbreviations, acronyms and the most common symbols, in their full expanded form are presented. The whole taxonomy for classification of non-standard words in Croatian language together with rule-based normalization methods combined with a lookup dictionary are proposed. Achieved token rate for normalization of Croatian texts is 95%, where 80% of expanded words are in correct morphological form.

📄 PDF Abstract BibTeX arXiv:1503.08167

Code (0)

등록된 구현이 없습니다.

Tasks

FormGeneral ClassificationSpeech SynthesisText Normalizationtext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

Non-Standard Words as Features for Text Categorization

2014-08-28 · Slobodan Beliga, Sanda Martinčić-Ipšić

This paper presents categorization of Croatian texts using Non-Standard Words (NSW) as features. Non-Standard Words are: numbers, dates, acronyms, abbreviations, currency, etc. NSWs in Croatian language are determined ac…

LemmatizationText Categorization

Complex Networks Measures for Differentiation between Normal and Shuffled Croatian Texts

2014-05-15 · Domagoj Margan, Ana Meštrović, Sanda Martinčić-Ipšić

This paper studies the properties of the Croatian texts via complex networks. We present network properties of normal and shuffled Croatian texts for different shuffling principles: on the sentence level and on the text …

Sentence

Initial Comparison of Linguistic Networks Measures for Parallel Texts

2014-05-08 · Kristina Ban, Ana Meštrović, Sanda Martinčić-Ipšić

This paper presents preliminary results of Croatian syllable networks analysis. Syllable network is a network in which nodes are syllables and links between them are constructed according to their connections within word…

Clustering

A preliminary study of Croatian Language Syllable Networks

2014-05-16 · Kristina Ban, Ivan Ivakić, Ana Meštrović

This paper presents preliminary results of Croatian syllable networks analysis. Syllable network is a network in which nodes are syllables and links between them are constructed according to their connections within word…

Clustering

Comparison of Short-Text Sentiment Analysis Methods for Croatian

2017-04-01 · WS 2017 4 · Leon Rotim, Jan {\v{S}}najder

We focus on the task of supervised sentiment classification of short and informal texts in Croatian, using two simple yet effective methods: word embeddings and string kernels. We investigate whether word embeddings offe…

General ClassificationSentiment AnalysisSentiment ClassificationStock Price Prediction+2