paper-with-me

홈 › Papers

ParsiNorm: A Persian Toolkit for Speech Processing Normalization

2021-11-01 · Romina Oji, Seyedeh Fatemeh Razavi, Sajjad Abdi Dehsorkh, Alireza Hariri, Hadi Asheri, Reshad Hosseini

In general, speech processing models consist of a language model along with an acoustic model. Regardless of the language model's complexity and variants, three critical pre-processing steps are needed in language models: cleaning, normalization, and tokenization. Among mentioned steps, the normalization step is so essential to format unification in pure textual applications. However, for embedded language models in speech processing modules, normalization is not limited to format unification. Moreover, it has to convert each readable symbol, number, etc., to how they are pronounced. To the best of our knowledge, there is no Persian normalization toolkits for embedded language models in speech processing modules, So in this paper, we propose an open-source normalization toolkit for text processing in speech applications. Briefly, we consider different readable Persian text like symbols (common currencies, #, @, URL, etc.), numbers (date, time, phone number, national code, etc.), and so on. Comparison with other available Persian textual normalization tools indicates the superiority of the proposed method in speech processing. Also, comparing the model's performance for one of the proposed functions (sentence separation) with other common natural language libraries such as HAZM and Parsivar indicates the proper performance of the proposed method. Besides, its evaluation of some Persian Wikipedia data confirms the proper performance of the proposed method.

📄 PDF Abstract BibTeX arXiv:2111.03470

Code (1)

haraai/normalization- 공식 구현

Tasks

Language ModellingSentence

Similar Papers 제목 키워드 기반

DadmaTools: Natural Language Processing Toolkit for Persian Language

2022-07-01 · NAACL (ACL) 2022 7 · Romina Etezadi, Mohammad Karrabi, Najmeh Zare, Mohamad Bagher Sajadi 외

We introduce DadmaTools, an open-source Python Natural Language Processing toolkit for the Persian language. The toolkit is a neural pipeline based on spaCy for several text processing tasks, including normalization, tok…

ChunkingConstituency ParsingDependency ParsingLemmatization

ParsiPy: NLP Toolkit for Historical Persian Texts in Python

2025-03-22 · Farhan Farsi, Parnian Fazel, Sepand Haghighi, Sadra Sabouri 외

The study of historical languages presents unique challenges due to their complex orthographic systems, fragmentary textual evidence, and the absence of standardized digital representations of text in those languages. Ta…

LemmatizationPart-Of-Speech TaggingTransliteration

Colloquial Persian POS (CPPOS) Corpus: A Novel Corpus for Colloquial Persian Part of Speech Tagging

2023-10-01 · Leyla Rabiei, Farzaneh Rahmani, Mohammad Khansari, Zeinab Rajabi 외

Introduction: Part-of-Speech (POS) Tagging, the process of classifying words into their respective parts of speech (e.g., verb or noun), is essential in various natural language processing applications. POS tagging is a …

Machine TranslationPart-Of-Speech TaggingPOSPOS Tagging+3

Parsivar: A Language Processing Toolkit for Persian

2018-05-01 · LREC 2018 5 · Salar Mohtaj, Behnam Roshanfekr, Atefeh Zafarian, Habibollah Asghari
Morphological Analysis

A Basic Language Resource Kit for Persian

2012-05-01 · LREC 2012 5 · Mojgan Seraji, Be{\'a}ta Megyesi, Joakim Nivre

Persian with its about 100,000,000 speakers in the world belongs to the group of languages with less developed linguistically annotated resources and tools. The few existing resources and tools are neither open source no…

Part-Of-Speech TaggingPOSSentenceSentence segmentation+1