paper-with-me

Papers

Romanized to Native Malayalam Script Transliteration Using an Encoder-Decoder Framework

2024-12-13 · Bajiyo Baiju, Kavya Manohar, Leena G Pillai, Elizabeth Sherly

In this work, we present the development of a reverse transliteration model to convert romanized Malayalam to native script using an encoder-decoder framework built with attention-based bidirectional Long Short Term Memory (Bi-LSTM) architecture. To train the model, we have used curated and combined collection of 4.3 million transliteration pairs derived from publicly available Indic language translitertion datasets, Dakshina and Aksharantar. We evaluated the model on two different test dataset provided by IndoNLP-2025-Shared-Task that contain, (1) General typing patterns and (2) Adhoc typing patterns, respectively. On the Test Set-1, we obtained a character error rate (CER) of 7.4%. However upon Test Set-2, with adhoc typing patterns, where most vowel indicators are missing, our model gave a CER of 22.7%.

📄 PDF Abstract BibTeX arXiv:2412.09957

Code (1)

vrclc-duk/ml-en-transliteration 공식 구현

Tasks

DecoderTransliteration

Similar Papers 제목 키워드 기반

IndoNLP 2025: Shared Task on Real-Time Reverse Transliteration for Romanized Indo-Aryan languages

2025-01-10 · Deshan Sumanathilaka, Isuri Anuradha, Ruvan Weerasinghe, Nicholas Micallef 외

The paper overviews the shared task on Real-Time Reverse Transliteration for Romanized Indo-Aryan languages. It focuses on the reverse transliteration of low-resourced languages in the Indo-Aryan family to their native s…

Transliteration

Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches

2024-12-31 · Yomal De Mel, Kasun Wickramasinghe, Nisansa de Silva, Surangika Ranathunga

Due to reasons of convenience and lack of tech literacy, transliteration (i.e., Romanizing native scripts instead of using localization tools) is eminently prevalent in the context of low-resource languages such as Sinha…

DecoderMachine TranslationNMTTransliteration

Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration

2025-11-27 · Kanchon Gharami, Quazi Sarwar Muhtaseem, Deepti Gupta, Lavanya Elluri 외 arxiv

The development of robust transliteration techniques to enhance the effectiveness of transforming Romanized scripts into native scripts is crucial for Natural Language Processing tasks, including sentiment analysis, spee…

Information RetrievalSpeech RecognitionSentiment Analysis

Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset

2020-07-02 · LREC 2020 5 · Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke 외

This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each language: 1) native script Wikipedia text; 2)…

Language ModelingLanguage ModellingSentenceTransliteration

Swa-bhasha Resource Hub: Romanized Sinhala to Sinhala Transliteration Systems and Data Resources

2025-07-12 · Deshan Sumanathilaka, Sameera Perera, Sachithya Dharmasiri, Maneesha Athukorala 외 arxiv

The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala transliteration between 2020 and 2025. These resources have played a significant…