paper-with-me

홈 › Papers

Criteria for Useful Automatic Romanization in South Asian Languages

2022-06-01 · LREC 2022 6 · Isin Demirsahin, Cibu Johny, Alexander Gutkin, Brian Roark

This paper presents a number of possible criteria for systems that transliterate South Asian languages from their native scripts into the Latin script, a process known as romanization. These criteria are related to either fidelity to human linguistic behavior (pronunciation transparency, naturalness and conventionality) or processing utility for people (ease of input) as well as under-the-hood in systems (invertibility and stability across languages and scripts). When addressing these differing criteria several linguistic considerations, such as modeling of prominent phonological processes and their relation to orthography, need to be taken into account. We discuss these key linguistic details in the context of Brahmic scripts and languages that use them, such as Hindi and Malayalam. We then present the core features of several romanization algorithms, implemented in a finite state transducer (FST) formalism, that address differing criteria. Implementations of these algorithms have been released as part of the Nisaba finite-state script processing library.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset

2020-07-02 · LREC 2020 5 · Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke 외

This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each language: 1) native script Wikipedia text; 2)…

Language ModelingLanguage ModellingSentenceTransliteration

IndicJR: A Judge-Free Benchmark of Jailbreak Robustness in South Asian Languages

2026-02-18 · Priyaranjan Pattnayak, Sanchari Chowdhuri arxiv

Safety alignment of large language models (LLMs) is mostly evaluated in English and contract-bound, leaving multilingual vulnerabilities understudied. We introduce \textbf{Indic Jailbreak Robustness (IJR)}, a judge-free …

Cloud-based Automatic Speech Recognition Systems for Southeast Asian Languages

2022-10-07 · Lei Wang, Rong Tong, Cheung Chi Leung, Sunil Sivadas 외

This paper provides an overall introduction of our Automatic Speech Recognition (ASR) systems for Southeast Asian languages. As not much existing work has been carried out on such regional languages, a few difficulties s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Bhasacitra: Visualising the dialect geography of South Asia

2021-05-28 · Aryaman Arora, Adam Farris, Gopalakrishnan R, Samopriya Basu

We present Bhasacitra, a dialect mapping system for South Asia built on a database of linguistic studies of languages of the region annotated for topic and location data. We analyse language coverage and look towards app…

Bhāṣācitra: Visualising the dialect geography of South Asia

2021-08-01 · ACL (LChange) 2021 8 · Aryaman Arora, Adam Farris, Gopalakrishnan R, Samopriya Basu

We present Bhāṣācitra, a dialect mapping system for South Asia built on a database of linguistic studies of languages of the region annotated for topic and location data. We analyse language coverage and look towards app…