paper-with-me

Papers

Latin script keyboards for South Asian languages with finite-state normalization

2019-09-01 · WS 2019 9 · Lawrence Wolf-Sonkin, Vlad Schogol, Brian Roark, Michael Riley

The use of the Latin script for text entry of South Asian languages is common, even though there is no standard orthography for these languages in the script. We explore several compact finite-state architectures that permit variable spellings of words during mobile text entry. We find that approaches making use of transliteration transducers provide large accuracy improvements over baselines, but that simpler approaches involving a compact representation of many attested alternatives yields much of the accuracy gain. This is particularly important when operating under constraints on model size (e.g., on inexpensive mobile devices with limited storage and memory for keyboard models), and on speed of inference, since people typing on mobile keyboards expect no perceptual delay in keyboard responsiveness.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Transliteration

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset

2020-07-02 · LREC 2020 5 · Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke 외

This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each language: 1) native script Wikipedia text; 2)…

Language ModelingLanguage ModellingSentenceTransliteration

Criteria for Useful Automatic Romanization in South Asian Languages

2022-06-01 · LREC 2022 6 · Isin Demirsahin, Cibu Johny, Alexander Gutkin, Brian Roark

This paper presents a number of possible criteria for systems that transliterate South Asian languages from their native scripts into the Latin script, a process known as romanization. These criteria are related to eithe…

A Breadth-First Catalog of Text Processing, Speech Processing and Multimodal Research in South Asian Languages

2024-12-20 · Pranav Gupta

We review the recent literature (January 2022- October 2024) in South Asian languages on text-based language processing, multimodal models, and speech processing, and provide a spotlight analysis focused on 21 low-resour…

Bhaasha, Bhasa, Zaban: A Survey for Low-Resourced Languages in South Asia -- Current Stage and Challenges

2025-09-15 · Sampoorna Poria, Xiaolei Huang arxiv

Rapid developments of large language models have revolutionized many NLP tasks for English data. Unfortunately, the models and their evaluations for low-resource languages are being overlooked, especially for languages i…

SailCompass: Towards Reproducible and Robust Evaluation for Southeast Asian Languages

2024-12-02 · Jia Guo, Longxu Dou, Guangtao Zeng, Stanley Kok 외

In this paper, we introduce SailCompass, a reproducible and robust evaluation benchmark for assessing Large Language Models (LLMs) on Southeast Asian Languages (SEA). SailCompass encompasses three main SEA languages, eig…

Multiple-choice