paper-with-me

Papers

Khmer Spellchecking: A Holistic Approach

2025-11-12 · Marry Kong, Rina Buoy, Sovisal Chenda, Nguonly Taing arxiv

Compared to English and other high-resource languages, spellchecking for Khmer remains an unresolved problem due to several challenges. First, there are misalignments between words in the lexicon and the word segmentation model. Second, a Khmer word can be written in different forms. Third, Khmer compound words are often loosely and easily formed, and these compound words are not always found in the lexicon. Fourth, some proper nouns may be flagged as misspellings due to the absence of a Khmer named-entity recognition (NER) model. Unfortunately, existing solutions do not adequately address these challenges. This paper proposes a holistic approach to the Khmer spellchecking problem by integrating Khmer subword segmentation, Khmer NER, Khmer grapheme-to-phoneme (G2P) conversion, and a Khmer language model to tackle these challenges, identify potential correction candidates, and rank the most suitable candidate. Experimental results show that the proposed approach achieves a state-of-the-art Khmer spellchecking accuracy of up to 94.4%, compared to existing solutions. The benchmark datasets for Khmer spellchecking and NER tasks in this study will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2511.09812

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LeSpell - A Multi-Lingual Benchmark Corpus of Spelling Errors to Develop Spellchecking Methods for Learner Language

2022-06-01 · LREC 2022 6 · Marie Bexte, Ronja Laarmann-Quante, Andrea Horbach, Torsten Zesch

Spellchecking text written by language learners is especially challenging because errors made by learners differ both quantitatively and qualitatively from errors made by already proficient learners. We introduce LeSpell…

Towards Explainable Khmer Polarity Classification

2025-11-12 · Marry Kong, Rina Buoy, Sovisal Chenda, Nguonly Taing arxiv

Khmer polarity classification is a fundamental natural language processing task that assigns a positive, negative, or neutral label to a given Khmer text input. Existing Khmer models typically predict the label without e…

A Survey on Importance of Homophones Spelling Correction Model for Khmer Authors

2024-11-11 · Seanghort Born, Madeth May, Claudine Piau-Toffolon, Sébastien Iksal

Homophones present a significant challenge to authors in any languages due to their similarities of pronunciations but different meanings and spellings. This issue is particularly pronounced in the Khmer language, rich i…

Spelling CorrectionSurvey

Khmer Word Search: Challenges, Solutions, and Semantic-Aware Search

2021-12-16 · Rina Buoy, Nguonly Taing, Sovisal Chenda

Search is one of the key functionalities in digital platforms and applications such as an electronic dictionary, a search engine, and an e-commerce platform. While the search function in some languages is trivial, Khmer …

PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation

2025-12-15 · Hour Kaing, Raj Dabre, Haiyue Song, Van-Hien Tran 외 arxiv

This work introduces {\it PrahokBART}, a compact pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora. We focus on improving the pre-training corpus qual…

Machine TranslationHeadline GenerationText SummarizationText Generation