paper-with-me

Papers

ACE: Automatic Colloquialism, Typographical and Orthographic Errors Detection for Chinese Language

2016-12-01 · COLING 2016 12 · Shichao Dong, Gabriel Pui Cheong Fung, Binyang Li, Baolin Peng, Ming Liao, Jia Zhu, Kam-Fai Wong

We present a system called ACE for Automatic Colloquialism and Errors detection for written Chinese. ACE is based on the combination of N-gram model and rule-base model. Although it focuses on detecting colloquial Cantonese (a dialect of Chinese) at the current stage, it can be extended to detect other dialects. We chose Cantonese becauase it has many interesting properties, such as unique grammar system and huge colloquial terms, that turn the detection task extremely challenging. We conducted experiments using real data and synthetic data. The results indicated that ACE is highly reliable and effective.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Persian Typographical Error Type Detection Using Deep Neural Networks on Algorithmically-Generated Misspellings

2023-05-19 · Mohammad Dehghani, Heshaam Faili

Spelling correction is a remarkable challenge in the field of natural language processing. The objective of spelling correction tasks is to recognize and rectify spelling errors automatically. The development of applicat…

Spelling Correctiontoken-classificationToken Classification

Correcting the Autocorrect: Context-Aware Typographical Error Correction via Training Data Augmentation

2020-05-03 · LREC 2020 5 · Kshitij Shah, Gerard de Melo

In this paper, we explore the artificial generation of typographical errors based on real-world statistics. We first draw on a small set of annotated data to compute spelling error statistics. These are then invoked to i…

BIG-bench Machine LearningData Augmentation

Correcting Arabic Soft Spelling Mistakes using BiLSTM-based Machine Learning

2021-08-02 · Gheith A. Abandah, Ashraf Suyyagh, Mohammed Z. Khedher

Soft spelling errors are a class of spelling mistakes that is widespread among native Arabic speakers and foreign learners alike. Some of these errors are typographical in nature. They occur due to orthographic variation…

BIG-bench Machine LearningDecoderSpelling Correction

E-Bench: Towards Evaluating the Ease-of-Use of Large Language Models

2024-06-16 · Zhenyu Zhang, Bingguang Hao, Jinpeng Li, Zekai Zhang 외

Most large language models (LLMs) are sensitive to prompts, and another synonymous expression or a typo may lead to unexpected results for the model. Composing an optimal prompt for a specific demand lacks theoretical su…

GM-RKB WikiText Error Correction Task and Baselines

2020-05-01 · LREC 2020 5 · Gabor Melli, Abdelrhman Eldallal, Bassim Lazem, Olga Moreira

We introduce the GM-RKB WikiText Error Correction Task for the automatic detection and correction of typographical errors in WikiText annotated pages. The included corpus is based on a snapshot of the GM-RKB domain-speci…

Language ModelingLanguage Modelling