paper-with-me

Papers

LeSpell - A Multi-Lingual Benchmark Corpus of Spelling Errors to Develop Spellchecking Methods for Learner Language

2022-06-01 · LREC 2022 6 · Marie Bexte, Ronja Laarmann-Quante, Andrea Horbach, Torsten Zesch

Spellchecking text written by language learners is especially challenging because errors made by learners differ both quantitatively and qualitatively from errors made by already proficient learners. We introduce LeSpell, a multi-lingual (English, German, Italian, and Czech) evaluation data set of spelling mistakes in context that we compiled from seven underlying learner corpora. Our experiments show that existing spellcheckers do not work well with learner data. Thus, we introduce a highly customizable spellchecking component for the DKPro architecture, which improves performance in many settings.

📄 PDF Abstract BibTeX

Code (1)

ltl-ude/ltl-spelling 공식 구현

Similar Papers 제목 키워드 기반

GitHub Typo Corpus: A Large-Scale Multilingual Dataset of Misspellings and Grammatical Errors

2019-11-28 · LREC 2020 5 · Masato Hagiwara, Masato Mita

The lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present …

Grammatical Error CorrectionSpelling Correction

A Benchmark Corpus of English Misspellings and a Minimally-supervised Model for Spelling Correction

2019-08-01 · WS 2019 8 · Michael Flor, Michael Fried, Alla Rozovskaya

Spelling correction has attracted a lot of attention in the NLP community. However, models have been usually evaluated on artificiallycreated or proprietary corpora. A publiclyavailable corpus of authentic misspellings, …

Spelling Correction

Multi-teacher Distillation for Multilingual Spelling Correction

2023-11-20 · Jingfen Zhang, Xuan Guo, Sravan Bodapati, Christopher Potts

Accurate spelling correction is a critical step in modern search interfaces, especially in an era of mobile devices and speech-to-text interfaces. For services that are deployed around the world, this poses a significant…

Multilingual NLPSpeech-to-TextSpelling Correction

Manually Annotated Spelling Error Corpus for Amharic

2021-06-25 · Andargachew Mekonnen Gezmu, Tirufat Tesifaye Lema, Binyam Ephrem Seyoum, Andreas Nürnberger

This paper presents a manually annotated spelling error corpus for Amharic, lingua franca in Ethiopia. The corpus is designed to be used for the evaluation of spelling error detection and correction. The misspellings are…

Text normalization for low-resource languages: the case of Ligurian

2022-06-16 · Stefano Lusito, Edoardo Ferrante, Jean Maillard

Text normalization is a crucial technology for low-resource languages which lack rigid spelling conventions or that have undergone multiple spelling reforms. Low-resource text normalization has so far relied upon hand-cr…

Text Normalization