Improving Native Language Identification by Using Spelling Errors
In this paper, we explore spelling errors as a source of information for detecting the native language of a writer, a previously under-explored area. We note that character n-grams from misspelled words are very indicative of the native language of the author. In combination with other lexical features, spelling error features lead to 1.2{\%} improvement in accuracy on classifying texts in the TOEFL11 corpus by the author{'}s native language, compared to systems participating in the NLI shared task.
Code (0)
등록된 구현이 없습니다.
Tasks
Language IdentificationNative Language IdentificationSimilar Papers 제목 키워드 기반
Anglicized Words and Misspelled Cognates in Native Language Identification
In this paper, we present experiments that estimate the impact of specific lexical choices of people writing in a second language (L2). In particular, we look at misspelled words that indicate lexical uncertainty on the …
Language IdentificationNative Language IdentificationNative Language Identification with Large Language Models
We present the first experiments on Native Language Identification (NLI) using LLMs such as GPT-4. NLI is the task of predicting a writer's first language by analyzing their writings in a second language, and is used in …
Language AcquisitionLanguage IdentificationNative Language IdentificationNeural Networks and Spelling Features for Native Language Identification
We present the RUG-SU team{'}s submission at the Native Language Identification Shared Task 2017. We combine several approaches into an ensemble, based on spelling error features, a simple neural network using word repre…
Language IdentificationNative Language IdentificationWord EmbeddingsCCTC: A Cross-Sentence Chinese Text Correction Dataset for Native Speakers
The Chinese text correction (CTC) focuses on detecting and correcting Chinese spelling errors and grammatical errors. Most existing datasets of Chinese spelling check (CSC) and Chinese grammatical error correction (GEC) …
Grammatical Error CorrectionSentenceSimilarity-Based Unsupervised Spelling Correction Using BioWordVec: Development and Usability Study of Bacterial Culture and Antimicrobial Susceptibility Reports
Background: Existing bacterial culture test results for infectious diseases are written in unrefined text, resulting in many problems, including typographical errors and stop words. Effective spelling correction process…
Cultural Vocal Bursts Intensity PredictionSpelling Correction