Native Language Identification with Large Language Models
We present the first experiments on Native Language Identification (NLI) using LLMs such as GPT-4. NLI is the task of predicting a writer's first language by analyzing their writings in a second language, and is used in second language acquisition and forensic linguistics. Our results show that GPT models are proficient at NLI classification, with GPT-4 setting a new performance record of 91.7% on the benchmark TOEFL11 test set in a zero-shot setting. We also show that unlike previous fully-supervised settings, LLMs can perform NLI without being limited to a set of known classes, which has practical implications for real-world applications. Finally, we also show that LLMs can provide justification for their choices, providing reasoning based on spelling errors, syntactic patterns, and usage of directly translated linguistic patterns.
Code (0)
등록된 구현이 없습니다.
Tasks
Language AcquisitionLanguage IdentificationNative Language IdentificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Improving Language Identification of Accented Speech
Language identification from speech is a common preprocessing step in many spoken language processing systems. In recent years, this field has seen fast progress, mostly due to the use of self-supervised models pretraine…
Language Identificationspeech-recognitionSpeech RecognitionSpoken language identificationNative Language Identification Using Large, Longitudinal Data
Native Language Identification (NLI) is a task aimed at determining the native language (L1) of learners of second language (L2) on the basis of their written texts. To date, research on NLI has focused on relatively sma…
BIG-bench Machine LearningLanguage IdentificationNative Language IdentificationText ClassificationNative Language Identification Using a Mixture of Character and Word N-grams
Native language identification (NLI) is the task of determining an author{'}s native language, based on a piece of his/her writing in a second language. In recent years, NLI has received much attention due to its challen…
Language AcquisitionLanguage IdentificationNative Language IdentificationMeasuring Interlanguage: Native Language Identification with L1-influence Metrics
The task of native language (L1) identification suffers from a relative paucity of useful training corpora, and standard within-corpus evaluation is often problematic due to topic bias. In this paper, we introduce a meth…
Language AcquisitionLanguage IdentificationMachine TranslationNative Language Identification+3On the Development of a Large Scale Corpus for Native Language Identification
Native Language Identification (NLI) is the task of identifying an author’s native language from their writings in a second language. In this paper, we introduce a new corpus (italki), which is larger than the current co…
BIG-bench Machine LearningLanguage IdentificationNative Language Identification