Measuring Interlanguage: Native Language Identification with L1-influence Metrics
The task of native language (L1) identification suffers from a relative paucity of useful training corpora, and standard within-corpus evaluation is often problematic due to topic bias. In this paper, we introduce a method for L1 identification in second language (L2) texts that relies only on much more plentiful L1 data, rather than the L2 texts that are traditionally used for training. In particular, we do word-by-word translation of large L1 blog corpora to create a mapping to L2 forms that are a possible result of language transfer, and then use that information for unsupervised classification. We show this method is effective in several different learner corpora, with bigram features being particularly useful.
Code (0)
등록된 구현이 없습니다.
Tasks
Language AcquisitionLanguage IdentificationMachine TranslationNative Language IdentificationText ClassificationTranslationWord TranslationSimilar Papers 제목 키워드 기반
Generating a Lexicon of Errors in Portuguese to Support an Error Identification System for Spanish Native Learners
Portuguese is a less resourced language in what concerns foreign language learning. Aiming to inform a module of a system designed to support scientific written production of Spanish native speakers learning Portuguese, …
TranslationUnravelling Interlanguage Facts via Explainable Machine Learning
Native language identification (NLI) is the task of training (via supervised machine learning) a classifier that guesses the native language of the author of a text. This task has been extensively researched in the last …
BIG-bench Machine LearningLanguage IdentificationNative Language IdentificationMeasuring Feature Diversity in Native Language Identification
Towards Offensive Language Identification for Dravidian Languages
Offensive speech identification in countries like India poses several challenges due to the usage of code-mixed and romanized variants of multiple languages by the users in their posts on social media. The challenge of o…
Few-Shot LearningLanguage IdentificationTransfer LearningTransliteration+1Semantic Role Labeling for Learner Chinese: the Importance of Syntactic Parsing and L2-L1 Parallel Data
This paper studies semantic parsing for interlanguage (L2), taking semantic role labeling (SRL) as a case task and learner Chinese as a case language. We first manually annotate the semantic roles for a set of learner te…
Semantic ParsingSemantic Role LabelingSentence