Using Classifier Features to Determine Language Transfer on Morphemes
The aim of this thesis is to perform a Native Language Identification (NLI) task where we identify an English learner{'}s native language background based only on the learner{'}s English writing samples. We focus on the use of English grammatical morphemes across four proficiency levels. The outcome of the computational task is connected to a position in second language acquisition research that holds all learners acquire English grammatical morphemes in the same order, regardless of native language background. We use the NLI task as a tool to uncover cross-linguistic influence on the developmental trajectory of morphemes. We perform a cross-corpus evaluation across proficiency levels to increase the reliability and validity of the linguistic features that predict the native language background. We include native English data to determine the different morpheme patterns used by native versus non-native English speakers. Furthermore, we conduct a human NLI task to determine the type and magnitude of language transfer cues used by human raters versus the classifier.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-corpusLanguage AcquisitionLanguage IdentificationNative Language IdentificationText ClassificationSimilar Papers 제목 키워드 기반
Enhancing Gender-Inclusive Machine Translation with Neomorphemes and Large Language Models
Machine translation (MT) models are known to suffer from gender bias, especially when translating into languages with extensive gendered morphology. Accordingly, they still fall short in using gender-inclusive language, …
Machine TranslationTranslationRational Communication Shapes Morphological Composition
Human languages expand vocabularies by combining existing morphemes rather than inventing arbitrary forms. Communicative efficiency shapes lexical systems at multiple levels (Gibson et al., 2019), yet morphological compo…
Automatic Detection of Morphological Processes in the Yorùbá Language
Automatic morphology induction is important for computational processing of natural language. In resource-scarce languages in particular, it offers the possibility of supplementing data-driven strategies of Natural Langu…
Morpheme Induction for Emergent Language
We introduce CSAR, an algorithm for inducing morphemes from emergent language corpora of parallel utterances and meanings. It is a greedy algorithm that (1) weights morphemes based on mutual information between forms and…
A Pointer Network Architecture for Joint Morphological Segmentation and Tagging
Morphologically Rich Languages (MRLs) such as Arabic, Hebrew and Turkish often require Morphological Disambiguation (MD), i.e., the prediction of morphological decomposition of tokens into morphemes, early in the pipelin…
Morphological Disambiguation