Predicting proficiency levels in learner writings by transferring a linguistic complexity model from expert-written coursebooks
The lack of a sufficient amount of data tailored for a task is a well-recognized problem for many statistical NLP methods. In this paper, we explore whether data sparsity can be successfully tackled when classifying language proficiency levels in the domain of learner-written output texts. We aim at overcoming data sparsity by incorporating knowledge in the trained model from another domain consisting of input texts written by teaching professionals for learners. We compare different domain adaptation techniques and find that a weighted combination of the two types of data performs best, which can even rival systems based on considerably larger amounts of in-domain data. Moreover, we show that normalizing errors in learners{'} texts can substantially improve classification when level-annotated in-domain data is not available.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain AdaptationLanguage AcquisitionTransfer LearningSimilar Papers 제목 키워드 기반
Coursebook Texts as a Helping Hand for Classifying Linguistic Complexity in Language Learners' Writings
We bring together knowledge from two different types of language learning data, texts learners read and texts they write, to improve linguistic complexity classification in the latter. Linguistic complexity in the foreig…
ClassificationDomain AdaptationGeneral ClassificationTowards interpretable models for language proficiency assessment: Predicting the CEFR level of Estonian learner texts
Using NLP to analyze authentic learner language helps to build automated assessment and feedback tools. It also offers new and extensive insights into the development of second language production. However, there is a la…
Assessing the validity of new paradigmatic complexity measures as criterial features for proficiency in L2 writings in English
This article addresses Second Language (L2) writing development through an investigation of new grammatical and structural complexity metrics. We explore the paradigmatic production in learner English by linking language…
Towards a Data Analytics Pipeline for the Visualisation of Complexity Metrics in L2 writings
We present the design of a tool for the visualisation of linguistic complexity in second language (L2) learner writings. We show how metrics can be exploited to visualise complexity in L2 writings in relation to CEFR lev…
RelationModeling language learning using specialized Elo rating
Automatic assessment of the proficiency levels of the learner is a critical part of Intelligent Tutoring Systems. We present methods for assessment in the context of language learning. We use a specialized Elo formula us…