Classifying German Language Proficiency Levels Using Large Language Models
Assessing language proficiency is essential for education, as it enables instruction tailored to learners needs. This paper investigates the use of Large Language Models (LLMs) for automatically classifying German texts according to the Common European Framework of Reference for Languages (CEFR) into different proficiency levels. To support robust training and evaluation, we construct a diverse dataset by combining multiple existing CEFR-annotated corpora with synthetic data. We then evaluate prompt-engineering strategies, fine-tuning of a LLaMA-3-8B-Instruct model and a probing-based approach that utilizes the internal neural state of the LLM for classification. Our results show a consistent performance improvement over prior methods, highlighting the potential of LLMs for reliable and scalable CEFR classification.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Coursebook Texts as a Helping Hand for Classifying Linguistic Complexity in Language Learners' Writings
We bring together knowledge from two different types of language learning data, texts learners read and texts they write, to improve linguistic complexity classification in the latter. Linguistic complexity in the foreig…
ClassificationDomain AdaptationGeneral ClassificationPredicting proficiency levels in learner writings by transferring a linguistic complexity model from expert-written coursebooks
The lack of a sufficient amount of data tailored for a task is a well-recognized problem for many statistical NLP methods. In this paper, we explore whether data sparsity can be successfully tackled when classifying lang…
Domain AdaptationLanguage AcquisitionTransfer LearningThe MERLIN corpus: Learner language and the CEFR
The MERLIN corpus is a written learner corpus for Czech, German,and Italian that has been designed to illustrate the Common European Framework of Reference for Languages (CEFR) with authentic learner data. The corpus con…
Language AcquisitionLanguage IdentificationNative Language IdentificationExperiments with Universal CEFR Classification
The Common European Framework of Reference (CEFR) guidelines describe language proficiency of learners on a scale of 6 levels. While the description of CEFR guidelines is generic across languages, the development of auto…
ClassificationGeneral ClassificationUnderstanding Editing Behaviors in Multilingual Wikipedia
Multilingualism is common offline, but we have a more limited understanding of the ways multilingualism is displayed online and the roles that multilinguals play in the spread of content between speakers of different lan…