paper-with-me

홈 › Papers

Classifying German Language Proficiency Levels Using Large Language Models

2025-12-06 · Elias-Leander Ahlers, Witold Brunsmann, Malte Schilling arxiv

Assessing language proficiency is essential for education, as it enables instruction tailored to learners needs. This paper investigates the use of Large Language Models (LLMs) for automatically classifying German texts according to the Common European Framework of Reference for Languages (CEFR) into different proficiency levels. To support robust training and evaluation, we construct a diverse dataset by combining multiple existing CEFR-annotated corpora with synthetic data. We then evaluate prompt-engineering strategies, fine-tuning of a LLaMA-3-8B-Instruct model and a probing-based approach that utilizes the internal neural state of the LLM for classification. Our results show a consistent performance improvement over prior methods, highlighting the potential of LLMs for reliable and scalable CEFR classification.

📄 PDF Abstract BibTeX arXiv:2512.06483

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Coursebook Texts as a Helping Hand for Classifying Linguistic Complexity in Language Learners' Writings

2016-12-01 · WS 2016 12 · Ildik{\'o} Pil{\'a}n, David Alfter, Elena Volodina

We bring together knowledge from two different types of language learning data, texts learners read and texts they write, to improve linguistic complexity classification in the latter. Linguistic complexity in the foreig…

ClassificationDomain AdaptationGeneral Classification

Predicting proficiency levels in learner writings by transferring a linguistic complexity model from expert-written coursebooks

2016-12-01 · COLING 2016 12 · Ildik{\'o} Pil{\'a}n, Elena Volodina, Torsten Zesch

The lack of a sufficient amount of data tailored for a task is a well-recognized problem for many statistical NLP methods. In this paper, we explore whether data sparsity can be successfully tackled when classifying lang…

Domain AdaptationLanguage AcquisitionTransfer Learning

The MERLIN corpus: Learner language and the CEFR

2014-05-01 · LREC 2014 5 · Adriane Boyd, Jirka Hana, Lionel Nicolas, Detmar Meurers 외

The MERLIN corpus is a written learner corpus for Czech, German,and Italian that has been designed to illustrate the Common European Framework of Reference for Languages (CEFR) with authentic learner data. The corpus con…

Language AcquisitionLanguage IdentificationNative Language Identification

Experiments with Universal CEFR Classification

2018-04-18 · WS 2018 6 · Sowmya Vajjala, Taraka Rama

The Common European Framework of Reference (CEFR) guidelines describe language proficiency of learners on a scale of 6 levels. While the description of CEFR guidelines is generic across languages, the development of auto…

ClassificationGeneral Classification

Understanding Editing Behaviors in Multilingual Wikipedia

2015-08-28 · Suin Kim, Sungjoon Park, Scott A. Hale, Sooyoung Kim 외

Multilingualism is common offline, but we have a more limited understanding of the ways multilingualism is displayed online and the roles that multilinguals play in the spread of content between speakers of different lan…