paper-with-me

홈 › Papers

Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays

2026-07-16 · John Maurice Gayed arxiv

This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in "AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models" (Gayed, 2026), which was fine-tuned on 480 argumentative essays from two prompts, we evaluate scoring accuracy on the full TOEFL11 corpus: 12,100 essays written by test-takers from 11 first-language backgrounds across eight prompts, none of which were seen during training. The model's raw scores (0.5-5.0) are mapped to the same three proficiency bands (low, medium, high) used by ETS, enabling direct comparison. The model achieved an overall band agreement of 77.79% and a quadratic weighted kappa of 0.702, with adjacent-band agreement of 99.98%. Accuracy was stable across all eight unseen prompts, with no advantage for prompts thematically related to the training data, indicating robust cross-prompt generalization. However, the model exhibited a systematic, L1-linked scoring offset. Within every proficiency band, essays from European-language backgrounds received consistently higher scores than essays from East-Asian-language backgrounds, a pattern not attributable to the composition of the fine-tuning data. This is the first large-scale L1 fairness analysis of a fine-tuned open-weight LLM for automated essay scoring.

📄 PDF Abstract BibTeX arXiv:2607.14605

Code (0)

등록된 구현이 없습니다.

Tasks

Automated Essay Scoring

Similar Papers 제목 키워드 기반

Automated Essay Scoring in the Presence of Biased Ratings

2018-06-01 · NAACL 2018 6 · Evelin Amorim, Marcia Can{\c{c}}ado, Adriano Veloso

Studies in Social Sciences have revealed that when people evaluate someone else, their evaluations often reflect their biases. As a result, rater bias may introduce highly subjective factors that make their evaluations i…

Automated Essay Scoring

Does the Prompt-based Large Language Model Recognize Students' Demographics and Introduce Bias in Essay Scoring?

2025-04-30 · Kaixun Yang, Mladen Raković, Dragan Gašević, Guanliang Chen

Large Language Models (LLMs) are widely used in Automated Essay Scoring (AES) due to their ability to capture semantic meaning. Traditional fine-tuning approaches required technical expertise, limiting accessibility for …

Automated Essay ScoringFairnessLanguage ModelingLanguage Modelling+1

Investigating neural architectures for short answer scoring

2017-09-01 · WS 2017 9 · Brian Riordan, Andrea Horbach, Aoife Cahill, Torsten Zesch 외

Neural approaches to automated essay scoring have recently shown state-of-the-art performance. The automated essay scoring task typically involves a broad notion of writing quality that encompasses content, grammar, orga…

Automated Essay ScoringReading ComprehensionWord Embeddings

Mitigating Bias in Automated Grading Systems for ESL Learners: A Contrastive Learning Approach

2026-01-23 · Kevin Fan, Eric Yun arxiv

As Automated Essay Scoring (AES) systems are increasingly used in high-stakes educational settings, concerns regarding algorithmic bias against English as a Second Language (ESL) learners have increased. Current Transfor…

Automated Essay ScoringContrastive Learning

Review of feedback in Automated Essay Scoring

2023-07-09 · You-Jin Jong, Yong-Jin Kim, Ok-Chol Ri

The first automated essay scoring system was developed 50 years ago. Automated essay scoring systems are developing into systems with richer functions than the previous simple scoring systems. Its purpose is not only to …

Automated Essay Scoring