paper-with-me

홈 › Papers

Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs

2026-05-05 · Stefano Bannò, Kate Knill, Mark Gales arxiv

Automated essay scoring (AES) research often relies on rank-based correlation metrics to validate analytic assessment. However, such metrics obscure both intrinsic intercorrelations among analytic dimensions that arise from the structure of writing proficiency itself and halo effects, whereby holistic impressions bleed into fine-grained component scores. As a result, high correlations may mask a system's true diagnostic behaviour. In this study, we propose a novel self-referential assessment evaluation framework that focuses on identifying intra-learner strengths and weaknesses rather than assessing inter-learner rankings. We conduct experiments on the publicly available ICNALE GRA, a uniquely dense second-language writing dataset annotated holistically and analytically by up to 80 trained raters. To obtain reliable reference scores, we apply two-facet Rasch modelling to calibrate rater severity and derive fair average scores across ten analytic aspects and holistic proficiency. We compare the analytic scoring performance of human operational raters and three large language models (LLMs) in a zero-shot setting. Our results show that LLMs tend to outperform single human raters in identifying relative weaknesses (negative feedback) across several proficiency aspects, while human raters remain stronger at identifying relative strengths (positive feedback). Overall, our findings highlight the limitations of rank-based evaluation for analytic assessment and demonstrate the value of intra-learner, profile-based methods for assessing and deploying LLMs in AES.

📄 PDF Abstract BibTeX arXiv:2605.04298

Code (0)

등록된 구현이 없습니다.

Tasks

Automated Essay Scoring

Similar Papers 제목 키워드 기반

LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English Writing

2025-02-17 · Zhengxiang Wang, Veronika Makarova, Zhi Li, Jordan Kodner 외

The paper explores the performance of LLMs in the context of multi-dimensional analytic writing assessments, i.e. their ability to provide both scores and comments based on multiple assessment criteria. Using a corpus of…

Towards a Formalisation of Value-based Actions and Consequentialist Ethics

2024-03-25 · Adam Wyner, Tomasz Zurek, DOrota Stachura-Zurek

Agents act to bring about a state of the world that is more compatible with their personal or institutional values. To formalise this intuition, the paper proposes an action framework based on the STRIPS formalisation. T…

Ethics

IFlyEA: A Chinese Essay Assessment System with Automated Rating, Review Generation, and Recommendation

2021-08-01 · ACL 2021 5 · Jiefu Gong, Xiao Hu, Wei Song, Ruiji Fu 외

Automated Essay Assessment (AEA) aims to judge students{'} writing proficiency in an automatic way. This paper presents a Chinese AEA system IFlyEssayAssess (IFlyEA), targeting on evaluating essays written by native Chin…

Review Generation

Evaluating AI-Generated Essays with GRE Analytical Writing Assessment

2024-10-22 · Yang Zhong, Jiangang Hao, Michael Fauss, Chen Li 외

The recent revolutionary advance in generative AI enables the generation of realistic and coherent texts by large language models (LLMs). Despite many existing evaluation metrics on the quality of the generated texts, th…

VerAs: Verify then Assess STEM Lab Reports

2024-02-07 · Berk Atıl, Mahsa Sheikhi Karizaki, Rebecca J. Passonneau

With an increasing focus in STEM education on critical thinking skills, science writing plays an ever more important role in curricula that stress inquiry skills. A recently published dataset of two sets of college level…

Automated Essay ScoringOpen-Domain Question AnsweringQuestion Answering