paper-with-me

홈 › Papers

Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models

2025-08-17 · Abhirup Chakravarty arxiv

Automated Essay Scoring (AES) has emerged to prominence in response to the growing demand for educational automation. Providing an objective and cost-effective solution, AES standardises the assessment of extended responses. Although substantial research has been conducted in this domain, recent investigations reveal that alternative deep-learning architectures outperform transformer-based models. Despite the successful dominance in the performance of the transformer architectures across various other tasks, this discrepancy has prompted a need to enrich transformer-based AES models through contextual enrichment. This study delves into diverse contextual factors using the ASAP-AES dataset, analysing their impact on transformer-based model performance. Our most effective model, augmented with multiple contextual dimensions, achieves a mean Quadratic Weighted Kappa score of 0.823 across the entire essay dataset and 0.8697 when trained on individual essay sets. Evidently surpassing prior transformer-based models, this augmented approach only underperforms relative to the state-of-the-art deep learning model trained essay-set-wise by an average of 3.83\% while exhibiting superior performance in three of the eight sets. Importantly, this enhancement is orthogonal to architecture-based advancements and seamlessly adaptable to any AES model. Consequently, this contextual augmentation methodology presents a versatile technique for refining AES capabilities, contributing to automated grading and evaluation evolution in educational settings.

📄 PDF Abstract BibTeX arXiv:2508.16638

Code (0)

등록된 구현이 없습니다.

Tasks

Automated Essay Scoring

Results from the Paper

RankTaskDatasetModelMetrics
#1 Automated Essay Scoring ASAP-AES Empirical Analysis of the Effect of Cont Quadratic Weighted Kappa: 0.823

Similar Papers 제목 키워드 기반

Empirical Study of Large Language Models as Automated Essay Scoring Tools in English Composition__Taking TOEFL Independent Writing Task for Example

2024-01-07 · Wei Xia, Shaoguang Mao, Chanjing Zheng

Large language models have demonstrated exceptional capabilities in tasks involving natural language generation, reasoning, and comprehension. This study aims to construct prompts and comments grounded in the diverse sco…

Automated Essay ScoringPrompt LearningText Generation

Fine-Tuning Multilingual Language Models for Code Review: An Empirical Study on Industrial C# Projects

2025-07-25 · Igli Begolli, Meltem Aksoy, Daniel Neider arxiv

Code review is essential for maintaining software quality but often time-consuming and cognitively demanding, especially in industrial environments. Recent advancements in language models (LMs) have opened new avenues fo…

RepoReviewer: A Local-First Multi-Agent Architecture for Repository-Level Code Review

2026-03-17 · Peng Zhang arxiv

Repository-level code review requires reasoning over project structure, repository context, and file-level implementation details. Existing automated review workflows often collapse these tasks into a single pass, which …

Evolutionary Context Search for Automated Skill Acquisition

2026-02-18 · Qi Sun, Stefan Nielsen, Rio Yokota, Yujin Tang arxiv

Large Language Models cannot reliably acquire new knowledge post-deployment -- even when relevant text resources exist, models fail to transform them into actionable knowledge without retraining. Retrieval-Augmented Gene…

Semantic SimilarityPrompt Engineering

Automated Boilerplate: Prevalence and Quality of Contract Generators in the Context of Swiss Privacy Policies

2025-10-07 · Luka Nenadic, David Rodriguez arxiv

It has become increasingly challenging for firms to comply with a plethora of novel digital regulations. This is especially true for smaller businesses that often lack both the resources and know-how to draft complex leg…