Should You Fine-Tune BERT for Automated Essay Scoring?
Most natural language processing research now recommends large Transformer-based models with fine-tuning for supervised classification tasks; older strategies like bag-of-words features and linear models have fallen out of favor. Here we investigate whether, in automated essay scoring (AES) research, deep neural models are an appropriate technological choice. We find that fine-tuning BERT produces similar performance to classical models at significant additional cost. We argue that while state-of-the-art strategies do match existing best results, they come with opportunity costs in computational resources. We conclude with a review of promising areas for research on student essays where the unique characteristics of Transformers may provide benefits over classical methods to justify the costs.
Code (0)
등록된 구현이 없습니다.
Tasks
Automated Essay ScoringMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Enhancing Arabic Automated Essay Scoring with Synthetic Data and Error Injection
Automated Essay Scoring (AES) plays a crucial role in assessing language learners' writing quality, reducing grading workload, and providing real-time feedback. Arabic AES systems are particularly challenged by the lack …
Automated Essay ScoringType predictionTransGAT: Transformer-Based Graph Neural Networks for Multi-Dimensional Automated Essay Scoring
Essay writing is a critical component of student assessment, yet manual scoring is labor-intensive and inconsistent. Automated Essay Scoring (AES) offers a promising alternative, but current approaches face limitations. …
Automated Essay ScoringMitigating Bias in Automated Grading Systems for ESL Learners: A Contrastive Learning Approach
As Automated Essay Scoring (AES) systems are increasingly used in high-stakes educational settings, concerns regarding algorithmic bias against English as a Second Language (ESL) learners have increased. Current Transfor…
Automated Essay ScoringContrastive LearningLong Context Automated Essay Scoring with Language Models
Transformer-based language models are architecturally constrained to process text of a fixed maximum length. Essays written by higher-grade students frequently exceed the maximum allowed length for many popular open-sour…
Automated Essay ScoringMulti-Stage Pre-training for Automated Chinese Essay Scoring
This paper proposes a pre-training based automated Chinese essay scoring method. The method involves three components: weakly supervised pre-training, supervised cross- prompt fine-tuning and supervised target- prompt fi…
Domain Adaptation