paper-with-me

홈 › Papers

Using Analytic Scoring Rubrics in the Automatic Assessment of College-Level Summary Writing Tasks in L2

2017-11-01 · IJCNLP 2017 11 · Tamara Sladoljev-Agejev, Jan {\v{S}}najder

Assessing summaries is a demanding, yet useful task which provides valuable information on language competence, especially for second language learners. We consider automated scoring of college-level summary writing task in English as a second language (EL2). We adopt the Reading-for-Understanding (RU) cognitive framework, extended with the Reading-to-Write (RW) element, and use analytic scoring with six rubrics covering content and writing quality. We show that regression models with reference-based and linguistic features considerably outperform the baselines across all the rubrics. Moreover, we find interesting correlations between summary features and analytic rubrics, revealing the links between the RU and RW constructs.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Reading Comprehensionregression

Similar Papers 제목 키워드 기반

Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring

2024-07-04 · Xuansheng Wu, Padmaja Pravin Saraf, Gyeonggeon Lee, Ehsan Latif 외

Large language models (LLMs) have demonstrated strong potential in performing automatic scoring for constructed response assessments. While constructed responses graded by humans are usually based on given grading rubric…

Logical Reasoning

VerAs: Verify then Assess STEM Lab Reports

2024-02-07 · Berk Atıl, Mahsa Sheikhi Karizaki, Rebecca J. Passonneau

With an increasing focus in STEM education on critical thinking skills, science writing plays an ever more important role in curricula that stress inquiry skills. A recently published dataset of two sets of college level…

Automated Essay ScoringOpen-Domain Question AnsweringQuestion Answering

Applying Large Language Models and Chain-of-Thought for Automatic Scoring

2023-11-30 · Gyeong-Geon Lee, Ehsan Latif, Xuansheng Wu, Ninghao Liu 외

This study investigates the application of large language models (LLMs), specifically GPT-3.5 and GPT-4, with Chain-of-Though (CoT) in the automatic scoring of student-written responses to science assessments. We focused…

Few-Shot LearningPrompt EngineeringZero-Shot Learning

Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement

2026-05-07 · Jessica Huynh, Alfredo Gomez, Athiya Deviyani, Renee Shelby 외 arxiv

Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifications in a rubric presented to both huma…

SedarEval: Automated Evaluation using Self-Adaptive Rubrics

2025-01-26 · Zhiyuan Fan, Weinong Wang, Xing Wu, Debing Zhang

The evaluation paradigm of LLM-as-judge gains popularity due to its significant reduction in human labor and time costs. This approach utilizes one or more large language models (LLMs) to assess the quality of outputs fr…

Logical Reasoning