paper-with-me

Papers

Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory

2025-07-26 · Dan Song, Won-Chan Lee, Hong Jiao arxiv

This study investigates the estimation of reliability for large language models (LLMs) in scoring writing tasks from the AP Chinese Language and Culture Exam. Using generalizability theory, the research evaluates and compares score consistency between human and AI raters across two types of AP Chinese free-response writing tasks: story narration and email response. These essays were independently scored by two trained human raters and seven AI raters. Each essay received four scores: one holistic score and three analytic scores corresponding to the domains of task completion, delivery, and language use. Results indicate that although human raters produced more reliable scores overall, LLMs demonstrated reasonable consistency under certain conditions, particularly for story narration tasks. Composite scoring that incorporates both human and AI raters improved reliability, which supports that hybrid scoring models may offer benefits for large-scale writing assessments.

📄 PDF Abstract BibTeX arXiv:2507.19980

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Machine-assisted writing evaluation: Exploring pre-trained language models in analyzing argumentative moves

2025-03-25 · Wenjuan Qin, Weiran Wang, Yuming Yang, Tao Gui

The study investigates the efficacy of pre-trained language models (PLMs) in analyzing argumentative moves in a longitudinal learner corpus. Prior studies on argumentative moves often rely on qualitative analysis and man…

Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild

2025-05-21 · Sheshera Mysore, Debarati Das, Hancheng Cao, Bahareh Sarrafzadeh

As large language models (LLMs) are used in complex writing workflows, users engage in multi-turn interactions to steer generations to better fit their needs. Rather than passively accepting output, users actively refine…

LLMs as Writing Assistants: Exploring Perspectives on Sense of Ownership and Reasoning

2024-03-20 · Azmine Toushik Wasi, Mst Rafia Islam, Raima Islam

Sense of ownership in writing confines our investment of thoughts, time, and contribution, leading to attachment to the output. However, using writing assistants introduces a mental dilemma, as some content isn't directl…

Exploring the Limitations of Detecting Machine-Generated Text

2024-06-16 · Jad Doughman, Osama Mohammed Afzal, Hawau Olamide Toyin, Shady Shehata 외

Recent improvements in the quality of the generations by large language models have spurred research into identifying machine-generated text. Such work often presents high-performing detectors. However, humans and machin…

Text Detection

ABScribe: Rapid Exploration & Organization of Multiple Writing Variations in Human-AI Co-Writing Tasks using Large Language Models

2023-09-29 · Mohi Reza, Nathan Laundry, Ilya Musabirov, Peter Dushniku 외

Exploring alternative ideas by rewriting text is integral to the writing process. State-of-the-art Large Language Models (LLMs) can simplify writing variation generation. However, current interfaces pose challenges for s…