paper-with-me

홈 › Papers

Towards Human-Like Grading: A Unified LLM-Enhanced Framework for Subjective Question Evaluation

2025-10-09 · Fanwei Zhua, Jiaxuan He, Xiaoxiao Chen, Zulong Chen, Quan Lu, Chenrui Mei arxiv

Automatic grading of subjective questions remains a significant challenge in examination assessment due to the diversity in question formats and the open-ended nature of student responses. Existing works primarily focus on a specific type of subjective question and lack the generality to support comprehensive exams that contain diverse question types. In this paper, we propose a unified Large Language Model (LLM)-enhanced auto-grading framework that provides human-like evaluation for all types of subjective questions across various domains. Our framework integrates four complementary modules to holistically evaluate student answers. In addition to a basic text matching module that provides a foundational assessment of content similarity, we leverage the powerful reasoning and generative capabilities of LLMs to: (1) compare key knowledge points extracted from both student and reference answers, (2) generate a pseudo-question from the student answer to assess its relevance to the original question, and (3) simulate human evaluation by identifying content-related and non-content strengths and weaknesses. Extensive experiments on both general-purpose and domain-specific datasets show that our framework consistently outperforms traditional and LLM-based baselines across multiple grading metrics. Moreover, the proposed system has been successfully deployed in real-world training and certification exams at a major e-commerce enterprise.

📄 PDF Abstract BibTeX arXiv:2510.07912

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AceTone: Bridging Words and Colors for Conditional Image Grading

2026-04-01 · Tianren Ma, Mingxiang Liao, Xijin Zhang, Qixiang Ye arxiv

Color affects how we interpret image style and emotion. Previous color grading methods rely on patch-wise recoloring or fixed filter banks, struggling to generalize across creative intents or align with human aesthetic p…

Reinforcement Learning

A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization

2024-10-03 · Yucheng Chu, Hang Li, Kaiqi Yang, Harry Shomer 외

Open-ended short-answer questions (SAGs) have been widely recognized as a powerful tool for providing deeper insights into learners' responses in the context of learning analytics (LA). However, SAGs often present challe…

automatic short answer grading

KIEGLFN: A unified acne grading framework on face images

2022-06-01 · Computer Methods and Programs in Biomedicine 2022 6 · Yi Lin, Jingchi Jiang, Zhaoyang Ma, Dongxin Chen 외

Grading the severity level is an extremely important procedure for correct diagnoses and personalized treatment schemes for acne. However, the acne grading criteria are not unified in the medical field. This work aims to…

Acne Severity Grading

Beyond human subjectivity and error: a novel AI grading system

2024-05-07 · Alexandra Gobrecht, Felix Tuma, Moritz Möller, Thomas Zöller 외

The grading of open-ended questions is a high-effort, high-impact task in education. Automating this task promises a significant reduction in workload for education professionals, as well as more consistent grading outco…

automatic short answer gradingFairness

Pre-Training BERT on Domain Resources for Short Answer Grading

2019-11-01 · IJCNLP 2019 11 · Chul Sung, Tejas Dhamecha, Swarnadeep Saha, Tengfei Ma 외

Pre-trained BERT contextualized representations have achieved state-of-the-art results on multiple downstream NLP tasks by fine-tuning with task-specific data. While there has been a lot of focus on task-specific fine-tu…

automatic short answer gradingLanguage ModelingLanguage Modelling