paper-with-me

홈 › Papers

Leveraging BART to Assess CS1 C++ Programming Assignments using Rubric-based Criteria

2026-06-02 · Kelsey Rainey, Jesse Roberts arxiv

This paper investigates rubric-aware, multitask fine-tuning of transformer models for automated grading of introductory C++ programming assignments, with the goal of producing grade predictions that better reflect instructor grading behavior than general-purpose LLMs. Using multi-semester CS1 data, student submissions are paired with numeric scores, letter-grade buckets, and assignment rubrics, then preprocessed into unified sequences for transformer input. A BART encoder-decoder with LoRA adaptation is trained to jointly predict numeric grades and grade buckets, augmented with a distribution-matching term to align predicted and empirical grade distributions, an evaluation dimension often overlooked in prior work. Experiments compare single-task and multitask training, hard one-hot versus fuzzy and boundary-based soft labels, and rubric versus no-rubric conditions, with additional T5 and pairwise-pretrained variants. Results show that multitask BART with boundary-based soft labels and rubric context achieves lower mean absolute error and stronger grade-distribution alignment than single-task, hard-label, or code-only baselines. Fully fine-tuned T5 further improves distributional fidelity, while pairwise pretraining reduces numeric error at the cost of minority-class sensitivity. Collectively, the findings suggest that calibration-aware, rubric-guided training produces more instructor-like grading behavior than accuracy-optimized alternatives.

📄 PDF Abstract BibTeX arXiv:2606.03814

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rubric Reliability and Annotation of Content and Argument in Source-Based Argument Essays

2019-08-01 · WS 2019 8 · Yanjun Gao, Alex Driban, Brennan Xavier McManus, Elena Musi 외

We present a unique dataset of student source-based argument essays to facilitate research on the relations between content, argumentation skills, and assessment. Two classroom writing assignments were given to college s…

AGACCI : Affiliated Grading Agents for Criteria-Centric Interface in Educational Coding Contexts

2025-07-07 · Kwangsuk Park, Jiwoong Yang arxiv

Recent advances in AI-assisted education have encouraged the integration of vision-language models (VLMs) into academic assessment, particularly for tasks that require both quantitative and qualitative evaluation. Howeve…

Zero Shot Learning for Code Education: Rubric Sampling with Deep Learning Inference

2018-09-05 · Mike Wu, Milan Mosse, Noah Goodman, Chris Piech

In modern computer science education, massive open online courses (MOOCs) log thousands of hours of data about how students solve coding challenges. Being so rich in data, these platforms have garnered the interest of th…

MisconceptionsZero-Shot Learning

Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics

2025-03-31 · Aditya Pathak, Rachit Gandhi, Vaibhav Uttam, Devansh 외

Since the emergence of Large Language Models (LLMs) popularized by the release of GPT-3 and ChatGPT, LLMs have shown remarkable promise in programming-related tasks. While code generation using LLMs has become a popular …

AllCode Generation

Building an Effective Automated Assessment System for C/C++ Introductory Programming Courses in ODL Environment

2022-05-24 · Muhammad Salman Khan, Adnan Ahmad, Muhammad Humayoun

Assessments help in evaluating the knowledge gained by a learner at any specific point as well as in continuous improvement of the curriculum design and the whole learning process. However, with the increase in students'…