paper-with-me

홈 › Papers

REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading

2026-04-24 · Chengshuai Zhao, Fan Zhang, Kumar Satvik Chaudhary, Yiwen Li, Lo Pang-Yun Ting, Ying-Chih Chen, Huan Liu arxiv

Open-ended grading is central to equitable and personalized education, yet manual grading remains time-consuming and costly, underscoring the need for automated grading systems. Although recent neural and large language model (LLM) based systems have demonstrated superior performance, they are typically black-box models whose scoring processes and rationales are difficult for educators to verify and trust. Concept bottleneck models (CBMs) have emerged as a promising approach by routing predictions through human-interpretable concepts, providing a mechanistic guarantee of transparency. However, standard CBMs are not tailored to open-ended grading: they do not explicitly model fine-grained rubric dimensions, inadequately capture the ordinal semantics of scoring scales, and neglect inherent reliability issues in human concept annotations. To address these limitations, we propose REC-CBM, a rubric-aware error-correction concept bottleneck model for trustworthy open-ended grading. REC-CBM introduces a rubric-aware concept encoder that learns concept-specific representations over responses and an ordinal pairwise calibration objective that preserves ranking structure among rubric dimensions. It further incorporates a latent concept error-correction module that denoises concept predictions before final grade prediction while preserving interpretability. Comprehensive experiments on publicly available datasets show that REC-CBM consistently improves grading performance and produces more faithful concept-level reasoning than both state-of-the-art baselines. Further analyses validate the contribution of each component and demonstrate the applicability in realistic educational settings. Overall, this work provides a practical, interpretable grading solution that enables educators to inspect, intervene in, and trust automated decisions, advancing more transparent and trustworthy education.

📄 PDF Abstract BibTeX arXiv:2605.27402

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EssayCBM: Rubric-Aligned Concept Bottleneck Models for Transparent Essay Grading

2025-12-23 · Kumar Satvik Chaudhary, Chengshuai Zhao, Fan Zhang, Garima Agrawal 외 arxiv

Automated essay scoring (AES) has advanced significantly with neural language models, yet most systems remain opaque, offering little visibility into how grades are produced. In educational settings, instructors must be …

Automated Essay Scoring

Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization

2026-05-28 · Yun Wang, Xin Xia, Xuansheng Wu, Xiaoming Zhai 외 arxiv

LLM-based automated scoring approaches near-human performance, but scaling to new tasks remains bottlenecked by the per-item human configuration of upstream stages such as rubric construction. Human experts bypass this b…

How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos

2026-04-16 · Olga Loginova, Frank Keller arxiv

Reliable procedural monitoring in video requires exposure to naturally occurring human errors and the recoveries that follow. In egocentric recordings, mistakes are often partially occluded by hands and revealed through …

Video Generation

Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring

2025-05-01 · Jayoung Song, Kyungtae Lim, Jungyeul Park

Despite growing global interest in Korean language education, there remains a significant lack of learner corpora tailored to Korean L2 writing. To address this gap, we enhance the KoLLA Korean learner corpus by adding m…

DiversityGrammatical Error Correction

Hyperbolic Concept Bottleneck Models

2026-05-07 · Daniel Uyterlinde, Swasti Shreya Mishra, Pascal Mettes arxiv

Concept Bottleneck Models (CBMs) have become a popular approach to enable interpretability in neural networks by constraining classifier inputs to a set of human-understandable concepts. While effective, current models e…