paper-with-me

홈 › Papers

Disentangling Learning from Judgment: Representation Learning for Open Response Analytics

2025-12-30 · Conrad Borchers, Manit Patel, Seiyon M. Lee, Anthony F. Botelho arxiv

Open-ended responses are central to learning, yet automated scoring often conflates what students wrote with how teachers grade. We present an analytics-first framework that separates content signals from rater tendencies, making judgments visible and auditable via analytics. Using de-identified ASSISTments mathematics responses, we model teacher histories as dynamic priors and represent text with sentence embeddings. We apply centroid normalization and response-problem embedding differences, and explicitly model teacher effects with priors to reduce problem- and teacher-related confounds. Temporally-validated linear models quantify the contributions of each signal, and model disagreements surface observations for qualitative inspection. Results show that teacher priors heavily influence grade predictions; the strongest results arise when priors are combined with content embeddings (AUC~0.815), while content-only models remain above chance but substantially weaker (AUC~0.626). Adjusting for rater effects sharpens the selection of features derived from content representations, retaining more informative embedding dimensions and revealing cases where semantic evidence supports understanding as opposed to surface-level differences in how students respond. The contribution presents a practical pipeline that transforms embeddings from mere features into learning analytics for reflection, enabling teachers and researchers to examine where grading practices align (or conflict) with evidence of student reasoning and learning.

📄 PDF Abstract BibTeX arXiv:2512.23941

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Deep opacity and AI: A threat to XAI and to privacy protection mechanisms

2025-08-30 · Vincent C. Müller arxiv

It is known that big data analytics and AI pose a threat to privacy, and that some of this is due to some kind of "black box problem" in AI. I explain how this becomes a problem in the context of justification for judgme…

Who Laughs with Whom? Disentangling Influential Factors in Humor Preferences across User Clusters and LLMs

2026-01-06 · Soichiro Murakami, Hidetaka Kamigaito, Hiroya Takamura, Manabu Okumura arxiv

Humor preferences vary widely across individuals and cultures, complicating the evaluation of humor using large language models (LLMs). In this study, we model heterogeneity in humor preferences in Oogiri, a Japanese cre…

uBLEU: Uncertainty-Aware Automatic Evaluation Method for Open-Domain Dialogue Systems

2020-07-01 · ACL 2020 6 · Tsuta Yuma, Naoki Yoshinaga, Masashi Toyoda

Because open-domain dialogues allow diverse responses, basic reference-based metrics such as BLEU do not work well unless we prepare a massive reference set of high-quality responses for input utterances. To reduce this …

A Unified Representation Underlying the Judgment of Large Language Models

2025-10-31 · Yi-Long Lu, Jiajun Song, Wei Wang arxiv

A central architectural question for both biological and artificial intelligence is whether judgment relies on specialized modules or a unified, domain-general resource. While the discovery of decodable neural representa…

CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models

2026-04-01 · Fan Huang, Songheng Zhang, Haewoon Kwak, Jisun An arxiv

Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making contexts. While prior work has shown that LLMs exhibit cognitive biases behaviorally, whether these biases correspond to identifiable …