paper-with-me

홈 › Papers

AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment

2026-08-01 · Garv Vikram Gursahaney, Baskhad Idrisov, Thorsten Fröhlich, Tim Schlippe arxiv

Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although little empirical evidence exists regarding how thesis supervisors actually prioritize evaluation criteria. Consequently, this study investigates supervisor-derived criterion weights in thesis assessment and evaluates their impact on AI-based assessment. We surveyed 84 thesis supervisors across four academic disciplines and collected weighting data for 35 thesis assessment criteria. Comparison with the default criterion weights of the AI assessment system RubiSCoT [1] revealed substantial divergences between supervisor-derived and default criterion weights. To evaluate the practical implications of these differences, the supervisor-derived weights were integrated into multiple calibration configurations and evaluated on a corpus of 80 German-language theses. The best-performing configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, although the improvement was not statistically significant. Human supervisors showed substantially stronger agreement with each other, exhibiting a mean inter-supervisor relative deviation of 4.44%. The findings indicate that criterion-weight calibration alone does not substantially improve alignment between AI-generated and human assessments.

📄 PDF Abstract BibTeX arXiv:2608.00717

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Augmenting Dysarthric Speech Severity Assessment with MOS Supervision

2026-06-17 · Kaimeng Jia, Minzhu Tu, Zengrui Jin, Siyin Wang 외 arxiv

Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness. Automatic utterance-level assessment of dysarthric speech can support scalable speech monitoring and therapy-related anal…

Speech Synthesis

The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges

2024-11-23 · Jiqun Liu, Jiangen He

Can AI be cognitively biased in automated information judgment tasks? Despite recent progresses in measuring and mitigating social and algorithmic biases in AI and large language models (LLMs), it is not clear to what ex…

Information RetrievalMisinformation

Empirical Study of Quality Image Assessment for Synthesis of Fetal Head Ultrasound Imaging with DCGANs

2022-06-01 · Thea Bautista, Jacqueline Matthew, Hamideh Kerdegari, Laura Peralta Pereira 외

In this work, we present an empirical study of DCGANs, including hyperparameter heuristics and image quality assessment, as a way to address the scarcity of datasets to investigate fetal head ultrasound. We present exper…

Image Quality Assessment

PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review

2026-02-01 · Yanki Margalit, Erni Avram, Ran Taig, Oded Margalit 외 arxiv

Evaluating large language models typically relies on human-authored benchmarks, reference answers, and human or single-model judgments, approaches that scale poorly, become quickly outdated, and mismatch open-world deplo…

Ethical Asymmetry in Human-Robot Interaction - An Empirical Test of Sparrow's Hypothesis

2026-02-02 · Minyi Wang, Christoph Bartneck, Michael-John Turp, David Kaber arxiv

The ethics of human-robot interaction (HRI) have been discussed extensively based on three traditional frameworks: deontology, consequentialism, and virtue ethics. We conducted a mixed within/between experiment to invest…

Moral Permissibility