paper-with-me

Papers

Evaluating LLM-Generated Q&A Test: a Student-Centered Study

2025-05-10 · Anna Wróblewska, Bartosz Grabek, Jakub Świstak, Daniel Dan

This research prepares an automatic pipeline for generating reliable question-answer (Q&A) tests using AI chatbots. We automatically generated a GPT-4o-mini-based Q&A test for a Natural Language Processing course and evaluated its psychometric and perceived-quality metrics with students and experts. A mixed-format IRT analysis showed that the generated items exhibit strong discrimination and appropriate difficulty, while student and expert star ratings reflect high overall quality. A uniform DIF check identified two items for review. These findings demonstrate that LLM-generated assessments can match human-authored tests in psychometric performance and user satisfaction, illustrating a scalable approach to AI-assisted assessment development.

📄 PDF Abstract BibTeX arXiv:2505.06591

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Is Solving Better Than Evaluating GenAI Solutions?

2026-07-30 · Ethan Dickey, Marios Mertzanidis, Alexandros Psomas arxiv

As Generative AI (GenAI) tools become increasingly capable of generating solutions to computing assignments, the computing education community is exploring pedagogical approaches that emphasize solution evaluation, verif…

Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading Efficiency

2023-10-10 · Eric Zelikman, Wanjing Anya Ma, Jasmine E. Tran, Diyi Yang 외

Developing an educational test can be expensive and time-consuming, as each item must be written by experts and then evaluated by collecting hundreds of student responses. Moreover, many tests require multiple distinct s…

Language ModelingLanguage ModellingSentence

SuperSkillsStack: Agency, Domain Knowledge, Imagination, and Taste in Human-AI Design Education

2026-03-07 · Qian Huang, King Wang Poon arxiv

This study examines how students integrate generative artificial intelligence (AI) into design projects through the lens of the SuperSkillsStack framework, which identifies four key human competencies for effective human…

Humanizing AI Grading: Student-Centered Insights on Fairness, Trust, Consistency and Transparency

2026-02-08 · Bahare Riahi, Viktoriia Storozhevykh, Veronica Catete arxiv

This study investigates students' perceptions of Artificial Intelligence (AI) grading systems in an undergraduate computer science course (n = 27), focusing on a block-based programming final project. Guided by the ethic…

How Real Is AI Tutoring? Comparing Simulated and Human Dialogues in One-on-One Instruction

2025-09-02 · Ruijia Li, Yuan-Hao Jiang, Jiatong Wang, Bo Jiang arxiv

Heuristic and scaffolded teacher-student dialogues are widely regarded as critical for fostering students' higher-order thinking and deep learning. However, large language models (LLMs) currently face challenges in gener…