paper-with-me

홈 › Papers

Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment

2024-03-03 · Jingshen Zhang, Jiajun Xie, Xinying Qiu

Item difficulty plays a crucial role in adaptive testing. However, few works have focused on generating questions of varying difficulty levels, especially for multiple-choice (MC) cloze tests. We propose training pre-trained language models (PLMs) as surrogate models to enable item response theory (IRT) assessment, avoiding the need for human test subjects. We also propose two strategies to control the difficulty levels of both the gaps and the distractors using ranking rules to reduce invalid distractors. Experimentation on a benchmark dataset demonstrates that our proposed framework and methods can effectively control and evaluate the difficulty levels of MC cloze tests.

📄 PDF Abstract BibTeX arXiv:2403.01456

Code (0)

등록된 구현이 없습니다.

Tasks

Cloze TestMultiple-choice

Similar Papers 제목 키워드 기반

Improving the Production Efficiency and Well-formedness of Automatically-Generated Multiple-Choice Cloze Vocabulary Questions

2020-05-01 · LREC 2020 5 · Ralph Rose

Multiple-choice cloze (fill-in-the-blank) questions are widely used in knowledge testing and are commonly used for testing vocabulary knowledge. Word Quiz Constructor (WQC) is a Java application that is designed to produ…

Multiple-choice

Automated Generation of Multiple-Choice Cloze Questions for Assessing English Vocabulary Using GPT-turbo 3.5

2024-03-04 · Qiao Wang, Ralph Rose, Naho Orita, Ayaka Sugawara

A common way of assessing language learners' mastery of vocabulary is via multiple-choice cloze (i.e., fill-in-the-blank) questions. But the creation of test items can be laborious for individual teachers or in large-sca…

Multiple-choicePart-Of-Speech TaggingSentence

Difficulty-Controllable Cloze Question Distractor Generation

2025-11-03 · Seokhoon Kang, Yejin Jeon, Seonjeong Hwang, Gary Geunbae Lee arxiv

Multiple-choice cloze questions are commonly used to assess linguistic proficiency and comprehension. However, generating high-quality distractors remains challenging, as existing methods often lack adaptability and cont…

Distractor GenerationData Augmentation

A System for Generating Cloze Test Items from Russian-Language Text

2013-09-01 · RANLP 2013 9 · Andrey Kurtasov
Cloze Test

Reliable and Efficient Amortized Model-based Evaluation

2025-03-17 · Sang Truong, Yuheng Tu, Percy Liang, Bo Li 외

Comprehensive evaluations of language models (LM) during both development and deployment phases are necessary because these models possess numerous capabilities (e.g., mathematical reasoning, legal support, or medical di…

DiagnosticMathematical ReasoningMisinformationmodel