paper-with-me

Papers

RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning

2025-11-06 · Xinyuan Li, Murong Xu, Wenbiao Tao, Hanlun Zhu, Yike Zhao, Jipeng Zhang, Yunshi Lan arxiv

Large language models (LLMs) achieve high performance on mathematical reasoning, but these results can be inflated by training data leakage or superficial pattern matching rather than genuine reasoning. To this end, an adversarial perturbation-based evaluation is needed to measure true mathematical reasoning ability. Current rule-based perturbation methods often generate ill-posed questions and impede the systematic evaluation of question difficulty and the evolution of benchmarks. To bridge this gap, we propose RIDE, a novel adversarial question-rewriting framework that leverages Item Response Theory (IRT) to rigorously measure question difficulty and to generate intrinsically more challenging, well-posed variations of mathematical problems. We employ 35 LLMs to simulate students and build a difficulty ranker from their responses. This ranker provides a reward signal during reinforcement learning and guides a question-rewriting model to reformulate existing questions across difficulty levels. Applying RIDE to competition-level mathematical benchmarks yields perturbed versions that degrade advanced LLM performance, with experiments showing an average 21.73% drop across 26 models, thereby exposing limited robustness in mathematical reasoning and confirming the validity of our evaluation approach.

📄 PDF Abstract BibTeX arXiv:2511.04120

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

Predicting the Difficulty and Response Time of Multiple Choice Questions Using Transfer Learning

2020-07-01 · WS 2020 7 · Kang Xue, Victoria Yaneva, Christopher Runyon, Peter Baldwin

This paper investigates whether transfer learning can improve the prediction of the difficulty and response time parameters for 18,000 multiple-choice questions from a high-stakes medical exam. The type the signal that b…

Multiple-choiceTransfer Learning

SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction

2025-07-07 · Alexander Scarlatos, Nigel Fernandez, Christopher Ormerod, Susan Lottridge 외

Item (question) difficulties play a crucial role in educational assessments, enabling accurate and efficient assessment of student abilities and personalization to maximize learning outcomes. Traditionally, estimating it…

Learning Latent Parameters without Human Response Patterns: Item Response Theory with Artificial Crowds

2019-08-29 · IJCNLP 2019 11 · John P. Lalor, Hao Wu, Hong Yu

Incorporating Item Response Theory (IRT) into NLP tasks can provide valuable information about model performance and behavior. Traditionally, IRT models are learned using human response pattern (RP) data, presenting a si…

Natural Language InferenceSentiment Analysis

Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory

2026-04-30 · Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus 외 arxiv

Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa. However, these metrics provide limited insight into how grading per…

$β^3$-IRT: A New Item Response Model and its Applications

2019-03-10 · Yu Chen, Telmo Silva Filho, Ricardo B. C. Prudêncio, Tom Diethe 외

Item Response Theory (IRT) aims to assess latent abilities of respondents based on the correctness of their answers in aptitude test items with different difficulty levels. In this paper, we propose the $\beta^3$-IRT mod…