paper-with-me

홈 › Papers

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information

2025-08-15 · Youcheng Huang, Bowen Qin, Chen Huang, Duanyu Feng, Xi Yang, Wenqiang Lei arxiv

Large Reasoning Models (LRMs) have demonstrated remarkable problem-solving abilities in mathematics, as evaluated by existing benchmarks exclusively on well-defined problems. However, such evaluation setup constitutes a critical gap, since a genuine intelligent agent should not only solve problems (as a math quiz solver), but also be able~to ask for information when the problems lack sufficient information, enabling proactivity in responding users' requests. To bridge such gap, we proposes a new dataset consisting of two types of incomplete problems with diverse contexts. Based on the dataset, our systematical evaluation of LRMs reveals their inability in proactively asking for information. In addition, we uncover the behaviors related to overthinking and hallucination of LRMs, and highlight the potential and challenges of supervised fine-tuning in learning such ability. We hope to provide new insights in developing LRMs with genuine intelligence, rather than just solving problems.

📄 PDF Abstract BibTeX arXiv:2508.11252

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark

2026-03-01 · Zhiqi Yu, Xingping Liu, Haobin Mao, Mingshuo Liu 외 arxiv

Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, handwritten single-variable calculus work fr…

Mathematical ReasoningCollege Mathematics

CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning

2025-12-21 · Zijun Gao, Zhikun Xu, Xiao Ye, Ben Zhou arxiv

Large language models (LLMs) often solve challenging math exercises yet fail to apply the concept right when the problem requires genuine understanding. Popular Reinforcement Learning with Verifiable Rewards (RLVR) pipel…

Reinforcement LearningMathematical Reasoning

Can an AI Win Ghana's National Science and Maths Quiz? An AI Grand Challenge for Education

2023-01-30 · George Boateng, Victor Kumbol, Elsie Effah Kaufmann

There is a lack of enough qualified teachers across Africa which hampers efforts to provide adequate learning support such as educational question answering (EQA) to students. An AI system that can enable students to ask…

MathPositionQuestion Answering

Towards an AI to Win Ghana's National Science and Maths Quiz

2023-08-08 · George Boateng, Jonathan Abrefah Mensah, Kevin Takyi Yeboah, William Edor 외

Can an AI win Ghana's National Science and Maths Quiz (NSMQ)? That is the question we seek to answer in the NSMQ AI project, an open-source project that is building AI to compete live in the NSMQ and win. The NSMQ is an …

MathQuestion AnsweringSpeech-to-Texttext-to-speech+1

Automating Turkish Educational Quiz Generation Using Large Language Models

2024-06-05 · Kamyar Zeinalipour, Yusuf Gökberk Keptiğ, Marco Maggini, Marco Gori

Crafting quizzes from educational content is a pivotal activity that benefits both teachers and students by reinforcing learning and evaluating understanding. In this study, we introduce a novel approach to generate quiz…

Multiple-choice