paper-with-me

Papers

LLMs cannot spot math errors, even when allowed to peek into the solution

2025-09-01 · KV Aditya Srivatsa, Kaushal Kumar Maurya, Ekaterina Kochmar arxiv

Large language models (LLMs) demonstrate remarkable performance on math word problems, yet they have been shown to struggle with meta-reasoning tasks such as identifying errors in student solutions. In this work, we investigate the challenge of locating the first error step in stepwise solutions using two error reasoning datasets: VtG and PRM800K. Our experiments show that state-of-the-art LLMs struggle to locate the first error step in student solutions even when given access to the reference solution. To that end, we propose an approach that generates an intermediate corrected student solution, aligning more closely with the original student's solution, which helps improve performance.

📄 PDF Abstract BibTeX arXiv:2509.01395

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors

2025-10-02 · Dane Williamson, Yangfeng Ji, Matthew Dwyer arxiv

Large Language Models (LLMs) demonstrate strong mathematical problem-solving abilities but frequently fail on problems that deviate syntactically from their training distribution. We identify a systematic failure mode, s…

Discovering Blind Spots in Reinforcement Learning

2018-05-23 · Ramya Ramakrishnan, Ece Kamar, Debadeepta Dey, Julie Shah 외

Agents trained in simulation may make errors in the real world due to mismatches between training and execution environments. These mistakes can be dangerous and difficult to discover because the agent cannot predict the…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

2025-07-03 · Ken Tsui arxiv

Although large language models (LLMs) have transformed AI, they still make mistakes and can explore unproductive reasoning paths. Self-correction capability is essential for deploying LLMs in safety-critical applications…

Reinforcement Learning

ESCAPE: Countering Systematic Errors from Machine's Blind Spots via Interactive Visual Analysis

2023-03-16 · Yongsu Ahn, Yu-Ru Lin, Panpan Xu, Zeng Dai

Classification models learn to generalize the associations between data samples and their target classes. However, researchers have increasingly observed that machine learning practice easily leads to systematic errors i…

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

2025-05-17 · Guijin Son, Jiwoo Hong, Honglu Fan, Heejeong Nam 외

Recent advances in large language models (LLMs) have fueled the vision of automated scientific discovery, often called AI Co-Scientists. To date, prior work casts these systems as generative co-authors responsible for cr…

Misconceptionsscientific discovery