paper-with-me

홈 › Papers

Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA

2025-09-30 · Raphael Schumann, Stefan Riezler arxiv

Reasoning quality in large language models depends not only on producing correct answers but also on generating valid intermediate steps. We study this through multiple-choice question answering (MCQA), which provides a controlled setting with fixed answer options. Our analysis shows that when questions are effectively unsolvable for a model, spurious chains of thought (CoTs) are more likely to appear, leading to false positives. By estimating the solvability of each question, we uncover an intermediate regime where learning is most effective. Building on this insight, we adapt outcome-supervised reward models and reinforcement learning with group-relative advantage to incorporate solvability into their objectives. Across experiments on math and multimodal datasets, these modifications consistently yield higher rates of process-correct reasoning and, in reinforcement learning, improved answer accuracy as well. Our results highlight solvability as a key factor for reducing hallucinations and increasing reliability in CoT reasoning.

📄 PDF Abstract BibTeX arXiv:2509.25941

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningQuestion Answering

Similar Papers 제목 키워드 기반

MiST: Understanding the Role of Mid-Stage Scientific Training in Developing Chemical Reasoning Models

2025-12-24 · Andres M Bran, Tong Xie, Shai Pranesh, Jeffrey Meng 외 arxiv

Large Language Models can develop reasoning capabilities through online fine-tuning with rule-based rewards. However, recent studies reveal a critical constraint: reinforcement learning succeeds only when the base model …

Reinforcement Learning

More Capable, Less Faithful: A Multilingual Analysis of Mathematical (Un)Solvability Detection in LLMs

2026-08-31 · Maria-Eleni Zoumpoulidi, Nikolaos Xiros, Georgios Paraskevopoulos arxiv

Solvability detection is one of the most challenging aspects of mathematical reasoning for Large Language Models (LLMs). While prior work has studied this capability extensively, these analyses have been limited to Engli…

Mathematical Reasoning

Code-driven Number Sequence Calculation: Enhancing the inductive Reasoning Abilities of Large Language Models

2025-10-16 · Kedi Chen, Zhikai Lei, Xu Guo, Xuecheng Wu 외 arxiv

Large language models (LLMs) make remarkable progress in reasoning tasks. Among different reasoning modes, inductive reasoning, due to its better alignment with human learning, attracts increasing interest. However, rese…

Reinforcement Learning

Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems

2025-12-01 · Dengyun Peng, Qiguang Chen, Bofei Liu, Jiannan Guan 외 arxiv

Ensuring large language model (LLM) reliability requires distinguishing objective unsolvability (inherent contradictions) from subjective capability limitations (tasks exceeding model competence). Current LLMs often conf…

Reinforcement Learning

Knowledge Knows, Verbalization Tells: Disentangling Latent Directions for Mathematical Solvability in LLMs

2026-07-06 · Nikolaos Xiros, Maria-Eleni Zoumpoulidi, Georgios Paraskevopoulos arxiv

Although LLMs have made significant progress in mathematical reasoning, determining whether a mathematical problem is solvable remains a fundamental yet challenging capability. While recent studies have probed internal r…

Mathematical Reasoning