paper-with-me

홈 › Papers

Know Thy Reasoner: Not All Language Models Explore Alike

2026-04-12 · Moulik Choraria, Argyrios Gerogiannis, Anirban Das, Supriyo Chakraborty, Sourya Basu, Sambit Sahu, Lav R. Varshney arxiv

Compute scaling for LLM reasoning trades off exploring solution approaches (\emph{breadth}) against refining promising ones (\emph{depth}), yet why a given trade-off works, and why it often fails to transfer across models, remains unclear. We argue that \textbf{the optimal strategy depends on the model's \emph{diversity profile}, the spread of probability mass across solution approaches, and that this must be characterized before any exploration strategy is adopted.} We formalize this with a framework decomposing reasoning uncertainty, deriving when depth-based refinement outperforms parallel sampling, and validate it across three model families at both inference and training. Our central finding is that the diversity regime dictates the strategy: low-diversity aligned models benefit from depth-based refinement with lightweight intrinsic signals, whereas high-diversity base models are often harmed by it, and instead need breadth or stronger signals to compensate.

📄 PDF Abstract BibTeX arXiv:2604.10827

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

KG-Reasoner: A Reinforced Model for End-to-End Multi-Hop Knowledge Graph Reasoning

2026-04-14 · Shuai Wang, Yinan Yu arxiv

Large Language Models (LLMs) exhibit strong abilities in natural language understanding and generation, yet they struggle with knowledge-intensive reasoning. Structured Knowledge Graphs (KGs) provide an effective form of…

Knowledge Base Question AnsweringNatural Language UnderstandingReinforcement LearningKnowledge Graphs

LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?

2025-10-09 · Jingyuan Wang, Yankai Chen, Zhonghang Li, Chao Huang arxiv

Large language models (LLMs) have demonstrated remarkable progress in reasoning, often through supervised fine-tuning (SFT). However, SFT is resource-intensive, relying on large curated datasets, rejection-sampled demons…

Physics Reasoner: Knowledge-Augmented Reasoning for Solving Physics Problems with Large Language Models

2024-12-18 · Xinyu Pang, Ruixin Hong, Zhanke Zhou, Fangrui Lv 외

Physics problems constitute a significant aspect of reasoning, necessitating complicated reasoning ability and abundant physics knowledge. However, existing large language models (LLMs) frequently fail due to a lack of k…

PathReasoner-R1: Instilling Structured Reasoning into Pathology Vision-Language Model via Knowledge-Guided Policy Optimization

2026-01-29 · Songhan Jiang, Fengchun Liu, Ziyue Wang, Linghan Cai 외 arxiv

Vision-Language Models (VLMs) are advancing computational pathology with superior visual understanding capabilities. However, current systems often reduce diagnosis to directly output conclusions without verifiable evide…

Reinforcement LearningKnowledge Graphs

MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs

2025-08-04 · Guojiang Zhao, Zixiang Lu, Yutang Ge, Sihang Li 외 arxiv

Large Language Models (LLMs) have shown impressive performance across various domains, but their ability to perform molecular reasoning remains underexplored. Existing methods mostly rely on general-purpose prompting, wh…