paper-with-me

홈 › Papers

Diverse Inference and Verification for Advanced Reasoning

2025-02-14 · Iddo Drori, Gaston Longhitano, Mao Mao, Seunghwan Hyun, Yuke Zhang, Sungjun Park, Zachary Meeks, Xin-Yu Zhang, Ben Segev, Howard Yong, Nakul Verma, Avi Shporer, Alon Amit, Madeleine Udell

Reasoning LLMs such as OpenAI o1, o3 and DeepSeek R1 have made significant progress in mathematics and coding, yet find challenging advanced tasks such as International Mathematical Olympiad (IMO) combinatorics problems, Abstraction and Reasoning Corpus (ARC) puzzles, and Humanity's Last Exam (HLE) questions. We use a diverse inference approach that combines multiple models and methods at test time. We find that verifying mathematics and code problems, and rejection sampling on other problems is simple and effective. We automatically verify correctness of solutions to IMO problems by Lean, and ARC puzzles by code, and find that best-of-N effectively answers HLE questions. Our approach increases answer accuracy on IMO combinatorics problems from 33.3% to 77.8%, accuracy on HLE questions from 8% to 37%, and solves 80% of ARC puzzles that 948 humans could not and 26.5% of ARC puzzles that o3 high compute does not. Test-time simulations, reinforcement learning, and meta-learning with inference feedback improve generalization by adapting agent graph representations and varying prompts, code, and datasets. Our approach is reliable, robust, and scalable, and in the spirit of reproducible research, we will make it publicly available upon publication.

📄 PDF Abstract BibTeX arXiv:2502.09955

Code (0)

등록된 구현이 없습니다.

Tasks

ARCHumanity's Last ExamMeta-Learning

Similar Papers 제목 키워드 기반

VerifiAgent: a Unified Verification Agent in Language Model Reasoning

2025-04-01 · Jiuzhou Han, Wray Buntine, Ehsan Shareghi

Large language models demonstrate remarkable reasoning capabilities but often produce unreliable or incorrect responses. Existing verification methods are typically model-specific or domain-restricted, requiring signific…

Language ModelingLanguage ModellingMathematical Reasoning

ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated Verification

2026-01-28 · Siran Liu, Cyril Y. He arxiv

Chain-of-Thought reasoning significantly improves the performance of large language models on complex tasks, but incurs high inference latency due to long generation traces. Step-level speculative reasoning aims to mitig…

LogiDynamics: Unraveling the Dynamics of Logical Inference in Large Language Model Reasoning

2025-02-16 · Tianshi Zheng, Jiayang Cheng, Chunyang Li, Haochen Shi 외

Modern large language models (LLMs) employ various forms of logical inference, both implicitly and explicitly, when addressing reasoning tasks. Understanding how to optimally leverage these inference paradigms is critica…

Analogical questionsIn-Context LearningLanguage ModelingLanguage Modelling+3

AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification

2025-02-17 · Xiaoyu Tan, Tianchu Yao, Chao Qu, Bin Li 외

The reasoning capabilities of advanced large language models (LLMs) like o1 have revolutionized artificial intelligence applications. Nevertheless, evaluating and optimizing complex reasoning processes remain significant…

Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights

2025-02-18 · Shubham Parashar, Blake Olson, Sambhav Khurana, Eric Li 외

We examine the reasoning and planning capabilities of large language models (LLMs) in solving complex tasks. Recent advances in inference-time techniques demonstrate the potential to enhance LLM reasoning without additio…

Arithmetic ReasoningCommon Sense ReasoningLogical Reasoning