paper-with-me

홈 › Papers

Reasoning over mathematical objects: on-policy reward modeling and test time aggregation

2026-03-19 · Pranjal Aggarwal, Marjan Ghazvininejad, Seungone Kim, Ilia Kulikov, Jack Lanchantin, Xian Li, Tianjian Li, Bo Liu, Graham Neubig, Anaelia Ovalle, Swarnadeep Saha, Sainbayar Sukhbaatar, Sean Welleck, Jason Weston, Chenxi Whitehouse, Adina Williams, Jing Xu, Ping Yu, Weizhe Yuan, Jingyu Zhang, Wenting Zhao arxiv

The ability to precisely derive mathematical objects is a core requirement for downstream STEM applications, including mathematics, physics, and chemistry, where reasoning must culminate in formally structured expressions. Yet, current LM evaluations of mathematical and scientific reasoning rely heavily on simplified answer formats such as numerical values or multiple choice options due to the convenience of automated assessment. In this paper we provide three contributions for improving reasoning over mathematical objects: (i) we build and release training data and benchmarks for deriving mathematical objects, the Principia suite; (ii) we provide training recipes with strong LLM-judges and verifiers, where we show that on-policy judge training boosts performance; (iii) we show how on-policy training can also be used to scale test-time compute via aggregation. We find that strong LMs such as Qwen3-235B and o3 struggle on Principia, while our training recipes can bring significant improvements over different LLM backbones, while simultaneously improving results on existing numerical and MCQA tasks, demonstrating cross-format generalization of reasoning abilities.

📄 PDF Abstract BibTeX arXiv:2603.18886

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rewarding Structural Conformance of Reasoning using Process Mining

2025-10-29 · Yongjae Lee, Taekhyun Park, Sunghyun Sim, Hyerim Bae arxiv

Recent advances in sparse reward policy gradient methods have enabled effective reinforcement learning (RL)-based language model post-training. However, for reasoning tasks such as mathematical problem solving, binarized…

Reinforcement LearningMathematical Reasoning

Adaptive Generate-Rank-Verify: Inference-Time Search with Costly Verification

2026-05-17 · Shaddin Dughmi, Mahdi Haghifam, Yusuf Hakan Kalayci arxiv

Many inference-time language-model pipelines combine a cheap reward signal with an expensive verifier, such as exact answer checking in mathematical reasoning or hidden-test execution in code generation. We formalize thi…

Mathematical ReasoningCode Generation

SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

2026-08-18 · Guozheng Sun arxiv

Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, t…

Reinforcement LearningMathematical Reasoning

Triviality Corrected Endogenous Reward

2026-04-13 · Xinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan 외 arxiv

Reinforcement learning for open-ended text generation is constrained by the lack of verifiable rewards, necessitating reliance on judge models that require either annotated data or powerful closed-source models. Inspired…

Reinforcement LearningMathematical ReasoningText Generation

Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning

2025-12-18 · Qihao Liu, Luoxin Ye, Wufei Ma, Yu-Cheng Chou 외 arxiv

Large language models (LLMs) with explicit reasoning capabilities excel at mathematical reasoning yet still commit process errors, such as incorrect calculations, brittle logic, and superficially plausible but invalid st…

Reinforcement LearningMathematical Reasoning