paper-with-me

홈 › Papers

Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

2025-04-08 · Qingyang Zhang, Haitao Wu, Changqing Zhang, Peilin Zhao, Yatao Bian

Existing methods to enhance the reasoning capability of large language models predominantly rely on supervised fine-tuning (SFT) followed by reinforcement learning (RL) on reasoning-specific data. These approaches critically depend on external supervisions--such as labeled reasoning traces, verified golden answers, or pre-trained reward models. In this work, we propose Entropy Minimized Policy Optimization (\ours), which makes an early attempt at fully unsupervised LLM reasoning incentivization. By continuously minimizing the predictive entropy of LLMs on unlabeled questions in a latent semantic space, \ours achieves competitive performance compared to supervised counterparts on both mathematical and free-form natural reasoning tasks. Specifically, without any supervised signals, \ours boosts the accuracy of Qwen2.5-Math-7B Base from 30.7\% to 48.1\% on mathematical benchmarks and improves the accuracy of Qwen2.5-7B Base from 32.1\% to 50.1\% on MMLU-Pro. Primary experiments and analysis are also provided to interpret the effectiveness of \ours. Code is available at https://github.com/QingyangZhang/EMPO.

📄 PDF Abstract BibTeX arXiv:2504.05812

Code (1)

qingyangzhang/empo 공식 구현 pytorch

Tasks

MathMathematical ReasoningMMLUReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Are we asking the right questions in MovieQA?

2019-11-08 · Bhavan Jasani, Rohit Girdhar, Deva Ramanan

Joint vision and language tasks like visual question answering are fascinating because they explore high-level understanding, but at the same time, can be more prone to language biases. In this paper, we explore the bias…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

2026-06-27 · Yong Yi Bay, Kathleen A. Yearick hf

People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning systems answer a hard question by sampling it many times (test-time scaling), and the more they draw, the m…

FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models

2025-04-20 · Mehrnoush Shamsfard, Zahra Saaberi, Mostafa Karimi manesh, Seyed Mohammad Hossein Hashemi 외

Research on evaluating and analyzing large language models (LLMs) has been extensive for resource-rich languages such as English, yet their performance in languages such as Persian has received considerably less attentio…

DescriptiveEthicsMultiple-choiceText Generation

LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

2026-05-27 · HuiMing Fan, Xiao Wang, Zheng Chu, Qianyu Wang 외 arxiv

Are LLM-based search agents genuinely searching, or using the web to verify what they already know? We study this question on BrowseComp with three diagnostics. Our analysis reveals Intrinsic Knowledge Dependence (IKD): …

More data speeds up training time in learning halfspaces over sparse vectors

2013-11-10 · NeurIPS 2013 12 · Amit Daniely, Nati Linial, Shai Shalev Shwartz

The increased availability of data in recent years has led several authors to ask whether it is possible to use data as a {\em computational} resource. That is, if more data is available, beyond the sample complexity lim…

PAC learning