paper-with-me

Papers

UpSkill: Mutual Information Skill Learning for Structured Response Diversity in LLMs

2026-02-25 · Devan Shah, Owen Yang, Daniel Yang, Chongyi Zheng, Benjamin Eysenbach arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning abilities of large language models (LLMs) on mathematics and programming tasks, but standard approaches that optimize single-attempt accuracy can inadvertently suppress response diversity across repeated attempts, narrowing exploration and overlooking underrepresented strategies. We introduce UpSkill, a training time method that adapts Mutual Information Skill Learning (MISL) to LLMs for optimizing pass@k correctness. We propose a novel reward that we implement within Group Relative Policy Optimization (GRPO): a token-level mutual information (MI) reward that encourages trajectory specificity to z. Experiments on GSM8K with three open-weight models, Llama 3.1-8B, Qwen 2.5-7B, and R1-Distilled-Qwen2.5-Math-1.5B, show that UpSkill improves multi-attempt metrics on the stronger base models, yielding mean gains of ~3% in pass@k for both Qwen and Llama without degrading pass@1. Additionally, we find both empirical and theoretical evidence that improvements in pass@k are closely tied to the mutual information objective.

📄 PDF Abstract BibTeX arXiv:2602.22296

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Slot Filling for Extracting Reskilling and Upskilling Options from the Web

2022-07-11 · Albert Weichselbraun, Roger Waldvogel, Andreas Fraefel, Alexander van Schie 외

Disturbances in the job market such as advances in science and technology, crisis and increased competition have triggered a surge in reskilling and upskilling programs. Information on suitable continuing education optio…

BenchmarkingEntity Linkingslot-fillingSlot Filling

Understanding Factors that Influence Upskilling

2021-03-22 · Eduardo Laguna-Muggenburg, Monica Bhole, Michael Meaney

We investigate the motivation and means through which individuals expand their skill-set by analyzing a survey of applicants from the Facebook Jobs product. Individuals who report being influenced by their networks or lo…

Upskilling with Generative AI: Practices and Challenges for Freelance Knowledge Workers

2026-04-29 · Kashif Imteyaz, Isabel Lopez, Nakul Rajpal, Hunjun Shin 외 arxiv

Freelance workers must continually acquire new skills to remain competitive in online labor markets, yet they lack the organizational training, mentorship, and infrastructure available to traditional employees. Generativ…

Epistemic Skills: Reasoning about Knowledge and Oblivion

2025-04-02 · Xiaolong Liang, Yì N. Wáng

This paper presents a class of epistemic logics that captures the dynamics of acquiring knowledge and descending into oblivion, while incorporating concepts of group knowledge. The approach is grounded in a system of wei…

The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era

2026-04-08 · Rudra Jadhav, Janhavi Danve arxiv

As Large Language Models reshape the global labor market, policymakers and workers need empirical data on which occupational skills may be most susceptible to automation. We present the Skill Automation Feasibility Index…

Reading Comprehension