paper-with-me

Papers

CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning

2025-12-21 · Zijun Gao, Zhikun Xu, Xiao Ye, Ben Zhou arxiv

Large language models (LLMs) often solve challenging math exercises yet fail to apply the concept right when the problem requires genuine understanding. Popular Reinforcement Learning with Verifiable Rewards (RLVR) pipelines reinforce final answers but provide little fine-grained conceptual signal, so models improve at pattern reuse rather than conceptual applications. We introduce CORE (Concept-Oriented REinforcement), an RL training framework that turns explicit concepts into a controllable supervision signal. Starting from a high-quality, low-contamination textbook resource that links verifiable exercises to concise concept descriptions, we run a sanity probe showing LLMs can restate definitions but fail concept-linked quizzes, quantifying the conceptual reasoning gap. CORE then (i) synthesizes concept-aligned quizzes, (ii) injects brief concept snippets during rollouts to elicit concept-primed trajectories, and (iii) reinforces conceptual reasoning via trajectory replacement after group failures, a lightweight forward-KL constraint that aligns unguided with concept-primed policies, or standard GRPO directly on concept-aligned quizzes. Across several models, CORE delivers consistent gains over vanilla and SFT baselines on both in-domain concept-exercise suites and diverse out-of-domain math benchmarks. CORE unifies direct training on concept-aligned quizzes and concept-injected rollouts under outcome regularization. It provides fine-grained conceptual supervision that bridges problem-solving competence and genuine conceptual reasoning, while remaining algorithm- and verifier-agnostic.

📄 PDF Abstract BibTeX arXiv:2512.18857

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

Term-Based Extraction of Medical Information: Pre-Operative Patient Education Use Case

2019-09-01 · RANLP 2019 9 · Martin Wolf, Volha Petukhova, Dietrich Klakow

The processing of medical information is not a trivial task for medical non-experts. The paper presents an artificial assistant designed to facilitate a reliable access to medical online contents. Interactions are modell…

Question AnsweringRetrievalTerm Extraction

The Reasons that Agents Act: Intention and Instrumental Goals

2024-02-11 · Francis Rhys Ward, Matt MacDermott, Francesco Belardinelli, Francesca Toni 외

Intention is an important and challenging concept in AI. It is important because it underlies many other concepts we care about, such as agency, manipulation, legal responsibility, and blame. However, ascribing intent to…

Philosophy

Unifying the Scope of Bridging Anaphora Types in English: Bridging Annotations in ARRAU and GUM

2024-10-02 · Lauren Levine, Amir Zeldes

Comparing bridging annotations across coreference resources is difficult, largely due to a lack of standardization across definitions and annotation schemas and narrow coverage of disparate text domains across resources.…

Defining Knowledge: Bridging Epistemology and Large Language Models

2024-10-03 · Constanza Fierro, Ruchira Dhar, Filippos Stamatiou, Nicolas Garneau 외

Knowledge claims are abundant in the literature on large language models (LLMs); but can we say that GPT-4 truly "knows" the Earth is round? To address this question, we review standard definitions of knowledge in episte…

Bridging resolution: Task definition, corpus resources and rule-based experiments

2018-08-01 · COLING 2018 8 · Ina Roesiger, Arndt Riester, Jonas Kuhn

Recent work on bridging resolution has so far been based on the corpus ISNotes (Markert et al. 2012), as this was the only corpus available with unrestricted bridging annotation. Hou et al. 2014{'}s rule-based system cur…

Coreference Resolution