paper-with-me

홈 › Papers

GRACE: Gradient-aligned Reasoning Data Curation for Efficient Post-training

2026-05-13 · Junjie Li, Ziao Wang, NingXuan Ma, Jianghong Ma, Xiaofeng Zhang arxiv

Existing reasoning data curation pipelines score whole samples, treating every intermediate step as equally valuable. In reality, steps within a trace contribute very unevenly, and selecting reasoning data well requires assessing them individually. We present GRACE, a gradient-aligned curation method that views each reasoning trace as a sequence of optimization events and scores every step by two complementary signals: its alignment with the answer-oriented gradient direction, and its consistency with the preceding reasoning trajectory. Step-level scores are aggregated into a sample-level value for subset selection, using only the model's internal optimization signals and no external reward models or step annotations. To make this scalable, GRACE introduces a representation-level gradient proxy that estimates step-level alignment from token-level upstream signals in a single forward pass. Post-training Qwen3-VL-2B-Instruct on MMathCoT-1M, GRACE reaches 108.8% of the full-data performance with 20% of the data and retains 100.2% with only 5%, with subsets that transfer effectively across model backbones.

📄 PDF Abstract BibTeX arXiv:2605.13130

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment

2026-01-15 · Felix Jahn, Yannic Muskalla, Lisa Dargasz, Patrick Schramowski 외 arxiv

As AI agents become increasingly autonomous, widely deployed in consequential contexts, and efficacious in bringing about real-world impacts, ensuring that their decisions are not only instrumentally effective but also n…

What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge

2026-01-16 · Yosub Shin, Michael Buriek, Boris Sobolev, Pavel Bushuyeu 외 arxiv

We study data curation for multimodal reasoning through the NeurIPS 2025 Data Curation for Vision-Language Reasoning (DCVLR) challenge, which isolates dataset selection by fixing the model and training protocol. Using a …

Multimodal Reasoning

GRACE: Discriminator-Guided Chain-of-Thought Reasoning

2023-05-24 · Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee 외

In the context of multi-step reasoning, e.g., with chain-of-thought, language models (LMs) can easily assign a high likelihood to incorrect steps. As a result, decoding strategies that optimize for solution likelihood of…

GSM8KMath

GRACE: Generative Representation Learning via Contrastive Policy Optimization

2025-10-06 · Jiashuo Sun, Shixuan Liu, Zhaochen Su, Xianrui Zhong 외 arxiv

Prevailing methods for training Large Language Models (LLMs) as text encoders rely on contrastive losses that treat the model as a black box function, discarding its generative and reasoning capabilities in favor of stat…

Representation Learning

GRACE: Generative Recommendation via Journey-Aware Sparse Attention on Chain-of-Thought Tokenization

2025-07-19 · Luyi Ma, Wanjia Zhang, Kai Zhao, Abhishek Kulkarni 외 arxiv

Generative models have recently demonstrated strong potential in multi-behavior recommendation systems, leveraging the expressive power of transformers and tokenization to generate personalized item sequences. However, t…

Sequential RecommendationRecommendation SystemsKnowledge Graphs