paper-with-me

홈 › Papers

Reasoning on a Budget: Miniaturizing DeepSeek R1 with SFT-GRPO Alignment for Instruction-Tuned LLMs

2025-05-16 · techrxiv 2025 5 · Esmaeil Narimissa

Large language models (LLMs) excel at general-purpose generation but often struggle with structured reasoning tasks. Recent methods like DeepSeek-R1 have shown that reinforcement learning with rule-based rewards can significantly enhance reasoning capabilities. However, reproducing such pipelines remains computationally intensive and inaccessible to most researchers. In this work, we present a modular, low-cost replication of the DeepSeek-R1 training methodology using Qwen2.5-0.5B-Instruct (a compact instruction-tuned LLM) optimized via a two-stage pipeline: Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). SFT aligns the base model to reasoning-structured prompts using LoRA-based parameter-efficient fine-tuning. GRPO then refines this policy using a critic-free reinforcement learning algorithm guided by five composable reward functions, including accuracy, reasoning presence, and formatting compliance. The entire training process was executed for under $100 USD on AWS SageMaker, demonstrating that high-impact reasoning alignment is achievable without large-scale compute. Quantitative metrics confirm strong convergence, high reward stability, and consistent output structure. This study contributes a scalable and reproducible template for aligning compact LLMs to reasoning-intensive tasks under constrained computational budgets.

📄 PDF Abstract BibTeX

Code (1)

EsmaeilNarimissa/aws-sft-grpo-budget-llm-finetune pytorch

Tasks

Deep Reinforcement LearningMathematical Reasoningparameter-efficient fine-tuningreinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

BASE 설명 없음
SFT Shrink and Fine-Tune, or SFT, is a type of distillation that avoids explicit distillation by copying parameters to a student student model and then fine-tuning.…

Similar Papers 제목 키워드 기반

iGRPO: Self-Feedback-Driven LLM Reasoning

2026-02-09 · Ali Hatamizadeh, Shrimai Prabhumoye, Igor Gitman, Ximing Lu 외 arxiv

Large Language Models (LLMs) have shown promise in solving complex mathematical problems, yet they still fall short of producing accurate and consistent solutions. Reinforcement Learning (RL) is a framework for aligning …

Reinforcement LearningMathematical Reasoning

Spectral Policy Optimization: Coloring your Incorrect Reasoning in GRPO

2025-05-16 · Peter Chen, Xiaopeng Li, Ziniu Li, Xi Chen 외

Reinforcement learning (RL) has demonstrated significant success in enhancing reasoning capabilities in large language models (LLMs). One of the most widely used RL methods is Group Relative Policy Optimization (GRPO)~\c…

AllDiversityReinforcement Learning (RL)

Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning

2026-03-29 · Feiding, Yongkang Zhang, Yuhao Liao, Zijian Zeng 외 arxiv

Vision--language models (VLMs) are increasingly aligned via Group Relative Policy Optimization (GRPO)-style training. However, relying solely on terminal outcome rewards yields sparse credit assignment in multi-step reas…

Reinforcement LearningMultimodal Reasoning

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models

2025-06-05 · Fei Ding, Baiqiao Wang, Zijian Zeng, Youwei Wang

The Group Relative Policy Optimization (GRPO) algorithm has demonstrated considerable success in enhancing the reasoning capabilities of large language models (LLMs), as evidenced by DeepSeek-R1. However, the absence of …

Mathematical Reasoning

Broken Chains: The Cost of Incomplete Reasoning in LLMs

2026-02-16 · Ian Su, Gaurav Purushothaman, Jey Narayan, Ruhika Goel 외 arxiv

Reasoning-specialized models like OpenAI's 5.1 and DeepSeek-V3.2 allocate substantial inference compute to extended chain-of-thought (CoT) traces, yet reasoning tokens incur significant costs. How do different reasoning …