paper-with-me

Papers

ACECODER: Acing Coder RL via Automated Test-Case Synthesis

2025-02-03 · Huaye Zeng, Dongfu Jiang, Haozhe Wang, Ping Nie, Xiaotong Chen, Wenhu Chen

Most progress in recent coder models has been driven by supervised fine-tuning (SFT), while the potential of reinforcement learning (RL) remains largely unexplored, primarily due to the lack of reliable reward data/model in the code domain. In this paper, we address this challenge by leveraging automated large-scale test-case synthesis to enhance code model training. Specifically, we design a pipeline that generates extensive (question, test-cases) pairs from existing code data. Using these test cases, we construct preference pairs based on pass rates over sampled programs to train reward models with Bradley-Terry loss. It shows an average of 10-point improvement for Llama-3.1-8B-Ins and 5-point improvement for Qwen2.5-Coder-7B-Ins through best-of-32 sampling, making the 7B model on par with 236B DeepSeek-V2.5. Furthermore, we conduct reinforcement learning with both reward models and test-case pass rewards, leading to consistent improvements across HumanEval, MBPP, BigCodeBench, and LiveCodeBench (V4). Notably, we follow the R1-style training to start from Qwen2.5-Coder-base directly and show that our RL training can improve model on HumanEval-plus by over 25\% and MBPP-plus by 6\% for merely 80 optimization steps. We believe our results highlight the huge potential of reinforcement learning in coder models.

📄 PDF Abstract BibTeX arXiv:2502.01718

Code (0)

등록된 구현이 없습니다.

Tasks

HumanEvalmbppreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

AceCoder: Utilizing Existing Code to Enhance Code Generation

2023-03-31 · Jia Li, YunFei Zhao, Yongmin Li, Ge Li 외

Large Language Models (LLMs) have shown great success in code generation. LLMs take as the input a prompt and output the code. A key question is how to make prompts (i.e., Prompting Techniques). Existing prompting techni…

Code GenerationmbppRetrievalText Generation

TraceCoder: Towards Traceable ICD Coding via Multi-Source Knowledge Integration

2025-10-17 · Mucheng Ren, He Chen, Yuchen Yan, Danqing Hu 외 arxiv

Automated International Classification of Diseases (ICD) coding assigns standardized diagnosis and procedure codes to clinical records, playing a critical role in healthcare systems. However, existing methods face challe…

TraceCoder: A Trace-Driven Multi-Agent Framework for Automated Debugging of LLM-Generated Code

2026-02-06 · Jiangping Huang, Wenguang Ye, Weisong Sun, Jian Zhang 외 arxiv

Large Language Models (LLMs) often generate code with subtle but critical bugs, especially for complex tasks. Existing automated repair methods typically rely on superficial pass/fail signals, offering limited visibility…

Automated Testing for Deep Learning Systems with Differential Behavior Criteria

2019-12-31 · Yuan Gao, Yiqiang Han

In this work, we conducted a study on building an automated testing system for deep learning systems based on differential behavior criteria. The automated testing goals were achieved by jointly optimizing two objective …

Deep Learning

Graph Learning for Bidirectional Disease Contact Tracing on Real Human Mobility Data

2025-01-30 · Sofia Hurtado, Radu Marculescu

For rapidly spreading diseases where many cases show no symptoms, swift and effective contact tracing is essential. While exposure notification applications provide alerts on potential exposures, a fully automated system…

Graph Learning