paper-with-me

홈 › Papers

Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs

2026-05-07 · Yujia Chen, Yang Ye, Xiao Chu, Yuchi Ma, Cuiyun Gao arxiv

Reinforcement learning (RL) with verifiable rewards has proven effective at post-training LLMs for coding, yet deploying separate task-specific specialists incurs costs that scale with the number of tasks, motivating a unified multi-task RL (MTRL) approach. However, existing MTRL methods treat all coding tasks uniformly, relying on fixed data curricula under a shared optimization strategy, ultimately limiting the effectiveness of multi-task training. To address these limitations, we propose ASTOR, a multi-tASk code reinforcement learning framework via uTility-driven coORdination. Centered on task utility, a signal capturing each task learning potential and cross-task synergy, ASTOR comprises two coupled modules: 1) Hierarchical Utility-Routed Data Scheduling module hierarchically allocates training budget and prioritizes informative prompts, steering training toward the most valuable data and 2) Adaptive Utility-Calibrated Policy Optimization module dynamically scales per-task KL regularization, matching update constraints to each tasks current training state. Experiments on two widely-used LLMs across four representative coding tasks demonstrate that ASTOR consistently improves a single model across all tasks, outperforming the best task-specific specialist by 9.0%-9.5% and surpassing the strongest MTRL baseline by 7.5%-12.8%.

📄 PDF Abstract BibTeX arXiv:2605.06111

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

VORTEX: Aligning Task Utility and Human Preferences through LLM-Guided Reward Shaping

2025-09-19 · Guojun Xiong, Milind Tambe arxiv

In social impact optimization, AI decision systems often rely on solvers that optimize well-calibrated mathematical objectives. However, these solvers cannot directly accommodate evolving human preferences, typically exp…

A Generative AI Framework for Intelligent Utility Billing CO 2 Analytics and Sustainable Resource Optimisation

2026-05-15 · Pavan Manjunath, Thomas Pruefer arxiv

Distribution utilities are now expected to deliver bills that customers can actually read attach a defensible carbon number to every kWh sold and schedule load against grid stress and emissions constraints We propose an …

Linking Perception, Confidence and Accuracy in MLLMs

2026-03-12 · Yuetian Du, Yucheng Wang, Rongyu Zhang, Zhijie Xu 외 arxiv

Recent advances in Multi-modal Large Language Models (MLLMs) have predominantly focused on enhancing visual perception to improve accuracy. However, a critical question remains unexplored: Do models know when they do not…

Reinforcement Learning

Difficulty-Calibrated Interpolation Paths for Conditional Flow Matching

2026-08-21 · Airin Akter Tania, Md Raihan Khan arxiv

Conditional Flow Matching trains generative models by regressing a network onto the velocity of a prescribed noise-to-data interpolation path. The interpolation schedule that shapes this path is known to affect convergen…

Class-frequency Guided Noise Schedule for Diffusion Models

2026-06-26 · Jiequan Cui, Beier Zhu, Qingshan Xu, Xiaojuan Qi 외 arxiv

In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion models. For score-based generative models, low-density regions often lead to inaccur…

Text-to-Image GenerationImage Classification