paper-with-me

홈 › Papers

RULER: Instance-aware Rubric Rewards for SVG Generation

2026-09-21 · Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng hf

Generating Scalable Vector Graphics (SVG) code from natural-language instructions is an open-ended task without absolute visual ground truth, leaving both evaluation and policy optimization without a faithful signal. Scalar metrics (CLIP, Aesthetic) calibrated on natural images transfer poorly to stylized vector content, and reusing them as RL rewards triggers reward hacking. We address both limitations with rubric-based scoring. We first establish empirically that prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics, both across samples and within instructions. Building on this finding, we introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning), which converts each instruction into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and the weighted satisfactions form a fine-grained reward optimized via Group Relative Policy Optimization. Because the rubric is derived from text alone, RULER requires neither paired SVG ground truth nor human preference labels. On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432/0.395 to 0.693/0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3, with ablations identifying rubric design as the active lever for RL on open-ended SVG generation. The project page is available at https://hangyuran.github.io/RULER/.

📄 PDF Abstract BibTeX arXiv:2609.25270

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

2026-01-13 · Yihan Hong, Huaiyuan Yao, Bolin Shen, Wanpeng Xu 외 arxiv

Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the same criteria inconsistently, produce score attributions that are difficu…

Text Generation

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

2026-07-06 · Mingqi Gao, Hongyuan Dong, Yifei Chen, Zhisheng Zhong 외 arxiv

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs …

Dense Video CaptioningReinforcement Learning

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning

2026-05-22 · Xiaoyuan Li, Keqin Bao, Moxin Li, Yubo Ma 외 arxiv

Rubric-based rewards offer a promising way to extend reinforcement learning (RL) for large language models beyond tasks with automatically verifiable answers. However, scaling rubric-based RL remains challenging: existin…

Reinforcement LearningContinual PretrainingInstruction Following

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

2026-05-19 · Utkarsh Tyagi, Xingang Guo, MohammadHossein Rezaei, Daniel George 외 arxiv

Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many important model behaviors require satisfying several qualitative criteri…

Reinforcement Learning

Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards

2026-05-26 · Yu Huang, Zihua Zhao, Zhaoxin Huan, Wanli Gu 외 arxiv

The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical dilemma inherent in this training paradig…

Reinforcement Learning