paper-with-me

Papers

Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts

2026-06-09 · Haoyu Dong arxiv

Code-generating large language models (LLMs) increasingly produce visual artifacts such as charts, web pages, and slides by writing programs that are executed by non-differentiable renderers, committing to code before observing the render. As a result, otherwise executable code often yields artifacts with visually salient defects, including overlapping elements, clipped text, broken alignment, low contrast, and overflow. We study visual-feedback self-distillation for code-generated visual artifacts. We propose Visual-SDPO, a self-distillation policy-optimization framework that treats rendered visual feedback as privileged context for a weight-sharing teacher and distills this feedback into a coding student. To make supervision spatially targeted rather than uniform, we introduce Visual-Grounded Code Credit Weighting, which traces each detected defect back to the code statements responsible for the affected elements and amplifies the distillation signal on those statements. A sequence-level GRPO (Group Relative Policy Optimization) term complements the dense token-level objective by rewarding executable, visually high-quality rollouts, while failed executions remain learnable through the self-distillation path by passing execution errors as privileged context to the teacher. We instantiate Visual-SDPO for chart, web/UI, and slide generation with a unified Qwen3-VL-8B-Instruct backbone. Across chart-to-code, UI-to-code, and slide-generation benchmarks (ChartMimic, Design2Code, and AeSlides), Visual-SDPO improves over the zero-shot base by more than 10 absolute points in the primary metric and over GRPO by at least 2.4 points, with fewer training steps and no added inference-time cost.

📄 PDF Abstract BibTeX arXiv:2606.10334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Policy Improvement Reinforcement Learning

2026-04-01 · Huaiyang Wang, Xiaojie Li, Xiaohan Wang, Zhixia Zhang 외 arxiv

Reinforcement learning has become a central post-training paradigm for improving LLM and agent capabilities. Yet existing RL post-training methods share a common blind spot: they construct local learning signals from sam…

Mathematical ReasoningReinforcement Learning

Learning from Language Feedback via Variational Policy Distillation

2026-05-14 · Yang Li, Erik Nijkamp, Semih Yavuz, Shafiq Joty arxiv

Reinforcement learning from verifiable rewards (RLVR) suffers from sparse outcome signals, creating severe exploration bottlenecks on complex reasoning tasks. Recent on-policy self-distillation methods attempt to address…

Reinforcement LearningMathematical ReasoningCode Generation

Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback

2026-06-12 · Woohyeon Byeon, Jiwon Jeon, Jeonghye Kim, Youngchul Sung arxiv

We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillation or single-model fine-tuning, our goal …

Reinforcement Learning via Self-Distillation

2026-01-28 · Jonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann 외 arxiv

Large language models are increasingly post-trained with reinforcement learning in verifiable domains such as code and math. Yet, current methods for reinforcement learning with verifiable rewards (RLVR) learn only from …

Reinforcement Learning

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity

2026-06-24 · Andrei Liviu Nicolicioiu, Mohammad Pezeshki, Aaron Courville arxiv

On-policy self-distillation achieves strong pass@1 accuracy by using a single model as both teacher and student, with the teacher conditioned on a correct demonstration to provide dense token-level feedback. We show that…

Reinforcement Learning