paper-with-me

홈 › Papers

OPOD: On-Policy Omni Distillation

2026-07-23 · Tong Zhao, Yuyang Hu, Reed Li, Yu Lu, Haibo Shi, Yutao Zhu, Zhicheng Dou arxiv

Omni-modal models provide a unified interface for text, images, and audio. However, improving these abilities together remains difficult, as post-training on pooled multimodal data often fails to preserve the strengths of modality teachers. On-policy distillation (OPD) has recently become popular in model post-training. It samples responses from the current student and compares the teacher's and student's next-token distributions along those responses, yielding dense supervision while reducing the mismatch between training and inference. Despite these advantages, standard OPD does not readily extend to several modality teachers. Their guidance may favor conflicting changes to the shared model, while matching each teacher's next-token distribution can prevent the student from moving beyond that teacher. To address these challenges, we propose On-Policy Omni Distillation (OPOD), which consolidates text, image, and audio teachers into one omni model. OPOD routes each response to the corresponding teacher, controls the teachers independently, and applies guidance only when the teacher assigns a higher probability to the generated token. The selected teacher also evaluates answer confidence and whether the reasoning increases support for the answer. Extensive experiments on twelve benchmarks show that OPOD achieves the best average at three model scales, reaching 70.8, 51.7, and 46.2 and outperforming the strongest comparator by 2.1, 1.8, and 1.7 points. At 30B, it surpasses the base model and pooled RL training on all twelve benchmarks, and ranks first or second on eleven even when the teachers are included. Only the student is retained for deployment.

📄 PDF Abstract BibTeX arXiv:2607.20918

Code (1)

Tavish9/awesome-daily-AI-arxiv ★ 112

Similar Papers 제목 키워드 기반

Enhancing Vision-Based Policies with Omni-View and Cross-Modality Knowledge Distillation for Mobile Robots

2026-03-21 · Kai Li, Shiyu Zhao arxiv

Vision-based policies are widely applied in robotics for tasks such as manipulation and locomotion. On lightweight mobile robots, however, they face a trilemma of limited scene transferability, restricted onboard computa…

Knowledge Distillation

Continual Reinforcement Learning deployed in Real-life using Policy Distillation and Sim2Real Transfer

2019-06-11 · René Traoré, Hugo Caselles-Dupré, Timothée Lesort, Te Sun 외

We focus on the problem of teaching a robot to solve tasks presented sequentially, i.e., in a continual learning scenario. The robot should be able to solve all tasks it has encountered, without forgetting past tasks. We…

Continual Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification

2026-05-31 · Yuhang Zhou, Lizhu Zhang, Yifan Wu, Mingyi Wang 외 arxiv

On-Policy Distillation (OPD) trains a student model on its own generative trajectories under dense token-level feedback from a stronger teacher, mitigating both the off-policy distribution shift of Supervised Fine-Tuning…

Reinforcement LearningSemantic Similarity

DisCoRL: Continual Reinforcement Learning via Policy Distillation

2019-07-11 · René Traoré, Hugo Caselles-Dupré, Timothée Lesort, Te Sun 외

In multi-task reinforcement learning there are two main challenges: at training time, the ability to learn different policies with a single model; at test time, inferring which of those policies applying without an exter…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Representation Learning

OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing

2026-06-14 · Zebang Cheng, Shuimu Chen, Boxue Yang, Yuanshen Guan 외 arxiv

Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This challenge is particularly pronounced in human-centered scenarios involving …

Reinforcement Learning