paper-with-me

Papers

AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization

2026-03-18 · Dailan He, Guanlin Feng, Xingtong Ge, Yi Zhang, Bingqi Ma, Guanglu Song, Yu Liu, Hongsheng Li arxiv

Streaming autoregressive (AR) video generators combined with few-step distillation achieve low-latency, high-quality synthesis, yet remain difficult to align via reinforcement learning from human feedback (RLHF). Existing SDE-based GRPO methods face challenges in this setting: few-step ODEs and consistency model samplers deviate from standard flow-matching ODEs, and their short, low-stochasticity trajectories are highly sensitive to initialization noise, rendering intermediate SDE exploration ineffective. We propose AR-CoPO (AutoRegressive Contrastive Policy Optimization), a framework that adapts the Neighbor GRPO contrastive perspective to streaming AR generation. AR-CoPO introduces chunk-level alignment via a forking mechanism that constructs neighborhood candidates at a randomly selected chunk, assigns sequence-level rewards, and performs localized GRPO updates. We further propose a semi-on-policy training strategy that complements on-policy exploration with exploitation over a replay buffer of reference rollouts, improving generation quality across domains. Experiments on Self-Forcing demonstrate that AR-CoPO improves both out-of-domain generalization and in-domain human preference alignment over the baseline, providing evidence of genuine alignment rather than reward hacking.

📄 PDF Abstract BibTeX arXiv:2603.17461

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningDomain GeneralizationVideo Generation

Similar Papers 제목 키워드 기반

The Past Mistake is the Future Wisdom: Error-driven Contrastive Probability Optimization for Chinese Spell Checking

2021-11-16 · ACL ARR September 2021 9 · Anonymous

Chinese Spell Checking (CSC) aims to detect and correct Chinese spelling errors, which are mainly caused by the phonological or visual similarity. Recently, pre-trained language models (PLMs) promote the progress of CSC …

Chinese Spell Checking

The Past Mistake is the Future Wisdom: Error-driven Contrastive Probability Optimization for Chinese Spell Checking

2022-03-02 · Findings (ACL) 2022 5 · Yinghui Li, Qingyu Zhou, Yangning Li, Zhongli Li 외

Chinese Spell Checking (CSC) aims to detect and correct Chinese spelling errors, which are mainly caused by the phonological or visual similarity. Recently, pre-trained language models (PLMs) promote the progress of CSC …

Chinese Spell Checking

VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos

2024-09-11 · Yan-Bo Lin, Yu Tian, Linjie Yang, Gedas Bertasius 외

We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large…

Contrastive LearningMusic Generation

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration

2026-05-14 · Ruicheng Zhang, Kaixi Cong, Jun Zhou, Zhizhou Zhong 외 arxiv

Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rely on noise-based exploration and SDE-based surrogate policies that a…

Reinforcement LearningVideo Alignment

OmniDRCA: Parallel Speech-Text Foundation Model via Dual-Resolution Speech Representations and Contrastive Alignment

2025-06-11 · Chao-Hong Tan, Qian Chen, Wen Wang, Chong Deng 외

Recent studies on end-to-end speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs to generate discrete speech tokens. Existing …

cross-modal alignmentQuestion AnsweringSpeech SynthesisText Generation