paper-with-me

홈 › Papers

Discrete Flow Matching Policy Optimization

2026-04-07 · Maojiang Su, Po-Chung Hsieh, Weimin Wu, Mingcheng Lu, Jiunhau Chen, Jerry Yao-Chieh Hu, Han Liu arxiv

We introduce Discrete flow Matching policy Optimization (DoMinO), a unified framework for Reinforcement Learning (RL) fine-tuning Discrete Flow Matching (DFM) models under a broad class of policy gradient methods. Our key idea is to view the DFM sampling procedure as a multi-step Markov Decision Process. This perspective provides a simple and transparent reformulation of fine-tuning reward maximization as a robust RL objective. Consequently, it not only preserves the original DFM samplers but also avoids biased auxiliary estimators and likelihood surrogates used by many prior RL fine-tuning methods. To prevent policy collapse, we also introduce new total-variation regularizers to keep the fine-tuned distribution close to the pretrained one. Theoretically, we establish an upper bound on the discretization error of DoMinO and tractable upper bounds for the regularizers. Experimentally, we evaluate DoMinO on regulatory DNA sequence design. DoMinO achieves stronger predicted enhancer activity and better sequence naturalness than the previous best reward-driven baselines. The regularization further improves alignment with the natural sequence distribution while preserving strong functional performance. These results establish DoMinO as an useful framework for controllable discrete sequence generation.

📄 PDF Abstract BibTeX arXiv:2604.06491

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation

2026-02-12 · Yeyao Ma, Chen Li, Xiaosong Zhang, Han Hu 외 arxiv

Post-training of flow matching models-aligning the output distribution with a high-quality target-is mathematically equivalent to imitation learning. While Supervised Fine-Tuning mimics expert demonstrations effectively,…

Video GenerationImage Generation

Tessellations of Semi-Discrete Flow Matching

2026-05-08 · Emile Pierret, Johannes Hertrich, Samuel Hurault, Julie Delon arxiv

We study Flow Matching in a semi-discrete setting where a Gaussian source is transported toward a discrete target supported on finitely many points. This semi-discrete regime is the theoretical setting behind the use of …

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

2025-10-24 · Yifu Luo, Haoyuan Sun, Xinhao Hu, Penghui Du 외 arxiv

Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccu…

Reinforcement Learning

Discrete Flow Matching for Offline-to-Online Reinforcement Learning

2026-05-12 · Fairoz Nower Khan, Nabuat Zaman Nahim, Peizhong Ju arxiv

Many reinforcement learning (RL) tasks have discrete action spaces, but most generative policy methods based on diffusion and flow matching are designed for continuous control. Meanwhile, generative policies usually rely…

Reinforcement LearningContinuous Control

Proximal Policy Optimization for Amortized Discrete Sampling

2026-06-14 · Anna Zykova-Myzina, Timofei Gritsaev, Daniil Tiapkin, Nikita Morozov arxiv

This paper explores policy gradient algorithms for training stochastic policies to sample from structured discrete probability distributions under the Generative Flow Network (GFlowNet) framework. Building on extensive t…

Reinforcement LearningGraph Generation