paper-with-me

홈 › Papers

STAR: SpatioTemporal Adaptive Reward Allocation for Text-to-Image RL Post-Training

2026-06-16 · Jinjie Shen, Wei Deng, Xian Hu, Daiguo Zhou, Jian Luan arxiv

Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the entire generative trajectory. However, text-to-image generation naturally has temporal and spatial structure: different denoising steps are responsible for different generation stages, and the content that truly determines text alignment often appears only in part of the image. This granularity mismatch makes it difficult for policy updates to focus on the generative components that actually affect the reward. To address this issue, we propose \textbf{SpatioTemporal Adaptive Reward (STAR) Allocation} for RL post-training of text-to-image diffusion and flow models. STAR uses text-image attention inside the generative model and starts from the core content that the user truly cares about in the prompt. It constructs spatial allocation maps that dynamically vary across denoising steps and rollouts, and allocates the same group-relative advantage to more relevant latent regions with almost no additional computational overhead. STAR then applies stronger policy updates to these regions through a spatially resolved policy objective. We use Stable Diffusion 3.5 Medium as the base model and evaluate on three tasks: GenEval, OCR text rendering, and PickScore. Experimental results show that STAR improves compositional semantic alignment, text rendering, and preference optimization without changing the external reward source, achieving $\mathbf{0.9759}$, $\mathbf{0.9757}$, and $\mathbf{23.60}$ on GenEval, OCR, and PickScore, respectively.

📄 PDF Abstract BibTeX arXiv:2606.17979

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

Spatiotemporal Modelling of Multi-Gateway LoRa Networks with Imperfect SF Orthogonality

2020-08-27 · Yathreb Bouazizi, Fatma Benkhelifa, Julie McCann

Meticulous modelling and performance analysis of Low-Power Wide-Area (LPWA) networks are essential for large scale dense Internet-of-Things (IoT) deployments. As Long Range (LoRa) is currently one of the most prominent L…

Resource allocation method using tug-of-war-based synchronization

2021-08-19 · Song-Ju Kim, Hiroyuki Yasuda, Ryoma Kitagawa, Mikio Hasegawa

We propose a simple channel-allocation method based on tug-of-war (TOW) dynamics, combined with the time scheduling based on nonlinear oscillator synchronization to efficiently use of the space (channel) and time resourc…

Scheduling

Coupling User Preference with External Rewards to Enable Driver-centered and Resource-aware EV Charging Recommendation

2022-10-23 · Chengyin Li, Zheng Dong, Nathan Fisher, Dongxiao Zhu

Electric Vehicle (EV) charging recommendation that both accommodates user preference and adapts to the ever-changing external environment arises as a cost-effective strategy to alleviate the range anxiety of private EV d…

Hierarchical Reinforcement Learning in StarCraft Micromanagement with Influence Maps and Cluster-based Scripts

2026-06-29 · Chunhui Bai, Changhe Li, Dequan Li, Xinye Cai 외 arxiv

Real-time strategy (RTS) games present significant AI challenges, characterized by expansive state-action spaces arising from multi-unit coordination in continuous battlefields, and sparse delayed rewards stemming from f…

Hierarchical Reinforcement Learning

Adaptive Bi-Level Multi-Robot Task Allocation and Learning under Uncertainty with Temporal Logic Constraints

2025-02-14 · Xiaoshan Lin, Roberto Tron

This work addresses the problem of multi-robot coordination under unknown robot transition models, ensuring that tasks specified by Time Window Temporal Logic are satisfied with user-defined probability thresholds. We pr…