paper-with-me

홈 › Papers

SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning

2026-02-07 · Yancheng Long, Yankai Yang, Hongyang Wei, Wei Chen, Tianke Zhang, Haonan fan, Changyi Liu, Kaiyu Jiang, Jiankang Chen, Kaiyu Tang, Bin Wen, Fan Yang, Tingting Gao, Han Li, Shuo Yang arxiv

Online Reinforcement Learning (RL) offers a promising avenue for complex image editing but is currently constrained by the scarcity of reliable and fine-grained reward signals. Existing evaluators frequently struggle with a critical perception gap we term "Attention Collapse," where models neglect cross-image comparisons and fail to capture fine-grained details, resulting in inaccurate perception and miscalibrated scores. To address these limitations, we propose SpatialReward, a reward model that enforces precise verification via explicit spatial reasoning. By anchoring reasoning to predicted edit regions, SpatialReward grounds semantic judgments in pixel-level evidence, significantly enhancing evaluative accuracy. Trained on a curated 260k spatial-aware dataset, our model achieves state-of-the-art performance on MMRB2 and EditReward-Bench, and outperforms proprietary evaluators on our proposed MultiEditReward-Bench. Furthermore, SpatialReward serves as a robust signal in online RL, boosting OmniGen2 by +0.90 on GEdit-Bench--surpassing the leading discriminative model and doubling the gain of GPT-4.1 (+0.45). These results demonstrate that spatial reasoning is essential for unlocking effective alignment in image editing.

📄 PDF Abstract BibTeX arXiv:2602.07458

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningSpatial ReasoningImage Editing

Similar Papers 제목 키워드 기반

SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation

2026-03-23 · Sashuai Zhou, Qiang Zhou, Junpeng Ma, Yue Cao 외 arxiv

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited atte…

Text-to-Image GenerationReinforcement LearningVisual Grounding

Has the Virtualization of the Face Changed Facial Perception? A Study of the Impact of Photo Editing and Augmented Reality on Facial Perception

2023-03-01 · Louisa Conwill, Sam English Anthony, Walter J. Scheirer

Augmented reality and other photo editing filters are popular methods used to modify faces online. Considering the important role of facial perception in communication, how do we perceive this increasing number of modifi…

AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations

2026-08-14 · Ying Huang, Wencan Zhang, Brian Y. Lim arxiv

Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human perception. While percep…

Enhancing Spatial Understanding in Image Generation via Reward Modeling

2026-02-27 · Zhenyu Tang, Chaoran Feng, Yufan Deng, Jie Wu 외 arxiv

Recent progress in text-to-image generation has greatly advanced visual fidelity and creativity, but it has also imposed higher demands on prompt complexity-particularly in encoding intricate spatial relationships. In su…

Text-to-Image GenerationReinforcement Learning

IE-Critic-R1: Advancing the Explanatory Measurement of Text-Driven Image Editing for Human Perception Alignment

2025-11-22 · Bowen Qu, Shangkun Sun, Xiaoyu Liang, Wei Gao arxiv

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different from the assessment of text-driven image…

Reinforcement LearningImage GenerationImage Editing