paper-with-me

Papers

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

2026-05-13 · Xuehai Bai, Yang Shi, Yi-Fan Zhang, Xuanyu Zhu, Yuran Wang, Yifan Dai, Xinyu Liu, Yiyan Ji, Xiaoling Gu, Yuanxing Zhang arxiv

Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, especially for strong frontier models, due to limited task difficulty and coarse-grained evaluation protocols. In parallel, reward models have become increasingly important for RL-based image editing optimization, yet existing reward model benchmarks still rely on unrealistic evaluation settings that deviate from practical RL scenarios. These limitations hinder reliable assessment of both image editing models and reward models. To address these challenges, we introduce Edit-Compass and EditReward-Compass, a unified evaluation suite for image editing and reward modeling. Edit-Compass contains 2,388 carefully annotated instances spanning six progressively challenging task categories, covering capabilities such as world knowledge reasoning, visual reasoning, and multi-image editing. Beyond broad task coverage, Edit-Compass adopts a fine-grained multidimensional evaluation framework based on structured reasoning and carefully designed scoring rubrics. In parallel, EditReward-Compass contains 2,251 preference pairs that simulate realistic reward modeling scenarios during RL optimization.

📄 PDF Abstract BibTeX arXiv:2605.13062

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingVisual ReasoningImage Editing

Similar Papers 제목 키워드 기반

EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing

2025-09-30 · Keming Wu, Sicong Jiang, Max Ku, Ping Nie 외 arxiv

Recently, we have witnessed great progress in image editing with natural language instructions. Several closed-source models like GPT-Image-1, Seedream, and Google-Nano-Banana have shown highly promising progress. Howeve…

Reinforcement LearningImage Editing

WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models

2026-04-20 · Xinping Lei, Xinyu Che, Junqi Xiong, Chenchen Zhang 외 arxiv

Large language models are rapidly evolving into interactive coding agents capable of end-to-end web coding, yet existing benchmarks evaluate only narrow slices of this capability, typically text-conditioned generation wi…

Ovis-U1 Technical Report

2025-06-29 · Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang, Liangfu Cao 외

In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Building on the foundation of the Ovis series,…

Image GenerationText to Image GenerationText-to-Image Generation

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

2026-07-15 · Kai Chen, Zichen Ding, Jiaye Ge, Shufan Jiang 외 arxiv

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hinderin…

SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning

2026-02-07 · Yancheng Long, Yankai Yang, Hongyang Wei, Wei Chen 외 arxiv

Online Reinforcement Learning (RL) offers a promising avenue for complex image editing but is currently constrained by the scarcity of reliable and fine-grained reward signals. Existing evaluators frequently struggle wit…

Reinforcement LearningSpatial ReasoningImage Editing