paper-with-me

Papers

GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing

2026-03-12 · Mingxin Liu, Ziqian Fan, Zhaokai Wang, Leyao Gu, Zirun Zhu, Yiguo He, Yuchen Yang, Changyao Tian, Xiangyu Zhao, Ning Liao, Shaofeng Zhang, Qibing Ren, Zhihang Zhong, Xuanhe Zhou, Junchi Yan, Xue Yang arxiv

Unified multimodal models target joint understanding, reasoning, and generation, but current image editing benchmarks are largely confined to natural images and shallow commonsense reasoning, offering limited assessment of this capability under structured, domain-specific constraints. In this work, we introduce GRADE, the first benchmark to assess discipline-informed knowledge and reasoning in image editing. GRADE comprises 520 carefully curated samples across 10 academic domains, spanning from natural science to social science. To support rigorous evaluation, we propose a multi-dimensional evaluation protocol that jointly assesses Discipline Reasoning, Visual Consistency, and Logical Readability. Extensive experiments on 20 state-of-the-art open-source and closed-source models reveal substantial limitations in current models under implicit, knowledge-intensive editing settings, leading to large performance gaps. Beyond quantitative scores, we conduct rigorous analyses and ablations to expose model shortcomings and identify the constraints within disciplinary editing. Together, GRADE pinpoints key directions for the future development of unified multimodal models, advancing the research on discipline-informed image editing and reasoning. Our benchmark and evaluation code are publicly released.

📄 PDF Abstract BibTeX arXiv:2603.12264

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing

2026-07-02 · Zhaokai Wang, Mingxin Liu, Zirun Zhu, Ziqian Fan 외 arxiv

Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a knowledge-intensive diagram whose correctness depends on disciplinary concep…

Text-to-Image GenerationImage Editing

T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation

2025-08-24 · Kaiyue Sun, Rongyao Fang, Chengqi Duan, Xian Liu 외 arxiv

We propose T2I-ReasonBench, a benchmark evaluating reasoning capabilities of text-to-image (T2I) models. It consists of four dimensions: Idiom Interpretation, Textual Image Design, Entity-Reasoning and Scientific-Reasoni…

Text-to-Image Generation

OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

2024-06-18 · Zhen Huang, Zengzhi Wang, Shijie Xia, Xuefeng Li 외

The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abil…

Benchmarkingscientific discovery

KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning

2025-05-14 · Peiqi Sui, Juan Diego Rodriguez, Philippe Laban, Dean Murphy 외

Each year, tens of millions of essays are written and graded in college-level English courses. Students are asked to analyze literary and cultural texts through a process known as close reading, in which they gather text…

BenchmarkingMMLUMultiple-choice

MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models

2025-04-08 · Pengfei Zhou, Fanrui Zhang, Xiaopeng Peng, Zhaopan Xu 외

Multimodal reasoning, which integrates language and visual cues into problem solving and decision making, is a fundamental aspect of human intelligence and a crucial step toward artificial general intelligence. However, …

MathMultimodal Reasoning