paper-with-me

홈 › Papers

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

2025-05-22 · Yongliang Wu, Zonghui Li, Xinting Hu, Xinyu Ye, Xianfang Zeng, Gang Yu, Wenbo Zhu, Bernt Schiele, Ming-Hsuan Yang, Xu Yang

Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, we introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens. Drawing from educational theory, KRIS-Bench categorizes editing tasks across three foundational knowledge types: Factual, Conceptual, and Procedural. Based on this taxonomy, we design 22 representative tasks spanning 7 reasoning dimensions and release 1,267 high-quality annotated editing instances. To support fine-grained evaluation, we propose a comprehensive protocol that incorporates a novel Knowledge Plausibility metric, enhanced by knowledge hints and calibrated through human studies. Empirical results on 10 state-of-the-art models reveal significant gaps in reasoning performance, highlighting the need for knowledge-centric benchmarks to advance the development of intelligent image editing systems.

📄 PDF Abstract BibTeX arXiv:2505.16707

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDiagnostic

Similar Papers 제목 키워드 기반

KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning

2025-05-14 · Peiqi Sui, Juan Diego Rodriguez, Philippe Laban, Dean Murphy 외

Each year, tens of millions of essays are written and graded in college-level English courses. Students are asked to analyze literary and cultural texts through a process known as close reading, in which they gather text…

BenchmarkingMMLUMultiple-choice

KrishokBondhu: A Retrieval-Augmented Voice-Based Agricultural Advisory Call Center for Bengali Farmers

2025-10-21 · Mohd Ruhul Ameen, Akif Islam, Farjana Aktar, M. Saifuzzaman Rafat arxiv

In Bangladesh, many farmers still struggle to access timely, expert-level agricultural guidance. This paper presents KrishokBondhu, a voice-enabled, call-centre-integrated advisory platform built on a Retrieval-Augmented…

Semantic Retrieval

VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context

2025-11-11 · Heyang Liu, Ziyang Cheng, Yuhao Wang, Hongcheng Liu 외 arxiv

The development of multi-modal large language models (LLMs) leads to intelligent approaches capable of speech interactions. As one of the most widely spoken languages globally, Mandarin is supported by most models to enh…

Pinchah Kristang: A Dictionary of Kristang

2020-05-01 · LREC 2020 5 · Lu{\'\i}s Morgado da Costa

This paper describes the development and current state of Pinchah Kristang {--} an online dictionary for Kristang. Kristang is a critically endangered language of the Portuguese-Eurasian communities residing mainly in Ma…

Dance2Hesitate: A Multi-Modal Dataset of Dancer-Taught Hesitancy for Understandable Robot Motion

2026-03-10 · Srikrishna Bangalore Raghu, Anna Soukhovei, Divya Sai Sindhuja Vankineni, Alexandra Bacula 외 arxiv

In human-robot collaboration, a robot's expression of hesitancy is a critical factor that shapes human coordination strategies, attention allocation, and safety-related judgments. However, designing hesitant robot motion…