paper-with-me

Papers

TECCI: Tricky Edits of Collected and Curated Images

2026-05-31 · Aishwarya Agrawal, Roy Hirsch, Yasumasa Onoe, Sherry Ben, Jason Baldridge arxiv

Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruction following, minimally editing the source image, and ensuring high visual quality. These problems are especially apparent when the requested edit is challenging, such as those that involve position, motion, viewpoint, scale and creative edits. To systematically test generative image editors, we propose a novel image editing benchmark -- TECCI: Tricky Edits of Collected and Curated Images. TECCI consists of a completely new set of images we are releasing. The images in TECCI span 7 image categories. The images and these categories were curated intentionally to target weaknesses of existing methods. The edit instructions in TECCI are automatically generated by Gemini, covering 5 edit types per source image. We also curated a set of 530 images for which we created challenging manually written edit instructions. Overall, TECCI contains 7550 pairs of images and edit instructions. We conduct human evaluations of five leading image editing models on TECCI. Humans judge outputs along three dimensions: 1) instruction following, 2) minimality of the edits, and 3) visual quality. To scale-up the evaluation, we also build an auto-rater using Gemini that achieves 74.7% accuracy in matching human evaluations. Our evaluations reveal that: 1) none of the models exceed a 22% overall success rate, demonstrating the challenging nature of TECCI, 2) Nano Banana Pro is the best performing model overall, 3) models perform significantly better at instruction following compared to minimal edits and visual quality, 4) models struggle with editing architecture and nature images which require strong understanding of spatial layout and intricate visual details. 5) reasoning and creative edits are the most difficult, whereas color and appearance edits are the easiest.

📄 PDF Abstract BibTeX arXiv:2606.01213

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingImage Editing

Similar Papers 제목 키워드 기반

Pattern Analogies: Learning to Perform Programmatic Image Edits by Analogy

2024-12-17 · CVPR 2025 1 · Aditya Ganeshan, Thibault Groueix, Paul Guerrero, Radomír Měch 외

Pattern images are everywhere in the digital and physical worlds, and tools to edit them are valuable. But editing pattern images is tricky: desired edits are often programmatic: structure-aware edits that alter the unde…

AmsterTime: A Visual Place Recognition Benchmark Dataset for Severe Domain Shift

2022-03-30 · Burak Yildiz, Seyran Khademi, Ronald Maria Siebes, Jan van Gemert

We introduce AmsterTime: a challenging dataset to benchmark visual place recognition (VPR) in presence of a severe domain shift. AmsterTime offers a collection of 2,500 well-curated images matching the same scene from a …

Image ClassificationImage RetrievalMetric LearningRetrieval+1

Fix-A-Step: Semi-supervised Learning from Uncurated Unlabeled Data

2022-08-25 · Zhe Huang, Mary-Joy Sidhom, Benjamin S. Wessler, Michael C. Hughes

Semi-supervised learning (SSL) promises improved accuracy compared to training classifiers on small labeled datasets by also training on many unlabeled images. In real applications like medical imaging, unlabeled data wi…

Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

2024-07-03 · Benno Krojer, Dheeraj Vattikonda, Luis Lara, Varun Jampani 외

An image editing model should be able to perform diverse edits, ranging from object replacement, changing attributes or style, to performing actions or movement, which require many forms of reasoning. Current general ins…

AttributeSpatial Reasoning

WikiAtomicEdits: A Multilingual Corpus of Wikipedia Edits for Modeling Language and Discourse

2018-08-28 · EMNLP 2018 10 · Manaal Faruqui, Ellie Pavlick, Ian Tenney, Dipanjan Das

We release a corpus of 43 million atomic edits across 8 languages. These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or delet…

Representation LearningSentence