paper-with-me

홈 › Papers

I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models

2025-12-04 · Juntong Wang, Jiarui Wang, Huiyu Duan, Jiaxiang Kang, Guangtao Zhai, Xiongkuo Min arxiv

Image editing models are advancing rapidly, yet comprehensive evaluation remains a significant challenge. Existing image editing benchmarks generally suffer from limited task scopes, insufficient evaluation dimensions, and heavy reliance on manual annotations, which significantly constrain their scalability and practical applicability. To address this, we propose \textbf{I2I-Bench}, a comprehensive benchmark for image-to-image editing models, which features (i) diverse tasks, encompassing 10 task categories across both single-image and multi-image editing tasks, (ii) comprehensive evaluation dimensions, including 30 decoupled and fine-grained evaluation dimensions with automated hybrid evaluation methods that integrate specialized tools and large multimodal models (LMMs), and (iii) rigorous alignment validation, justifying the consistency between our benchmark evaluations and human preferences. Using I2I-Bench, we benchmark numerous mainstream image editing models, investigating the gaps and trade-offs between editing models across various dimensions. We will open-source all components of I2I-Bench to facilitate future research.

📄 PDF Abstract BibTeX arXiv:2512.04660

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

2024-11-20 · Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He 외

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do…

BenchmarkingImage GenerationImage to Video GenerationVideo Generation

Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection

2026-06-15 · Markus Bujotzek, Dimitrios Bounias, Stefan Denner, Ralf Floca 외 arxiv

While federated learning (FL) enables collaborative medical image segmentation without centralizing sensitive data, real-world deployment is frequently complicated by cross-site label imperfections such as contour disagr…

Medical Image SegmentationFederated Learning

Benchmarking Suite for Synthetic Aperture Radar Imagery Anomaly Detection (SARIAD) Algorithms

2025-04-10 · Lucian Chauvina, Somil Guptac, Angelina Ibarrac, Joshua Peeples

Anomaly detection is a key research challenge in computer vision and machine learning with applications in many fields from quality control to radar imaging. In radar imaging, specifically synthetic aperture radar (SAR),…

Anomaly DetectionBenchmarking

ImgEdit: A Unified Image Editing Dataset and Benchmark

2025-05-26 · Yang Ye, Xianyi He, Zongjian Li, Bin Lin 외

Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quali…

Image EditingImage GenerationLanguage Modeling+3

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

2026-05-26 · Guixian Xu, Yide Liang, Zeli Su, Xuexian Song 외 arxiv

Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastructure. To fill this gap, we introduce FTib…

Continual Pretraining