Towards Scalable Human-aligned Benchmark for Text-guided Image Editing
A variety of text-guided image editing models have been proposed recently. However, there is no widely-accepted standard evaluation method mainly due to the subjective nature of the task, letting researchers rely on manual user study. To address this, we introduce a novel Human-Aligned benchmark for Text-guided Image Editing (HATIE). Providing a large-scale benchmark set covering a wide range of editing tasks, it allows reliable evaluation, not limited to specific easy-to-evaluate cases. Also, HATIE provides a fully-automated and omnidirectional evaluation pipeline. Particularly, we combine multiple scores measuring various aspects of editing so as to align with human perception. We empirically verify that the evaluation of HATIE is indeed human-aligned in various aspects, and provide benchmark results on several state-of-the-art models to provide deeper insights on their performance.
Code (1)
Tasks
text-guided-image-editingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation
We introduce MILE-RefHumEval, a reference-free framework for evaluating Large Language Models (LLMs) without ground-truth annotations or evaluator coordination. It leverages an ensemble of independently prompted evaluato…
Image CaptioningDial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service Dialogues
Discovering customer intentions in dialogue conversations is crucial for automated service agents. However, existing intent clustering methods often fail to align with human perceptions due to a heavy reliance on embeddi…
ClusteringCoherence EvaluationIntent DiscoveryText ClusteringEditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing
Recent text-guided image editing (TIE) models have achieved remarkable progress, while many edited images still suffer from issues such as artifacts, unexpected editings, unaesthetic contents. Although some benchmarks an…
Reinforcement LearningImage EditingReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and robotics. However, this task faces critic…
Motion GenerationUReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models
Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent textual and visual modalities are aligned. To investigate thi…
Image Generation