paper-with-me

홈 › Papers

HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing

2024-12-05 · Jinbin Bai, Wei Chow, Ling Yang, Xiangtai Li, Juncheng Li, Hanwang Zhang, Shuicheng Yan

We present HumanEdit, a high-quality, human-rewarded dataset specifically designed for instruction-guided image editing, enabling precise and diverse image manipulations through open-form language instructions. Previous large-scale editing datasets often incorporate minimal human feedback, leading to challenges in aligning datasets with human preferences. HumanEdit bridges this gap by employing human annotators to construct data pairs and administrators to provide feedback. With meticulously curation, HumanEdit comprises 5,751 images and requires more than 2,500 hours of human effort across four stages, ensuring both accuracy and reliability for a wide range of image editing tasks. The dataset includes six distinct types of editing instructions: Action, Add, Counting, Relation, Remove, and Replace, encompassing a broad spectrum of real-world scenarios. All images in the dataset are accompanied by masks, and for a subset of the data, we ensure that the instructions are sufficiently detailed to support mask-free editing. Furthermore, HumanEdit offers comprehensive diversity and high-resolution $1024 \times 1024$ content sourced from various domains, setting a new versatile benchmark for instructional image editing datasets. With the aim of advancing future research and establishing evaluation benchmarks in the field of image editing, we release HumanEdit at https://huggingface.co/datasets/BryanW/HumanEdit.

📄 PDF Abstract BibTeX arXiv:2412.04280

Code (1)

viiika/humanedit 공식 구현

Similar Papers 제목 키워드 기반

Diversity-Rewarded CFG Distillation

2024-10-08 · Geoffrey Cideron, Andrea Agostinelli, Johan Ferret, Sertan Girgin 외

Generative models are transforming creative domains such as music generation, with inference-time strategies like Classifier-Free Guidance (CFG) playing a crucial role. However, CFG doubles inference cost while limiting …

DiversityMusic Generation

WARP: On the Benefits of Weight Averaged Rewarded Policies

2024-06-24 · Alexandre Ramé, Johan Ferret, Nino Vieillard, Robert Dadashi 외

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) by encouraging their generations to have high rewards, using a reward model trained on human preferences. To prevent the forgetting of…

Code-A1: Adversarial Evolving of Code LLM and Test LLM via Reinforcement Learning

2026-03-16 · Aozhe Wang, Yuchen Yan, Nan Zhou, Zhengxi Lu 외 arxiv

Reinforcement learning for code generation relies on verifiable rewards from unit test pass rates. Yet high-quality test suites are scarce, existing datasets offer limited coverage, and static rewards fail to adapt as mo…

Reinforcement LearningCode Generation

Incentives shape how humans co-create with generative AI

2026-04-04 · Nathanael Jo, Manish Raghavan arxiv

Generative AI is quickly becoming an integral part of people's everyday workflows. Early evidence has shown that while generative AI can increase individual-level productivity, it does so at the cost of collective divers…

SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization

2025-11-17 · Xuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou 외 arxiv

Multimodal large language models (MLLMs) have demonstrated impressive reasoning and instruction-following capabilities, yet their expanded modality space introduces new compositional safety risks that emerge from complex…