paper-with-me

Papers

Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation

2024-06-13 · Yufan Zhou, Ruiyi Zhang, Kaizhi Zheng, Nanxuan Zhao, Jiuxiang Gu, Zichao Wang, Xin Eric Wang, Tong Sun

In subject-driven text-to-image generation, recent works have achieved superior performance by training the model on synthetic datasets containing numerous image pairs. Trained on these datasets, generative models can produce text-aligned images for specific subject from arbitrary testing image in a zero-shot manner. They even outperform methods which require additional fine-tuning on testing images. However, the cost of creating such datasets is prohibitive for most researchers. To generate a single training pair, current methods fine-tune a pre-trained text-to-image model on the subject image to capture fine-grained details, then use the fine-tuned model to create images for the same subject based on creative text prompts. Consequently, constructing a large-scale dataset with millions of subjects can require hundreds of thousands of GPU hours. To tackle this problem, we propose Toffee, an efficient method to construct datasets for subject-driven editing and generation. Specifically, our dataset construction does not need any subject-level fine-tuning. After pre-training two generative models, we are able to generate infinite number of high-quality samples. We construct the first large-scale dataset for subject-driven image editing and generation, which contains 5 million image pairs, text prompts, and masks. Our dataset is 5 times the size of previous largest dataset, yet our cost is tens of thousands of GPU hours lower. To test the proposed dataset, we also propose a model which is capable of both subject-driven image editing and generation. By simply training the model on our proposed dataset, it obtains competitive results, illustrating the effectiveness of the proposed dataset construction framework.

📄 PDF Abstract BibTeX arXiv:2406.09305

Code (0)

등록된 구현이 없습니다.

Tasks

GPUImage GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale

2026-07-07 · Ziting Wang, Yin Li, Zuhao Yang, Xiuchang Li 외 arxiv

LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in …

Decision Making

Temporal Network Embedding via Tensor Factorization

2021-08-22 · Jing Ma, Qiuchen Zhang, Jian Lou, Li Xiong 외

Representation learning on static graph-structured data has shown a significant impact on many real-world applications. However, less attention has been paid to the evolving nature of temporal networks, in which the edge…

Link PredictionNetwork EmbeddingRepresentation LearningTensor Decomposition

Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report

2025-07-09 · Li Du, Hanyu Zhao, Yiming Ju, Tengfei Pan arxiv

Instruction tuning has become a foundation for unlocking the capabilities of large-scale pretrained models and improving their performance on complex tasks. Thus, the construction of high-quality instruction datasets is …

Instruction Following

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

2025-05-26 · Shenghai Yuan, Xianyi He, Yufan Deng, Yang Ye 외

Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we pr…

Human-Domain Subject-to-VideoOpen-Domain Subject-to-VideoSingle-Domain Subject-to-VideoVideo Generation

Subjective Image Quality Assessment with Boosted Triplet Comparisons

2021-07-31 · Hui Men, Hanhe Lin, Mohsen Jenadeleh, Dietmar Saupe

In subjective full-reference image quality assessment, differences between perceptual image qualities of the reference image and its distorted versions are evaluated, often using degradation category ratings (DCR). Howev…

Full reference image quality assessmentFull-Reference Image Quality AssessmentImage Quality AssessmentTriplet