paper-with-me

Papers

OneActor: Consistent Character Generation via Cluster-Conditioned Guidance

2024-04-16 · Jiahao Wang, Caixia Yan, Haonan Lin, Weizhan Zhang, Mengmeng Wang, Tieliang Gong, Guang Dai, Hao Sun

Text-to-image diffusion models benefit artists with high-quality image generation. Yet their stochastic nature hinders artists from creating consistent images of the same subject. Existing methods try to tackle this challenge and generate consistent content in various ways. However, they either depend on external restricted data or require expensive tuning of the diffusion model. For this issue, we propose a novel one-shot tuning paradigm, termed OneActor. It efficiently performs consistent subject generation solely driven by prompts via a learned semantic guidance to bypass the laborious backbone tuning. We lead the way to formalize the objective of consistent subject generation from a clustering perspective, and thus design a cluster-conditioned model. To mitigate the overfitting challenge shared by one-shot tuning pipelines, we augment the tuning with auxiliary samples and devise two inference strategies: semantic interpolation and cluster guidance. These techniques are later verified to significantly improve the generation quality. Comprehensive experiments show that our method outperforms a variety of baselines with satisfactory subject consistency, superior prompt conformity as well as high image quality. Our method is capable of multi-subject generation and compatible with popular diffusion extensions. Besides, we achieve a 4 times faster tuning speed than tuning-based baselines and, if desired, avoid increasing the inference time. Furthermore, our method can be naturally utilized to pre-train a consistent subject generation network from scratch, which will implement this research task into more practical applications. (Project page: https://johnneywang.github.io/OneActor-webpage/)

📄 PDF Abstract BibTeX arXiv:2404.10267

Code (0)

등록된 구현이 없습니다.

Tasks

Consistent Character GenerationDenoisingImage Generation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Double cycle-consistent generative adversarial network for unsupervised conditional generation

2019-11-13 · Fei Ding, Feng Luo, Yin Yang

Conditional generative models have achieved considerable success in the past few years, but usually require a lot of labeled data. Recently, ClusterGAN combines GAN with an encoder to achieve remarkable clustering perfor…

ClusteringDisentanglementDiversityGenerative Adversarial Network

Image Clustering Conditioned on Text Criteria

2023-10-27 · Sehyun Kwon, Jaeseung Park, Minkyu Kim, Jaewoong Cho 외

Classical clustering methods do not provide users with direct control of the clustering results, and the clustering results may not be consistent with the relevant criterion that a user has in mind. In this work, we pres…

ClusteringImage Clustering

Game Level Clustering and Generation using Gaussian Mixture VAEs

2020-08-22 · Yang Zhihan, Sarkar Anurag, Cooper Seth

Variational autoencoders (VAEs) have been shown to be able to generate game levels but require manual exploration of the learned latent space to generate outputs with desired attributes. While conditional VAEs address th…

Clustering

Make-A-Story: Visual Memory Conditioned Consistent Story Generation

2022-11-23 · CVPR 2023 1 · Tanzila Rahman, Hsin-Ying Lee, Jian Ren, Sergey Tulyakov 외

There has been a recent explosion of impressive generative models that can produce high quality images (or videos) conditioned on text descriptions. However, all such approaches rely on conditional sentences that contain…

SentenceStory GenerationStory Visualization

Exploring Intra-Class Variation Factors With Learnable Cluster Prompts for Semi-Supervised Image Synthesis

2023-01-01 · CVPR 2023 1 · Yunfei Zhang, Xiaoyang Huo, Tianyi Chen, Si Wu 외

Semi-supervised class-conditional image synthesis is typically performed by inferring and injecting class labels into a conditional Generative Adversarial Network (GAN). The supervision in the form of class identity …

Conditional Image GenerationGenerative Adversarial NetworkImage Generation