paper-with-me

홈 › Papers

FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models

2024-12-10 · Tong Wu, Yinghao Xu, Ryan Po, Mengchen Zhang, Guandao Yang, Jiaqi Wang, Ziwei Liu, Dahua Lin, Gordon Wetzstein

Recent advances in text-to-image generation have enabled the creation of high-quality images with diverse applications. However, accurately describing desired visual attributes can be challenging, especially for non-experts in art and photography. An intuitive solution involves adopting favorable attributes from the source images. Current methods attempt to distill identity and style from source images. However, "style" is a broad concept that includes texture, color, and artistic elements, but does not cover other important attributes such as lighting and dynamics. Additionally, a simplified "style" adaptation prevents combining multiple attributes from different sources into one generated image. In this work, we formulate a more effective approach to decompose the aesthetics of a picture into specific visual attributes, allowing users to apply characteristics such as lighting, texture, and dynamics from different images. To achieve this goal, we constructed the first fine-grained visual attributes dataset (FiVA) to the best of our knowledge. This FiVA dataset features a well-organized taxonomy for visual attributes and includes around 1 M high-quality generated images with visual attribute annotations. Leveraging this dataset, we propose a fine-grained visual attribute adaptation framework (FiVA-Adapter), which decouples and adapts visual attributes from one or more source images into a generated one. This approach enhances user-friendly customization, allowing users to selectively apply desired attributes to create images that meet their unique preferences and specific content requirements.

📄 PDF Abstract BibTeX arXiv:2412.07674

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

FIVA: Facial Image and Video Anonymization and Anonymization Defense

2023-09-08 · Felix Rosberg, Eren Erdal Aksoy, Cristofer Englund, Fernando Alonso-Fernandez

In this paper, we present a new approach for facial anonymization in images and videos, abbreviated as FIVA. Our proposed method is able to maintain the same face anonymization consistently over frames with our suggested…

Face AnonymizationFace SwappingReconstruction Attack

A$^2$-Net: Learning Attribute-Aware Hash Codes for Large-Scale Fine-Grained Image Retrieval

2021-12-01 · NeurIPS 2021 12 · Xiu-Shen Wei, Yang shen, Xuhao Sun, Han-Jia Ye 외

Our work focuses on tackling large-scale fine-grained image retrieval as ranking the images depicting the concept of interests (i.e., the same sub-category labels) highest based on the fine-grained details in the query. …

AttributeDecoderImage RetrievalRetrieval

Fine-Grained Visual Comparisons with Local Learning

2014-06-01 · CVPR 2014 6 · Aron Yu, Kristen Grauman

Given two images, we want to predict which exhibits a particular visual attribute more than the other---even when the two images are quite similar. Existing relative attribute methods rely on global ranking functions; y…

Attribute

Attribute-Aware Deep Hashing with Self-Consistency for Large-Scale Fine-Grained Image Retrieval

2023-11-21 · Xiu-Shen Wei, Yang shen, Xuhao Sun, Peng Wang 외

Our work focuses on tackling large-scale fine-grained image retrieval as ranking the images depicting the concept of interests (i.e., the same sub-category labels) highest based on the fine-grained details in the query. …

AttributeDeep HashingImage ReconstructionImage Retrieval+1

Learning to Parameterize Visual Attributes for Open-set Fine-grained Retrieval

2023-09-21 · NeurIPS 2023 11

Open-set fine-grained retrieval is an emerging challenging task that allows to retrieve unknown categories beyond the training set. The best solution for handling unknown categories is to represent them using a set of v…