paper-with-me

홈 › Papers

FashionSAP: Symbols and Attributes Prompt for Fine-grained Fashion Vision-Language Pre-training

2023-04-11 · CVPR 2023 1 · Yunpeng Han, Lisai Zhang, Qingcai Chen, Zhijian Chen, Zhonghua Li, Jianxin Yang, Zhao Cao

Fashion vision-language pre-training models have shown efficacy for a wide range of downstream tasks. However, general vision-language pre-training models pay less attention to fine-grained domain features, while these features are important in distinguishing the specific domain tasks from general tasks. We propose a method for fine-grained fashion vision-language pre-training based on fashion Symbols and Attributes Prompt (FashionSAP) to model fine-grained multi-modalities fashion attributes and characteristics. Firstly, we propose the fashion symbols, a novel abstract fashion concept layer, to represent different fashion items and to generalize various kinds of fine-grained fashion features, making modelling fine-grained attributes more effective. Secondly, the attributes prompt method is proposed to make the model learn specific attributes of fashion items explicitly. We design proper prompt templates according to the format of fashion data. Comprehensive experiments are conducted on two public fashion benchmarks, i.e., FashionGen and FashionIQ, and FashionSAP gets SOTA performances for four popular fashion tasks. The ablation study also shows the proposed abstract fashion symbols, and the attribute prompt method enables the model to acquire fine-grained semantics in the fashion domain effectively. The obvious performance gains from FashionSAP provide a new baseline for future fashion task research.

📄 PDF Abstract BibTeX arXiv:2304.05051

Code (1)

hssip/fashionsap 공식 구현 pytorch

Tasks

Attribute

Similar Papers 제목 키워드 기반

Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions

2024-03-25 · CVPR 2025 1 · Stefan Andreas Baumann, Felix Krause, Michael Neumayr, Nick Stracke 외

In recent years, advances in text-to-image (T2I) diffusion models have substantially elevated the quality of their generated images. However, achieving fine-grained control over attributes remains a challenge due to the …

Attribute

Causality-guided Prompt Learning for Vision-language Models via Visual Granulation

2025-09-04 · Mengyu Gao, Qiulei Dong arxiv

Prompt learning has recently attracted much attention for adapting pre-trained vision-language models (e.g., CLIP) to downstream recognition tasks. However, most of the existing CLIP-based prompt learning methods only sh…

Causal Inference

Generating Compositional Scenes via Text-to-image RGBA Instance Generation

2024-11-16 · Alessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang 외

Text-to-image diffusion generative models can generate high quality images at the cost of tedious prompt engineering. Controllability can be improved by introducing layout conditioning, however existing methods lack layo…

ObjectPrompt Engineering

Composing Parts for Expressive Object Generation

2025-01-01 · CVPR 2025 1 · Harsh Rangwani, Aishwarya Agarwal, Kuldeep Kulkarni, R. Venkatesh Babu 외

Image composition and generation are processes where the artists need control over various parts of the generated images. However, the current state-of-the-art generation models, like Stable Diffusion, cannot handle …

AttributeDenoisingImage GenerationObject

GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection

2026-03-27 · Jiaming Li, Zhijia Liang, Weikai Chen, Lin Ma 외 arxiv

Fine-grained open-vocabulary object detection (FG-OVD) aims to detect novel object categories described by attribute-rich texts. While existing open-vocabulary detectors show promise at the base-category level, they unde…

Object LocalizationObject Detection