paper-with-me

Papers

A Contrastive Compositional Benchmark for Text-to-Image Synthesis: A Study with Unified Text-to-Image Fidelity Metrics

2023-12-04 · Xiangru Zhu, Penglei Sun, Chengyu Wang, Jingping Liu, Zhixu Li, Yanghua Xiao, Jun Huang

Text-to-image (T2I) synthesis has recently achieved significant advancements. However, challenges remain in the model's compositionality, which is the ability to create new combinations from known components. We introduce Winoground-T2I, a benchmark designed to evaluate the compositionality of T2I models. This benchmark includes 11K complex, high-quality contrastive sentence pairs spanning 20 categories. These contrastive sentence pairs with subtle differences enable fine-grained evaluations of T2I synthesis models. Additionally, to address the inconsistency across different metrics, we propose a strategy that evaluates the reliability of various metrics by using comparative sentence pairs. We use Winoground-T2I with a dual objective: to evaluate the performance of T2I models and the metrics used for their evaluation. Finally, we provide insights into the strengths and weaknesses of these metrics and the capabilities of current T2I models in tackling challenges across a range of complex compositional categories. Our benchmark is publicly available at https://github.com/zhuxiangru/Winoground-T2I .

📄 PDF Abstract BibTeX arXiv:2312.02338

Code (1)

zhuxiangru/winoground-t2i 공식 구현

Tasks

Image GenerationSentence

Similar Papers 제목 키워드 기반

Progressive Compositionality In Text-to-Image Generative Models

2024-10-22 · Xu Han, Linghao Jin, Xiaofeng Liu, Paul Pu Liang

Despite the impressive text-to-image (T2I) synthesis capabilities of diffusion models, they often struggle to understand compositional relationships between objects and attributes, especially in complex settings. Existin…

AttributeContrastive LearningQuestion AnsweringVisual Question Answering+1

StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis

2022-03-29 · CVPR 2022 1 · Zhiheng Li, Martin Renqiang Min, Kai Li, Chenliang Xu

Although progress has been made for text-to-image synthesis, previous methods fall short of generalizing to unseen or underrepresented attribute compositions in the input text. Lacking compositionality could have severe …

AttributeFairnessImage GenerationVocal Bursts Intensity Prediction

SCOT: Self-Supervised Contrastive Pretraining For Zero-Shot Compositional Retrieval

2025-01-12 · WACV 2025 3 · Bhavin Jawade, Joao V. B. Soares, Kapil Thadani, Deen Dayal Mohan 외

Compositional image retrieval (CIR) is a multimodal learning task where a model combines a query image with a user-provided text modification to retrieve a target image. CIR finds applications in a variety of domains inc…

Image RetrievalRetrievalTripletZero-Shot Composed Image Retrieval (ZS-CIR)

TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

2024-11-04 · Maitreya Patel, Abhiram Kusumba, Sheng Cheng, Changhoon Kim 외

Contrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the training data a significant factor in the eff…

Diversityimage-classificationImage ClassificationImage Retrieval+2

Text encoders bottleneck compositionality in contrastive vision-language models

2023-05-24 · Amita Kamath, Jack Hessel, Kai-Wei Chang

Performant vision-language (VL) models like CLIP represent captions using a single vector. How much information about language is lost in this bottleneck? We first curate CompPrompts, a set of increasingly compositional …

AttributeImage CaptioningObject