paper-with-me

홈 › Papers

Training Priors Predict Text-To-Image Model Performance

2023-05-23 · Charles Lovering, Ellie Pavlick

Text-to-image models can often generate some relations, i.e., "astronaut riding horse", but fail to generate other relations composed of the same basic parts, i.e., "horse riding astronaut". These failures are often taken as evidence that models rely on training priors rather than constructing novel images compositionally. This paper tests this intuition on the stablediffusion 2.1 text-to-image model. By looking at the subject-verb-object (SVO) triads that underlie these prompts (e.g., "astronaut", "ride", "horse"), we find that the more often an SVO triad appears in the training data, the better the model can generate an image aligned with that triad. Here, by aligned we mean that each of the terms appears in the generated image in the proper relation to each other. Surprisingly, this increased frequency also diminishes how well the model can generate an image aligned with the flipped triad. For example, if "astronaut riding horse" appears frequently in the training data, the image for "horse riding astronaut" will tend to be poorly aligned. Our results thus show that current models are biased to generate images with relations seen in training, and provide new data to the ongoing debate on whether these text-to-image models employ abstract compositional structure in a traditional sense, or rather, interpolate between relations explicitly seen in the training data.

📄 PDF Abstract BibTeX arXiv:2306.01755

Code (0)

등록된 구현이 없습니다.

Tasks

model

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Learning Visual Generative Priors without Text

2024-12-10 · CVPR 2025 1 · Shuailei Ma, Kecheng Zheng, Ying WEI, Wei Wu 외

Although text-to-image (T2I) models have recently thrived as visual generative priors, their reliance on high-quality text-image pairs makes scaling up expensive. We argue that grasping the cross-modality alignment is no…

Image to 3DPhilosophy

VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image Analysis

2024-02-27 · CVPR 2024 1 · Linshan Wu, Jiaxin Zhuang, Hao Chen

Self-Supervised Learning (SSL) has demonstrated promising results in 3D medical image analysis. However, the lack of high-level semantics in pre-training still heavily hinders the performance of downstream tasks. We obse…

Contrastive LearningMedical Image AnalysisPositionSelf-Supervised Learning

Cross-Image Contrastive Decoding: Precise, Lossless Suppression of Language Priors in Large Vision-Language Models

2025-05-15 · Jianfei Zhao, Feng Zhang, Xin Sun, Chong Feng

Language priors are a major cause of hallucinations in Large Vision-Language Models (LVLMs), often leading to text that is linguistically plausible but visually inconsistent. Recent work explores contrastive decoding as …

Image CaptioningLanguage ModelingLanguage ModellingLarge Language Model+1

Text-driven Visual Synthesis with Latent Diffusion Prior

2023-02-16 · Ting-Hsuan Liao, Songwei Ge, Yiran Xu, Yao-Chih Lee 외

There has been tremendous progress in large-scale text-to-image synthesis driven by diffusion models enabling versatile downstream applications such as 3D object synthesis from texts, image editing, and customized genera…

DecoderImage GenerationText to 3D

The Power of Triply Complementary Priors for Image Compressive Sensing

2020-05-16 · Zhiyuan Zha, Xin Yuan, Joey Tianyi Zhou, Jiantao Zhou 외

Recent works that utilized deep models have achieved superior results in various image restoration applications. Such approach is typically supervised which requires a corpus of training images with distribution similar …

Compressive SensingImage Restoration