paper-with-me

홈 › Papers

Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models

2024-08-01 · Juntu Zhao, Junyu Deng, Yixin Ye, Chongxuan Li, Zhijie Deng, Dequan Wang

Advancements in text-to-image diffusion models have broadened extensive downstream practical applications, but such models often encounter misalignment issues between text and image. Taking the generation of a combination of two disentangled concepts as an example, say given the prompt "a tea cup of iced coke", existing models usually generate a glass cup of iced coke because the iced coke usually co-occurs with the glass cup instead of the tea one during model training. The root of such misalignment is attributed to the confusion in the latent semantic space of text-to-image diffusion models, and hence we refer to the "a tea cup of iced coke" phenomenon as Latent Concept Misalignment (LC-Mis). We leverage large language models (LLMs) to thoroughly investigate the scope of LC-Mis, and develop an automated pipeline for aligning the latent semantics of diffusion models to text prompts. Empirical assessments confirm the effectiveness of our approach, substantially reducing LC-Mis errors and enhancing the robustness and versatility of text-to-image diffusion models. The code and dataset are here: https://github.com/RossoneriZhao/iced_coke.

📄 PDF Abstract BibTeX arXiv:2408.00230

Code (1)

rossonerizhao/iced_coke 공식 구현 jax

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts

2024-03-17 · Michael Saxon, Yiran Luo, Sharon Levy, Chitta Baral 외

Benchmarks of the multilingual capabilities of text-to-image (T2I) models compare generated images prompted in a test language to an expected image distribution over a concept set. One such benchmark, "Conceptual Coverag…

Translation

Lost in Translations? Building Sentiment Lexicons using Context Based Machine Translation

2012-12-01 · COLING 2012 12 · Xinfan Meng, Furu Wei, Ge Xu, Longkai Zhang 외
Machine TranslationSentiment AnalysisTranslation

Aligning Translation-Specific Understanding to General Understanding in Large Language Models

2024-01-10 · Yichong Huang, Baohang Li, Xiaocheng Feng, Chengpeng Fu 외

Large Language models (LLMs) have exhibited remarkable abilities in understanding complex texts, offering a promising path towards human-like translation performance. However, this study reveals the misalignment between …

Machine TranslationTranslation

Associative Texture Is Lost In Translation

2013-08-01 · WS 2013 8 · Beata Beigman Klebanov, Michael Flor
Machine TranslationTranslation

On the Value of Cross-Modal Misalignment in Multimodal Representation Learning

2025-04-14 · Yichao Cai, Yuhang Liu, Erdun Gao, Tianjiao Jiang 외

Multimodal representation learning, exemplified by multimodal contrastive learning (MMCL) using image-text pairs, aims to learn powerful representations by aligning cues across modalities. This approach relies on the cor…

Contrastive LearningRepresentation LearningSelection bias