paper-with-me

홈 › Papers

CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

2024-04-04 · Dongzhi Jiang, Guanglu Song, Xiaoshi Wu, Renrui Zhang, Dazhong Shen, Zhuofan Zong, Yu Liu, Hongsheng Li

Diffusion models have demonstrated great success in the field of text-to-image generation. However, alleviating the misalignment between the text prompts and images is still challenging. The root reason behind the misalignment has not been extensively investigated. We observe that the misalignment is caused by inadequate token attention activation. We further attribute this phenomenon to the diffusion model's insufficient condition utilization, which is caused by its training paradigm. To address the issue, we propose CoMat, an end-to-end diffusion model fine-tuning strategy with an image-to-text concept matching mechanism. We leverage an image captioning model to measure image-to-text alignment and guide the diffusion model to revisit ignored tokens. A novel attribute concentration module is also proposed to address the attribute binding problem. Without any image or human preference data, we use only 20K text prompts to fine-tune SDXL to obtain CoMat-SDXL. Extensive experiments show that CoMat-SDXL significantly outperforms the baseline model SDXL in two text-to-image alignment benchmarks and achieves start-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2404.03653

Code (2)

caraj7/comat 공식 구현 pytorch
mlpc-ucsd/TokenCompose pytorch

Tasks

AttributeImage CaptioningImage GenerationImage to textText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Aligning Diffusion Models by Optimizing Human Utility

2024-04-06 · Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato 외

We present Diffusion-KTO, a novel approach for aligning text-to-image diffusion models by formulating the alignment objective as the maximization of expected human utility. Since this objective applies to each generation…

CoMatch: Semi-supervised Learning with Contrastive Graph Regularization

2020-11-23 · ICCV 2021 10 · Junnan Li, Caiming Xiong, Steven Hoi

Semi-supervised learning has been an effective paradigm for leveraging unlabeled data to reduce the reliance on labeled data. We propose CoMatch, a new semi-supervised learning method that unifies dominant approaches and…

Contrastive LearningRepresentation LearningSelf-Supervised LearningSemi-Supervised Image Classification

CoMatcher: Multi-View Collaborative Feature Matching

2025-04-02 · CVPR 2025 1 · Jintao Zhang, Zimin Xia, Mingyue Dong, Shuhan Shen 외

This paper proposes a multi-view collaborative matching strategy for reliable track construction in complex scenarios. We observe that the pairwise matching paradigms applied to image set matching often result in ambiguo…

Scene Understandingset matching

Real-World Image Variation by Aligning Diffusion Inversion Chain

2023-05-30 · NeurIPS 2023 11 · Yuechen Zhang, Jinbo Xing, Eric Lo, Jiaya Jia

Recent diffusion model advancements have enabled high-fidelity images to be generated using text prompts. However, a domain gap exists between generated images and real-world images, which poses a challenge in generating…

Image GenerationImage-VariationSemantic SimilaritySemantic Textual Similarity+2

A Novel Approach to Image EEG Sleep Data for Improving Quality of Life in Patients Suffering From Brain Injuries Using DreamDiffusion

2024-07-02 · David Fahim, Joshveer Grewal, Ritvik Ellendula

Those experiencing strokes, traumatic brain injuries, and drug complications can often end up hospitalized and diagnosed with coma or locked-in syndrome. Such mental impediments can permanently alter the neurological pat…

EEG