paper-with-me

Papers

RealignDiff: Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment

2023-05-31 · Zutao Jiang, Guian Fang, Jianhua Han, Guansong Lu, Hang Xu, Shengcai Liao, Xiaojun Chang, Xiaodan Liang

Recent advances in text-to-image diffusion models have achieved remarkable success in generating high-quality, realistic images from textual descriptions. However, these approaches have faced challenges in precisely aligning the generated visual content with the textual concepts described in the prompts. In this paper, we propose a two-stage coarse-to-fine semantic re-alignment method, named RealignDiff, aimed at improving the alignment between text and images in text-to-image diffusion models. In the coarse semantic re-alignment phase, a novel caption reward, leveraging the BLIP-2 model, is proposed to evaluate the semantic discrepancy between the generated image caption and the given text prompt. Subsequently, the fine semantic re-alignment stage employs a local dense caption generation module and a re-weighting attention modulation module to refine the previously generated images from a local semantic view. Experimental results on the MS-COCO and ViLG-300 datasets demonstrate that the proposed two-stage coarse-to-fine semantic re-alignment method outperforms other baseline re-alignment techniques by a substantial margin in both visual quality and semantic similarity with the input prompt.

📄 PDF Abstract BibTeX arXiv:2305.19599

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationLanguage ModellingLarge Language ModelSemantic SimilaritySemantic Textual Similarity

Methods 이 논문이 사용한 방법론

fail 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Sherpa3D: Boosting High-Fidelity Text-to-3D Generation via Coarse 3D Prior

2023-12-11 · CVPR 2024 1 · Fangfu Liu, Diankun Wu, Yi Wei, Yongming Rao 외

Recently, 3D content creation from text prompts has demonstrated remarkable progress by utilizing 2D and 3D diffusion models. While 3D diffusion models ensure great multi-view consistency, their ability to generate high-…

3D GenerationText to 3D

Dreamix: Video Diffusion Models are General Video Editors

2023-02-02 · Eyal Molad, Eliahu Horwitz, Dani Valevski, Alex Rav Acha 외

Text-driven image and video diffusion models have recently achieved unprecedented generation realism. While diffusion models have been successfully applied for image editing, very few works have done so for video editing…

Image AnimationImage to Video GenerationSubject-driven Video GenerationText-to-Video Editing+2

Boosting Diffusion Model for Spectrogram Up-sampling in Text-to-speech: An Empirical Study

2024-06-07 · Chong Zhang, Yanqing Liu, Yang Zheng, Sheng Zhao

Scaling text-to-speech (TTS) with autoregressive language model (LM) to large-scale datasets by quantizing waveform into discrete speech tokens is making great progress to capture the diversity and expressiveness in huma…

DiversityLanguage ModelingLanguage Modellingtext-to-speech+1

Inversion Circle Interpolation: Diffusion-based Image Augmentation for Data-scarce Classification

2024-08-29 · CVPR 2025 1 · Yanghao Wang, Long Chen

Data Augmentation (DA), i.e., synthesizing faithful and diverse samples to expand the original training set, is a prevalent and effective strategy to improve the performance of various data-scarce tasks. With the powerfu…

ClassificationData AugmentationDenoisingDiversity+4

Texture Generation on 3D Meshes with Point-UV Diffusion

2023-08-21 · ICCV 2023 1 · Xin Yu, Peng Dai, Wenbo Li, Lan Ma 외

In this work, we focus on synthesizing high-quality textures on 3D meshes. We present Point-UV diffusion, a coarse-to-fine pipeline that marries the denoising diffusion model with UV mapping to generate 3D consistent and…

DenoisingTexture Synthesis