paper-with-me

홈 › Papers

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models

2025-01-01 · Emily Johnson, Noah Wilson

Text-to-image generation has witnessed significant advancements with the integration of Large Vision-Language Models (LVLMs), yet challenges remain in aligning complex textual descriptions with high-quality, visually coherent images. This paper introduces the Vision-Language Aligned Diffusion (VLAD) model, a generative framework that addresses these challenges through a dual-stream strategy combining semantic alignment and hierarchical diffusion. VLAD utilizes a Contextual Composition Module (CCM) to decompose textual prompts into global and local representations, ensuring precise alignment with visual features. Furthermore, it incorporates a multi-stage diffusion process with hierarchical guidance to generate high-fidelity images. Experiments conducted on MARIO-Eval and INNOVATOR-Eval benchmarks demonstrate that VLAD significantly outperforms state-of-the-art methods in terms of image quality, semantic alignment, and text rendering accuracy. Human evaluations further validate the superior performance of VLAD, making it a promising approach for text-to-image generation in complex scenarios.

📄 PDF Abstract BibTeX arXiv:2501.00917

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Step-Wise Hierarchical Alignment Network for Image-Text Matching

2021-06-11 · Zhong Ji, Kexin Chen, Haoran Wang

Image-text matching plays a central role in bridging the semantic gap between vision and language. The key point to achieve precise visual-semantic alignment lies in capturing the fine-grained cross-modal correspondence …

Image-text matchingText Matching

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

2026-01-28 · Kun Wang, Xiao Feng, Mingcheng Qu, Tonghua Su arxiv

Vision Language Action (VLA) models have recently shown great potential in bridging multimodal perception with robotic control. However, existing methods often rely on direct fine-tuning of pre-trained Vision-Language Mo…

Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning

2024-03-21 · Hasindri Watawana, Kanchana Ranasinghe, Tariq Mahmood, Muzammal Naseer 외

Self-supervised representation learning has been highly promising for histopathology image analysis with numerous approaches leveraging their patient-slide-patch hierarchy to learn better representations. In this paper, …

Representation LearningSelf-Supervised Learning

IMITATE: Clinical Prior Guided Hierarchical Vision-Language Pre-training

2023-10-11 · Che Liu, Sibo Cheng, Miaojing Shi, Anand Shah 외

In the field of medical Vision-Language Pre-training (VLP), significant efforts have been devoted to deriving text and image features from both clinical reports and associated medical images. However, most existing metho…

Contrastive LearningDescriptive

GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning

2026-03-10 · Xiao Yang, Ronghao Fu, Zhuoran Duan, Zhiwen Lin 외 arxiv

Vision-language pretraining models have made significant progress in bridging remote sensing imagery with natural language. However, existing approaches often fail to effectively integrate multi-granular visual and textu…