paper-with-me

홈 › Papers

HCMA: Hierarchical Cross-model Alignment for Grounded Text-to-Image Generation

2025-05-10 · Hang Wang, Zhi-Qi Cheng, Chenhao Lin, Chao Shen, Lei Zhang

Text-to-image synthesis has progressed to the point where models can generate visually compelling images from natural language prompts. Yet, existing methods often fail to reconcile high-level semantic fidelity with explicit spatial control, particularly in scenes involving multiple objects, nuanced relations, or complex layouts. To bridge this gap, we propose a Hierarchical Cross-Modal Alignment (HCMA) framework for grounded text-to-image generation. HCMA integrates two alignment modules into each diffusion sampling step: a global module that continuously aligns latent representations with textual descriptions to ensure scene-level coherence, and a local module that employs bounding-box layouts to anchor objects at specified locations, enabling fine-grained spatial control. Extensive experiments on the MS-COCO 2014 validation set show that HCMA surpasses state-of-the-art baselines, achieving a 0.69 improvement in Frechet Inception Distance (FID) and a 0.0295 gain in CLIP Score. These results demonstrate HCMA's effectiveness in faithfully capturing intricate textual semantics while adhering to user-defined spatial constraints, offering a robust solution for semantically grounded image generation. Our code is available at https://github.com/hwang-cs-ime/HCMA.

📄 PDF Abstract BibTeX arXiv:2505.06512

Code (1)

hwang-cs-ime/hcma 공식 구현

Tasks

cross-modal alignmentImage GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

2025-03-10 · Youjun Zhao, Jiaying Lin, Rynson W. H. Lau

Open-vocabulary 3D object detection (OV-3DOD) aims at localizing and classifying novel objects beyond closed sets. The recent success of vision-language models (VLMs) has demonstrated their remarkable capabilities to und…

3D Object Detectioncross-modal alignmentData IntegrationObject+2

Collective Large-scale Wind Farm Multivariate Power Output Control Based on Hierarchical Communication Multi-Agent Proximal Policy Optimization

2023-05-17 · Yubao Zhang, Xin Chen, Sumei Gong, Haojie Chen

Wind power is becoming an increasingly important source of renewable energy worldwide. However, wind farm power control faces significant challenges due to the high system complexity inherent in these farms. A novel comm…

Deep Reinforcement Learning

Efficiently Deploying LLMs with Controlled Risk

2024-10-03 · Michael J. Zellinger, Matt Thomson

Deploying large language models in production requires simultaneous attention to efficiency and risk control. Prior work has shown the possibility to cut costs while maintaining similar accuracy, but has neglected to foc…

MMLUTruthfulQA

Hierarchical Prototype-based Domain Priors for Multiple Instance Learning in Multimodal Histopathology Analysis

2026-04-27 · Xuemei Qiu, Dawei Fan, Yebin Huang, Yanping Chen 외 arxiv

Digital pathology has fundamentally altered diagnostic workflows by enabling the computational analysis of gigapixel Whole Slide Images (WSIs), yet effectively deciphering their complex tumor microenvironments remains a …

Multiple Instance Learning

HCMA-UNet: A Hybrid CNN-Mamba UNet with Axial Self-Attention for Efficient Breast Cancer Segmentation

2025-01-01 · Haoxuan Li, Wei Song, Peiwu Qin, Xi Yuan 외

Breast cancer lesion segmentation in DCE-MRI remains challenging due to heterogeneous tumor morphology and indistinct boundaries. To address these challenges, this study proposes a novel hybrid segmentation network, HCMA…

Computational EfficiencyLesion SegmentationMambaSegmentation