paper-with-me

Papers

Step-Wise Hierarchical Alignment Network for Image-Text Matching

2021-06-11 · Zhong Ji, Kexin Chen, Haoran Wang

Image-text matching plays a central role in bridging the semantic gap between vision and language. The key point to achieve precise visual-semantic alignment lies in capturing the fine-grained cross-modal correspondence between image and text. Most previous methods rely on single-step reasoning to discover the visual-semantic interactions, which lacks the ability of exploiting the multi-level information to locate the hierarchical fine-grained relevance. Different from them, in this work, we propose a step-wise hierarchical alignment network (SHAN) that decomposes image-text matching into multi-step cross-modal reasoning process. Specifically, we first achieve local-to-local alignment at fragment level, following by performing global-to-local and global-to-global alignment at context level sequentially. This progressive alignment strategy supplies our model with more complementary and sufficient semantic clues to understand the hierarchical correlations between image and text. The experimental results on two benchmark datasets demonstrate the superiority of our proposed method.

📄 PDF Abstract BibTeX arXiv:2106.06509

Code (0)

등록된 구현이 없습니다.

Tasks

Image-text matchingText Matching

Similar Papers 제목 키워드 기반

Hierarchical and Step-Layer-Wise Tuning of Attention Specialty for Multi-Instance Synthesis in Diffusion Transformers

2025-04-14 · ChunYang Zhang, Zhenhong Sun, Zhicheng Zhang, Junyan Wang 외

Text-to-image (T2I) generation models often struggle with multi-instance synthesis (MIS), where they must accurately depict multiple distinct instances in a single image based on complex prompts detailing individual feat…

AttributeLayout Generation

HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation

2026-08-14 · Yin Li, Ziyang Hu, Zhiyu Guo, Xiangyu Liu 외 arxiv

Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faithful evidence selection and placement. We…

EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment

2024-10-08 · Yifei Xing, Xiangyuan Lan, Ruiping Wang, Dongmei Jiang 외

Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language m…

cross-modal alignmentHallucinationMamba

HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding

2026-07-01 · Ji Ha Jang, Hayeon Kim, Chulwon Lee, Junghun James Kim 외 arxiv

CLIP (Contrastive Language-Image Pre-training) has become a de facto paradigm for image-text alignment, but it struggles with long-context descriptions (>77 tokens) due to absolute positional encoding and pretraining on …

Long-Context UnderstandingCross-Modal Retrieval

Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models

2025-04-22 · Dasol Jeong, Donggoo Kang, Jiwon Park, Hyebean Lee 외

We propose a diffusion-based framework for zero-shot image editing that unifies text-guided and reference-guided approaches without requiring fine-tuning. Our method leverages diffusion inversion and timestep-specific nu…

Attribute