paper-with-me

Papers

FAR-Net: Multi-Stage Fusion Network with Enhanced Semantic Alignment and Adaptive Reconciliation for Composed Image Retrieval

2025-07-17 · Jeong-Woo Park, Young-Eun Kim, Seong-Whan Lee

Composed image retrieval (CIR) is a vision language task that retrieves a target image using a reference image and modification text, enabling intuitive specification of desired changes. While effectively fusing visual and textual modalities is crucial, existing methods typically adopt either early or late fusion. Early fusion tends to excessively focus on explicitly mentioned textual details and neglect visual context, whereas late fusion struggles to capture fine-grained semantic alignments between image regions and textual tokens. To address these issues, we propose FAR-Net, a multi-stage fusion framework designed with enhanced semantic alignment and adaptive reconciliation, integrating two complementary modules. The enhanced semantic alignment module (ESAM) employs late fusion with cross-attention to capture fine-grained semantic relationships, while the adaptive reconciliation module (ARM) applies early fusion with uncertainty embeddings to enhance robustness and adaptability. Experiments on CIRR and FashionIQ show consistent performance gains, improving Recall@1 by up to 2.4% and Recall@50 by 1.04% over existing state-of-the-art methods, empirically demonstrating that FAR Net provides a robust and scalable solution to CIR tasks.

📄 PDF Abstract BibTeX arXiv:2507.12823

Code (0)

등록된 구현이 없습니다.

Tasks

Image Retrieval

Similar Papers 제목 키워드 기반

SETR: A Two-Stage Semantic-Enhanced Framework for Zero-Shot Composed Image Retrieval

2025-09-30 · Yuqi Xiao, Yingying Zhu arxiv

Zero-shot Composed Image Retrieval (ZS-CIR) aims to retrieve a target image given a reference image and a relative text, without relying on costly triplet annotations. Existing CLIP-based methods face two core challenges…

Image Retrieval

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

2024-03-08 · XiWei Hu, Rui Wang, Yixiao Fang, Bin Fu 외

Diffusion models have demonstrated remarkable performance in the domain of text-to-image generation. However, most widely used models still employ CLIP as their text encoder, which constrains their ability to comprehend …

DenoisingImage GenerationLanguage ModellingLarge Language Model+2

RealignDiff: Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment

2023-05-31 · Zutao Jiang, Guian Fang, Jianhua Han, Guansong Lu 외

Recent advances in text-to-image diffusion models have achieved remarkable success in generating high-quality, realistic images from textual descriptions. However, these approaches have faced challenges in precisely alig…

Caption GenerationLanguage ModellingLarge Language ModelSemantic Similarity+1

Training-Free Representation Guidance for Diffusion Models with a Representation Alignment Projector

2026-01-30 · Wenqiang Zu, Shenghao Xie, Bo Lei, Lei Ma arxiv

Recent progress in generative modeling has enabled high-quality visual synthesis with diffusion-based frameworks, supporting controllable sampling and large-scale training. Inference-time guidance methods such as classif…

VIBE: Video Instruction-aligned Background music gEneration

2026-08-31 · Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj, Gouthaman KV 외 arxiv

Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck of static cross-modal …

Instruction FollowingMusic Generation