paper-with-me

홈 › Papers

SAC: Semantic Attention Composition for Text-Conditioned Image Retrieval

2020-09-03 · Surgan Jandial, Pinkesh Badjatiya, Pranit Chawla, Ayush Chopra, Mausoom Sarkar, Balaji Krishnamurthy

The ability to efficiently search for images is essential for improving the user experiences across various products. Incorporating user feedback, via multi-modal inputs, to navigate visual search can help tailor retrieved results to specific user queries. We focus on the task of text-conditioned image retrieval that utilizes support text feedback alongside a reference image to retrieve images that concurrently satisfy constraints imposed by both inputs. The task is challenging since it requires learning composite image-text features by incorporating multiple cross-granular semantic edits from text feedback and then applying the same to visual features. To address this, we propose a novel framework SAC which resolves the above in two major steps: "where to see" (Semantic Feature Attention) and "how to change" (Semantic Feature Modification). We systematically show how our architecture streamlines the generation of text-aware image features by removing the need for various modules required by other state-of-art techniques. We present extensive quantitative, qualitative analysis, and ablation studies, to show that our architecture SAC outperforms existing techniques by achieving state-of-the-art performance on 3 benchmark datasets: FashionIQ, Shoes, and Birds-to-Words, while supporting natural language feedback of varying lengths.

📄 PDF Abstract BibTeX arXiv:2009.01485

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalNavigateRetrieval

Similar Papers 제목 키워드 기반

Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models

2023-05-23 · Ruichen Wang, Zekang Chen, Chen Chen, Jian Ma 외

Recent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts. However, they fail to semantically align the generated images with the prompts du…

AttributeImage Generation

The Stable Artist: Steering Semantics in Diffusion Latent Space

2022-12-12 · Manuel Brack, Patrick Schramowski, Felix Friedrich, Dominik Hintersdorf 외

Large, text-conditioned generative diffusion models have recently gained a lot of attention for their impressive performance in generating high-fidelity images from text alone. However, achieving high-quality results is …

Image Generation

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval

2026-07-01 · Jingjing Zhang, Lei Zhang, Zheren Fu, Zhendong Mao arxiv

Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on costly triplets, Zero-Shot CIR (ZS-CIR) alleviates this reliance through proxy tas…

Image Retrieval

Generating Description for Sequential Images with Local-Object Attention Conditioned on Global Semantic Context

2018-11-01 · WS 2018 11 · Jing Su, Chenghua Lin, Mian Zhou, QingYun Dai 외
Image CaptioningText Generation

Training-Free Occluded Text Rendering via Glyph Priors and Attention-Guided Semantic Blending

2026-05-16 · Jingqi Hou, Hongtian Wang arxiv

We present a training-free framework for occluded text rendering with a pretrained FLUX.1-dev backbone. The task requires a model to render recognizable typography and place an occluding object over the intended text reg…