paper-with-me

Papers

Text-to-Image Diffusion Models are Great Sketch-Photo Matchmakers

2024-03-12 · CVPR 2024 1 · Subhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, Yi-Zhe Song

This paper, for the first time, explores text-to-image diffusion models for Zero-Shot Sketch-based Image Retrieval (ZS-SBIR). We highlight a pivotal discovery: the capacity of text-to-image diffusion models to seamlessly bridge the gap between sketches and photos. This proficiency is underpinned by their robust cross-modal capabilities and shape bias, findings that are substantiated through our pilot studies. In order to harness pre-trained diffusion models effectively, we introduce a straightforward yet powerful strategy focused on two key aspects: selecting optimal feature layers and utilising visual and textual prompts. For the former, we identify which layers are most enriched with information and are best suited for the specific retrieval requirements (category-level or fine-grained). Then we employ visual and textual prompts to guide the model's feature extraction process, enabling it to generate more discriminative and contextually relevant cross-modal representations. Extensive experiments on several benchmark datasets validate significant performance improvements.

📄 PDF Abstract BibTeX arXiv:2403.07214

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalRetrievalSketch-Based Image Retrieval

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

One-shot Face Sketch Synthesis in the Wild via Generative Diffusion Prior and Instruction Tuning

2025-06-18 · Han Wu, Junyao Li, Kangbo Zhao, Sen Zhang 외

Face sketch synthesis is a technique aimed at converting face photos into sketches. Existing face sketch synthesis research mainly relies on training with numerous photo-sketch sample pairs from existing datasets. Howeve…

Face Sketch Synthesis

Text-Guided Scene Sketch-to-Photo Synthesis

2023-02-14 · AprilPyone MaungMaung, Makoto Shing, Kentaro Mitsui, Kei Sawada 외

We propose a method for scene-level sketch-to-photo synthesis with text guidance. Although object-level sketch-to-photo synthesis has been widely studied, whole-scene synthesis is still challenging without reference phot…

Self-Supervised Learning

Inversion-by-Inversion: Exemplar-based Sketch-to-Photo Synthesis via Stochastic Differential Equations without Training

2023-08-15 · XiMing Xing, Chuang Wang, Haitao Zhou, Zhihao Hu 외

Exemplar-based sketch-to-photo synthesis allows users to generate photo-realistic images based on sketches. Recently, diffusion-based methods have achieved impressive performance on image generation tasks, enabling highl…

Image Generation

Controllable Text-to-Image Generation with GPT-4

2023-05-29 · Tianjun Zhang, Yi Zhang, Vibhav Vineet, Neel Joshi 외

Current text-to-image generation models often struggle to follow textual instructions, especially the ones requiring spatial reasoning. On the other hand, Large Language Models (LLMs), such as GPT-4, have shown remarkabl…

Image GenerationInstruction FollowingSpatial ReasoningText to Image Generation+1

DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models

2023-06-26 · NeurIPS 2023 11 · XiMing Xing, Chuang Wang, Haitao Zhou, Jing Zhang 외

Even though trained mainly on images, we discover that pretrained diffusion models show impressive power in guiding sketch synthesis. In this paper, we present DiffSketcher, an innovative algorithm that creates \textit{v…