Image Retrieval with Multi-Modal Query
3개 벤치마크 · 논문 10편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Show and Tell: A Neural Image Caption Generator
A simple neural network module for relational reasoning
FiLM: Visual Reasoning with a General Conditioning Layer
Composing Text and Image for Image Retrieval - An Empirical Odyssey
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
Compositional Learning of Image-Text Query for Image Retrieval
Papers
Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations
Composed image retrieval extends content-based image retrieval systems by enabling users to search using reference images and captions that describe their intention. Despite great progress in developing image-text compos…
Content-Based Image RetrievalImage RetrievalImage Retrieval with Multi-Modal QueryRetrieval+1Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
We investigate composed image retrieval with text feedback. Users gradually look for the target of interest by moving from coarse to fine-grained feedback. However, existing methods merely focus on the latter, i.e., fine…
Composed Image Retrieval (CoIR)Image RetrievalImage Retrieval with Multi-Modal QueryRetrievalCompositional Learning of Image-Text Query for Image Retrieval
In this paper, we investigate the problem of retrieving images from a database based on a multi-modal (image-text) query. Specifically, the query text prompts some modification in the query image and the task is to retri…
Image RetrievalImage Retrieval with Multi-Modal QueryMetric Learning+1Composing Text and Image for Image Retrieval - An Empirical Odyssey
In this paper, we study the task of image retrieval, where the input query is specified in the form of an image plus some text that describes desired modifications to the input image. For example, we may present an image…
Image RetrievalImage Retrieval with Multi-Modal QueryRetrievalAttributes as Operators: Factorizing Unseen Attribute-Object Compositions
We present a new approach to modeling visual attributes. Prior work casts attributes in a similar role as objects, learning a latent representation where properties (e.g., sliced) are recognized by classifiers much in th…
AttributeCompositional Zero-Shot LearningImage Retrieval with Multi-Modal QueryObjectFiLM: Visual Reasoning with a General Conditioning Layer
We introduce a general-purpose conditioning method for neural networks called FiLM: Feature-wise Linear Modulation. FiLM layers influence neural network computation via a simple, feature-wise affine transformation based …
Image Retrieval with Multi-Modal QueryVisual Question Answering (VQA)Visual Question Answering (VQA) Split AVisual Question Answering (VQA) Split B+1