Papers Image Retrieval with Multi-Modal Query
“Image Retrieval with Multi-Modal Query” 태그가 달린 논문 10편 · 필터 해제
Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations
Composed image retrieval extends content-based image retrieval systems by enabling users to search using reference images and captions that describe their intention. Despite great progress in developing image-text compos…
Content-Based Image RetrievalImage RetrievalImage Retrieval with Multi-Modal QueryRetrieval+1Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
We investigate composed image retrieval with text feedback. Users gradually look for the target of interest by moving from coarse to fine-grained feedback. However, existing methods merely focus on the latter, i.e., fine…
Composed Image Retrieval (CoIR)Image RetrievalImage Retrieval with Multi-Modal QueryRetrievalCompositional Learning of Image-Text Query for Image Retrieval
In this paper, we investigate the problem of retrieving images from a database based on a multi-modal (image-text) query. Specifically, the query text prompts some modification in the query image and the task is to retri…
Image RetrievalImage Retrieval with Multi-Modal QueryMetric Learning+1Composing Text and Image for Image Retrieval - An Empirical Odyssey
In this paper, we study the task of image retrieval, where the input query is specified in the form of an image plus some text that describes desired modifications to the input image. For example, we may present an image…
Image RetrievalImage Retrieval with Multi-Modal QueryRetrievalAttributes as Operators: Factorizing Unseen Attribute-Object Compositions
We present a new approach to modeling visual attributes. Prior work casts attributes in a similar role as objects, learning a latent representation where properties (e.g., sliced) are recognized by classifiers much in th…
AttributeCompositional Zero-Shot LearningImage Retrieval with Multi-Modal QueryObjectFiLM: Visual Reasoning with a General Conditioning Layer
We introduce a general-purpose conditioning method for neural networks called FiLM: Feature-wise Linear Modulation. FiLM layers influence neural network computation via a simple, feature-wise affine transformation based …
Image Retrieval with Multi-Modal QueryVisual Question Answering (VQA)Visual Question Answering (VQA) Split AVisual Question Answering (VQA) Split B+1Automatic Spatially-aware Fashion Concept Discovery
This paper proposes an automatic spatially-aware concept discovery approach using weakly labeled image-text data from shopping websites. We first fine-tune GoogleNet by jointly modeling clothing images and their correspo…
AttributeClusteringImage Retrieval with Multi-Modal QueryRetrievalA simple neural network module for relational reasoning
Relational reasoning is a central component of generally intelligent behavior, but has proven difficult for neural networks to learn. In this paper we describe how to use Relation Networks (RNs) as a simple plug-and-play…
Image Retrieval with Multi-Modal QueryQuestion AnsweringRelational ReasoningVisual Question Answering+1Image Question Answering using Convolutional Neural Network with Dynamic Parameter Prediction
We tackle image question answering (ImageQA) problem by learning a convolutional neural network (CNN) with a dynamic parameter layer whose weights are determined adaptively based on questions. For the adaptive parameter …
Image Retrieval with Multi-Modal QueryParameter PredictionPredictionQuestion Answering+1Show and Tell: A Neural Image Caption Generator
Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a …
Image CaptioningImage Retrieval with Multi-Modal QuerySentenceText Generation+2