Semantic-Preserving Augmentation for Robust Image-Text Retrieval
Image text retrieval is a task to search for the proper textual descriptions of the visual world and vice versa. One challenge of this task is the vulnerability to input image and text corruptions. Such corruptions are often unobserved during the training, and degrade the retrieval model decision quality substantially. In this paper, we propose a novel image text retrieval technique, referred to as robust visual semantic embedding (RVSE), which consists of novel image-based and text-based augmentation techniques called semantic preserving augmentation for image (SPAugI) and text (SPAugT). Since SPAugI and SPAugT change the original data in a way that its semantic information is preserved, we enforce the feature extractors to generate semantic aware embedding vectors regardless of the corruption, improving the model robustness significantly. From extensive experiments using benchmark datasets, we show that RVSE outperforms conventional retrieval schemes in terms of image-text retrieval performance.
Code (1)
Tasks
Image-text RetrievalRetrievalText RetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval
Every hour, huge amounts of visual contents are posted on social media and user-generated content platforms. To find relevant videos by means of a natural language query, text-video retrieval methods have received increa…
Data AugmentationRetrievalVideo RetrievalSemantic-aware Data Augmentation for Text-to-image Synthesis
Data augmentation has been recently leveraged as an effective regularizer in various vision-language deep neural networks. However, in text-to-image synthesis (T2Isyn), current augmentation wisdom still suffers from the …
Data AugmentationImage GenerationDeep Semantic Multimodal Hashing Network for Scalable Image-Text and Video-Text Retrievals
Hashing has been widely applied to multimodal retrieval on large-scale multimedia data due to its efficiency in computation and storage. In this article, we propose a novel deep semantic multimodal hashing network (DSMHN…
Cross-Modal RetrievalDeep HashingImage-text RetrievalRepresentation Learning+3The Efficacy of Semantics-Preserving Transformations in Self-Supervised Learning for Medical Ultrasound
Data augmentation is a central component of joint embedding self-supervised learning (SSL). Approaches that work for natural images may not always be effective in medical imaging tasks. This study systematically investig…
ClassificationData AugmentationDiagnosticLine Detection+1Paired Cross-Modal Data Augmentation for Fine-Grained Image-to-Text Retrieval
This paper investigates an open research problem of generating text-image pairs to improve the training of fine-grained image-to-text cross-modal retrieval task, and proposes a novel framework for paired data augmentatio…
Cross-Modal RetrievalData AugmentationImage to textImage-to-Text Retrieval+2