paper-with-me

Papers

Semantic-Preserving Augmentation for Robust Image-Text Retrieval

2023-03-10 · Sunwoo Kim, Kyuhong Shim, Luong Trung Nguyen, Byonghyo Shim

Image text retrieval is a task to search for the proper textual descriptions of the visual world and vice versa. One challenge of this task is the vulnerability to input image and text corruptions. Such corruptions are often unobserved during the training, and degrade the retrieval model decision quality substantially. In this paper, we propose a novel image text retrieval technique, referred to as robust visual semantic embedding (RVSE), which consists of novel image-based and text-based augmentation techniques called semantic preserving augmentation for image (SPAugI) and text (SPAugT). Since SPAugI and SPAugT change the original data in a way that its semantic information is preserved, we enforce the feature extractors to generate semantic aware embedding vectors regardless of the corruption, improving the model robustness significantly. From extensive experiments using benchmark datasets, we show that RVSE outperforms conventional retrieval schemes in terms of image-text retrieval performance.

📄 PDF Abstract BibTeX arXiv:2303.05692

Code (1)

islab-github/cdatasets 공식 구현

Tasks

Image-text RetrievalRetrievalText Retrieval

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval

2022-08-03 · Alex Falcon, Giuseppe Serra, Oswald Lanz

Every hour, huge amounts of visual contents are posted on social media and user-generated content platforms. To find relevant videos by means of a natural language query, text-video retrieval methods have received increa…

Data AugmentationRetrievalVideo Retrieval

Semantic-aware Data Augmentation for Text-to-image Synthesis

2023-12-13 · Zhaorui Tan, Xi Yang, Kaizhu Huang

Data augmentation has been recently leveraged as an effective regularizer in various vision-language deep neural networks. However, in text-to-image synthesis (T2Isyn), current augmentation wisdom still suffers from the …

Data AugmentationImage Generation

Deep Semantic Multimodal Hashing Network for Scalable Image-Text and Video-Text Retrievals

2019-01-09 · Lu Jin, Zechao Li, Jinhui Tang

Hashing has been widely applied to multimodal retrieval on large-scale multimedia data due to its efficiency in computation and storage. In this article, we propose a novel deep semantic multimodal hashing network (DSMHN…

Cross-Modal RetrievalDeep HashingImage-text RetrievalRepresentation Learning+3

The Efficacy of Semantics-Preserving Transformations in Self-Supervised Learning for Medical Ultrasound

2025-04-10 · Blake VanBerlo, Alexander Wong, Jesse Hoey, Robert Arntfield

Data augmentation is a central component of joint embedding self-supervised learning (SSL). Approaches that work for natural images may not always be effective in medical imaging tasks. This study systematically investig…

ClassificationData AugmentationDiagnosticLine Detection+1

Paired Cross-Modal Data Augmentation for Fine-Grained Image-to-Text Retrieval

2022-07-29 · Hao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan Miao

This paper investigates an open research problem of generating text-image pairs to improve the training of fine-grained image-to-text cross-modal retrieval task, and proposes a novel framework for paired data augmentatio…

Cross-Modal RetrievalData AugmentationImage to textImage-to-Text Retrieval+2