paper-with-me

Papers

A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval

2022-08-03 · Alex Falcon, Giuseppe Serra, Oswald Lanz

Every hour, huge amounts of visual contents are posted on social media and user-generated content platforms. To find relevant videos by means of a natural language query, text-video retrieval methods have received increased attention over the past few years. Data augmentation techniques were introduced to increase the performance on unseen test examples by creating new training samples with the application of semantics-preserving techniques, such as color space or geometric transformations on images. Yet, these techniques are usually applied on raw data, leading to more resource-demanding solutions and also requiring the shareability of the raw data, which may not always be true, e.g. copyright issues with clips from movies or TV series. To address this shortcoming, we propose a multimodal data augmentation technique which works in the feature space and creates new videos and captions by mixing semantically similar samples. We experiment our solution on a large scale public dataset, EPIC-Kitchens-100, and achieve considerable improvements over a baseline method, improved state-of-the-art performance, while at the same time performing multiple ablation studies. We release code and pretrained models on Github at https://github.com/aranciokov/FSMMDA_VideoRetrieval.

📄 PDF Abstract BibTeX arXiv:2208.02080

Code (1)

aranciokov/fsmmda_videoretrieval 공식 구현 pytorch

Tasks

Data AugmentationRetrievalVideo Retrieval

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Learning Multimodal Data Augmentation in Feature Space

2022-12-29 · Zichang Liu, Zhiqiang Tang, Xingjian Shi, Aston Zhang 외

The ability to jointly learn from multiple modalities, such as text, audio, and visual data, is a defining feature of intelligent systems. While there have been promising advances in designing neural networks to harness …

Data Augmentationimage-classificationImage ClassificationMultimodal Deep Learning

Batch Augmentation with Unimodal Fine-tuning for Multimodal Learning

2025-05-10 · H M Dipu Kabir, Subrota Kumar Mondal, Mohammad Ali Moni

This paper proposes batch augmentation with unimodal fine-tuning to detect the fetus's organs from ultrasound images and associated clinical textual information. We also prescribe pre-training initial layers with investi…

Image AugmentationLarge Language ModelMultimodal Large Language Model

LIIR at SemEval-2021 task 6: Detection of Persuasion Techniques In Texts and Images using CLIP features

2021-05-31 · SEMEVAL 2021 · Erfan Ghadery, Damien Sileo, Marie-Francine Moens

We describe our approach for SemEval-2021 task 6 on detection of persuasion techniques in multimodal content (memes). Our system combines pretrained multimodal models (CLIP) and chained classifiers. Also, we propose to e…

Data Augmentation

Data Clustering as an Emergent Consensus of Autonomous Agents

2022-04-22 · Piotr Minakowski, Jan Peszek

We present a data segmentation method based on a first-order density-induced consensus protocol. We provide a mathematically rigorous analysis of the consensus model leading to the stopping criteria of the data segmentat…

ClusteringSegmentation

TextAug: Test time Text Augmentation for Multimodal Person Re-identification

2023-12-04 · Mulham Fawakherji, Eduard Vazquez, Pasquale Giampa, Binod Bhattarai

Multimodal Person Reidentification is gaining popularity in the research community due to its effectiveness compared to counter-part unimodal frameworks. However, the bottleneck for multimodal deep learning is the need f…

Data AugmentationMultimodal Deep LearningPerson Re-IdentificationSentence+1