paper-with-me

홈 › Papers

Cross-modal Search Method of Technology Video based on Adversarial Learning and Feature Fusion

2022-10-11 · Xiangbin Liu, Junping Du, Meiyu Liang, Ang Li

Technology videos contain rich multi-modal information. In cross-modal information search, the data features of different modalities cannot be compared directly, so the semantic gap between different modalities is a key problem that needs to be solved. To address the above problems, this paper proposes a novel Feature Fusion based Adversarial Cross-modal Retrieval method (FFACR) to achieve text-to-video matching, ranking and searching. The proposed method uses the framework of adversarial learning to construct a video multimodal feature fusion network and a feature mapping network as generator, a modality discrimination network as discriminator. Multi-modal features of videos are obtained by the feature fusion network. The feature mapping network projects multi-modal features into the same semantic space based on semantics and similarity. The modality discrimination network is responsible for determining the original modality of features. Generator and discriminator are trained alternately based on adversarial learning, so that the data obtained by the feature mapping network is semantically consistent with the original data and the modal features are eliminated, and finally the similarity is used to rank and obtain the search results in the semantic space. Experimental results demonstrate that the proposed method performs better in text-to-video search than other existing methods, and validate the effectiveness of the method on the self-built datasets of technology videos.

📄 PDF Abstract BibTeX arXiv:2210.05243

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalRetrievalText-to-video search

Similar Papers 제목 키워드 기반

Adversarial Attacks in Multimodal Systems: A Practitioner's Survey

2025-05-06 · Shashank Kapoor, Sanjay Surendranath Girija, Lakshit Arora, Dipen Pradhan 외

The introduction of multimodal models is a huge step forward in Artificial Intelligence. A single model is trained to understand multiple modalities: text, image, video, and audio. Open-source multimodal models have made…

Adversarial AttackSurvey

Cross-Modal Transferable Adversarial Attacks from Images to Videos

2021-12-10 · CVPR 2022 1 · Zhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang

Recent studies have shown that adversarial examples hand-crafted on one white-box model can be used to attack other black-box models. Such cross-model transferability makes it feasible to perform black-box attacks, which…

Video Recognition

Tree-based Text-Vision BERT for Video Search in Baidu Video Advertising

2022-09-19 · Tan Yu, Jie Liu, Yi Yang, Yi Li 외

The advancement of the communication technology and the popularity of the smart phones foster the booming of video ads. Baidu, as one of the leading search engine companies in the world, receives billions of search queri…

Image RetrievalRetrievalVideo Retrieval

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs

2025-07-01 · Jiaming Zhang, Rui Hu, Qing Guo, Wei Yang Bryan Lim

Video Multimodal Large Language Models (V-MLLMs) have shown impressive capabilities in temporal reasoning and cross-modal understanding, yet their vulnerability to adversarial attacks remains underexplored due to unique …

Text GenerationVideo Understanding

Revealing the Truth with ConLLM for Detecting Multi-Modal Deepfakes

2026-01-24 · Gautam Siddharth Kashyap, Harsh Joshi, Niharika Jain, Ebad Shabbir 외 arxiv

The rapid rise of deepfake technology poses a severe threat to social and political stability by enabling hyper-realistic synthetic media capable of manipulating public perception. However, existing detection methods str…

Contrastive LearningDeepFake Detection