paper-with-me

홈 › Papers

Learning Cross-modal Contrastive Features for Video Domain Adaptation

2021-08-26 · ICCV 2021 10 · Donghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu, Stan Sclaroff, Kate Saenko, Manmohan Chandraker

Learning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation methods mainly rely on adversarial feature alignment, which has been derived from the RGB image space. However, video data is usually associated with multi-modal information, e.g., RGB and optical flow, and thus it remains a challenge to design a better method that considers the cross-modal inputs under the cross-domain adaptation setting. To this end, we propose a unified framework for video domain adaptation, which simultaneously regularizes cross-modal and cross-domain feature representations. Specifically, we treat each modality in a domain as a view and leverage the contrastive learning technique with properly designed sampling strategies. As a result, our objectives regularize feature spaces, which originally lack the connection across modalities or have less alignment across domains. We conduct experiments on domain adaptive action recognition benchmark datasets, i.e., UCF, HMDB, and EPIC-Kitchens, and demonstrate the effectiveness of our components against state-of-the-art algorithms.

📄 PDF Abstract BibTeX arXiv:2108.11974

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionContrastive LearningDomain AdaptationOptical Flow Estimation

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Video Question Answering Using CLIP-Guided Visual-Text Attention

2023-03-06 · Shuhong Ye, Weikai Kong, Chenglin Yao, Jianfeng Ren 외

Cross-modal learning of video and text plays a key role in Video Question Answering (VideoQA). In this paper, we propose a visual-text attention mechanism to utilize the Contrastive Language-Image Pre-training (CLIP) tra…

General KnowledgeQuestion AnsweringVideo Question Answering

Cross-View Cross-Modal Unsupervised Domain Adaptation for Driver Monitoring System

2025-11-15 · Aditi Bhalla, Christian Hellert, Enkelejda Kasneci arxiv

Driver distraction remains a leading cause of road traffic accidents, contributing to thousands of fatalities annually across the globe. While deep learning-based driver activity recognition methods have shown promise in…

Unsupervised Domain AdaptationContrastive LearningActivity Recognition

Spatio-temporal Contrastive Domain Adaptation for Action Recognition

2021-06-19 · CVPR 2021 1 · Xiaolin Song, Sicheng Zhao, Jingyu Yang, Huanjing Yue 외

Unsupervised domain adaptation (UDA) for human action recognition is a practical and challenging problem. Compared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial rep…

Action RecognitionContrastive LearningDomain AdaptationSelf-Supervised Learning+2

Multi-Modal Video Topic Segmentation with Dual-Contrastive Domain Adaptation

2023-11-30 · Linzi Xing, Quan Tran, Fabian Caba, Franck Dernoncourt 외

Video topic segmentation unveils the coarse-grained semantic structure underlying videos and is essential for other video understanding tasks. Given the recent surge in multi-modal, relying solely on a single modality is…

Contrastive LearningDomain AdaptationSegmentationUnsupervised Domain Adaptation+1

Towards Contrastive Learning in Music Video Domain

2023-09-01 · Karel Veldkamp, Mariya Hendriksen, Zoltán Szlávik, Alexander Keijser

Contrastive learning is a powerful way of learning multimodal representations across various domains such as image-caption retrieval and audio-visual representation learning. In this work, we investigate if these finding…

Contrastive LearningGenre classificationMusic TaggingRepresentation Learning+1