Unseen Action Recognition with Unpaired Adversarial Multimodal Learning
In this paper, we present a method to learn a joint multimodal representation space that allows for the recognition of unseen activities in videos. We compare the effect of placing various constraints on the embedding space using paired text and video data. Additionally, we propose a method to improve the joint embedding space using an adversarial formulation with unpaired text and video data. In addition to testing on publicly available datasets, we introduce a new, large-scale text/video dataset. We experimentally confirm that learning such shared embedding space benefits three difficult tasks (i) zero-shot activity classification, (ii) unsupervised activity discovery, and (iii) unseen activity captioning.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionGeneral ClassificationTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Learning Multimodal Representations for Unseen Activities
We present a method to learn a joint multimodal representation space that enables recognition of unseen activities in videos. We first compare the effect of placing various constraints on the embedding space using paired…
General ClassificationTemporal Action LocalizationUnpaired Speech Enhancement by Acoustic and Adversarial Supervision for Speech Recognition
Many speech enhancement methods try to learn the relationship between noisy and clean speech, obtained using an acoustic room simulator. We point out several limitations of enhancement methods relying on clean speech tar…
Generative Adversarial NetworkSpeech Enhancementspeech-recognitionSpeech RecognitionMAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-based Image Captioning
Text-based image captioning (TextCap) requires simultaneous comprehension of visual content and reading the text of images to generate a natural language description. Although a task can teach machines to understand the …
Caption GenerationDescriptiveDiversityGenerative Adversarial Network+2Improving Alignment and Robustness with Circuit Breakers
AI systems can take harmful actions and are highly vulnerable to adversarial attacks. We present an approach, inspired by recent advances in representation engineering, that interrupts the models as they respond with har…
Adversarial RobustnessUnified Attentional Generative Adversarial Network for Brain Tumor Segmentation From Multimodal Unpaired Images
In medical applications, the same anatomical structures may be observed in multiple modalities despite the different image characteristics. Currently, most deep models for multimodal segmentation rely on paired registere…
Brain Tumor SegmentationGenerative Adversarial NetworkSegmentationTranslation+1