paper-with-me

Papers

Unseen Action Recognition with Unpaired Adversarial Multimodal Learning

2019-05-01 · ICLR 2019 5 · AJ Piergiovanni, Michael S. Ryoo

In this paper, we present a method to learn a joint multimodal representation space that allows for the recognition of unseen activities in videos. We compare the effect of placing various constraints on the embedding space using paired text and video data. Additionally, we propose a method to improve the joint embedding space using an adversarial formulation with unpaired text and video data. In addition to testing on publicly available datasets, we introduce a new, large-scale text/video dataset. We experimentally confirm that learning such shared embedding space benefits three difficult tasks (i) zero-shot activity classification, (ii) unsupervised activity discovery, and (iii) unseen activity captioning.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionGeneral ClassificationTemporal Action Localization

Similar Papers 제목 키워드 기반

Learning Multimodal Representations for Unseen Activities

2018-06-21 · AJ Piergiovanni, Michael S. Ryoo

We present a method to learn a joint multimodal representation space that enables recognition of unseen activities in videos. We first compare the effect of placing various constraints on the embedding space using paired…

General ClassificationTemporal Action Localization

Unpaired Speech Enhancement by Acoustic and Adversarial Supervision for Speech Recognition

2018-11-06 · Geonmin Kim, Hwaran Lee, Bo-Kyeong Kim, Sang-Hoon Oh 외

Many speech enhancement methods try to learn the relationship between noisy and clean speech, obtained using an acoustic room simulator. We point out several limitations of enhancement methods relying on clean speech tar…

Generative Adversarial NetworkSpeech Enhancementspeech-recognitionSpeech Recognition

MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-based Image Captioning

2021-12-13 · Wenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang 외

Text-based image captioning (TextCap) requires simultaneous comprehension of visual content and reading the text of images to generate a natural language description. Although a task can teach machines to understand the …

Caption GenerationDescriptiveDiversityGenerative Adversarial Network+2

Improving Alignment and Robustness with Circuit Breakers

2024-06-06 · Andy Zou, Long Phan, Justin Wang, Derek Duenas 외

AI systems can take harmful actions and are highly vulnerable to adversarial attacks. We present an approach, inspired by recent advances in representation engineering, that interrupts the models as they respond with har…

Adversarial Robustness

Unified Attentional Generative Adversarial Network for Brain Tumor Segmentation From Multimodal Unpaired Images

2019-07-08 · Wenguang Yuan, Jia Wei, Jiabing Wang, Qianli Ma 외

In medical applications, the same anatomical structures may be observed in multiple modalities despite the different image characteristics. Currently, most deep models for multimodal segmentation rely on paired registere…

Brain Tumor SegmentationGenerative Adversarial NetworkSegmentationTranslation+1