paper-with-me

Papers

Ani-GIFs: A benchmark dataset for domain generalization of action recognition from GIFs

2022-09-26 · Frontiers of Computer Science 2022 9 · Shoumik Sovan, Majumdar Shubhangi Jain, Isidora Chara Tourni, Arsenii Mustafin, Diala Lteif, Stan Sclaroff, Kate Saenko, Sarah Adel Bargal

Deep learning models perform remarkably well for the same task under the assumption that data is always coming from the same distribution. However, this is generally violated in practice, mainly due to the differences in data acquisition techniques and the lack of information about the underlying source of new data. Domain generalization targets the ability to generalize to test data of an unseen domain; while this problem is well-studied for images, such studies are significantly lacking in spatiotemporal visual content—videos and GIFs. This is due to (1) the challenging nature of misalignment of temporal features and the varying appearance/motion of actors and actions in different domains, and (2) spatiotemporal datasets being laborious to collect and annotate for multiple domains. We collect and present the first synthetic video dataset of Animated GIFs for domain generalization, Ani-GIFs, that is used to study the domain gap of videos vs. GIFs, and animated vs. real GIFs, for the task of action recognition. We provide a training and testing setting for Ani-GIFs, and extend two domain generalization baseline approaches, based on data augmentation and explainability, to the spatiotemporal domain to catalyze research in this direction.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionAnimated GIF GenerationData AugmentationDomain AdaptationDomain Generalization

Similar Papers 제목 키워드 기반

GIFGuard: Proactive Forensics against Deepfakes in Facial GIFs via Spatiotemporal Watermarking

2026-04-29 · Shupeng Che, Zhiqing Guo, Changtao Miao, Dan Ma 외 arxiv

The rapid evolution of deepfake technology poses an unprecedented threat to the authenticity of Graphics Interchange Format (GIF) imagery, which serves as a representative of short-loop temporal media in social networks.…

Video2GIF: Automatic Generation of Animated GIFs from Video

2016-05-16 · CVPR 2016 6 · Michael Gygli, Yale Song, Liangliang Cao

We introduce the novel problem of automatically generating animated GIFs from video. GIFs are short looping video with no sound, and a perfect combination between image and video that really capture our attention. GIFs t…

Happy Dance, Slow Clap: Using Reaction GIFs to Predict Induced Affect on Twitter

2021-05-20 · ACL 2021 5 · Boaz Shmueli, Soumya Ray, Lun-Wei Ku

Datasets with induced emotion labels are scarce but of utmost importance for many NLP tasks. We present a new, automated method for collecting texts along with their induced reaction labels. The method exploits the onlin…

TGIF: A New Dataset and Benchmark on Animated GIF Description

2016-04-10 · CVPR 2016 6 · Yuncheng Li, Yale Song, Liangliang Cao, Joel Tetreault 외

With the recent popularity of animated GIFs on social media, there is need for ways to index them with rich metadata. To advance research on animated GIF understanding, we collected a new dataset, Tumblr GIF (TGIF), with…

Image CaptioningMachine TranslationText GenerationTranslation+1

Cross-Modal Retrieval with Implicit Concept Association

2018-04-12 · Yale Song, Mohammad Soleymani

Traditional cross-modal retrieval assumes explicit association of concepts across modalities, where there is no ambiguity in how the concepts are linked to each other, e.g., when we do the image search with a query "dogs…

Cross-Modal RetrievalImage RetrievalMultiple Instance LearningRetrieval+1