Generating Videos of Zero-Shot Compositions of Actions and Objects
Human activity videos involve rich, varied interactions between people and objects. In this paper we develop methods for generating such videos -- making progress toward addressing the important, open problem of video generation in complex scenes. In particular, we introduce the task of generating human-object interaction videos in a zero-shot compositional setting, i.e., generating videos for action-object compositions that are unseen during training, having seen the target action and target object separately. This setting is particularly important for generalization in human activity video generation, obviating the need to observe every possible action-object combination in training and thus avoiding the combinatorial explosion involved in modeling complex scenes. To generate human-object interaction videos, we propose a novel adversarial framework HOI-GAN which includes multiple discriminators focusing on different aspects of a video. To demonstrate the effectiveness of our proposed framework, we perform extensive quantitative and qualitative evaluation on two challenging datasets: EPIC-Kitchens and 20BN-Something-Something v2.
Code (0)
등록된 구현이 없습니다.
Tasks
Human-Object Interaction DetectionObjectVideo GenerationSimilar Papers 제목 키워드 기반
Zero-Shot Action Recognition from Diverse Object-Scene Compositions
This paper investigates the problem of zero-shot action recognition, in the setting where no training videos with seen actions are available. For this challenging scenario, the current leading approach is to transfer kno…
Action RecognitionObjectTransfer LearningZero-Shot Action RecognitionCompositional Video Synthesis with Action Graphs
Videos of actions are complex signals containing rich compositional structure in space and time. Current video generation methods lack the ability to condition the generation on multiple coordinated and potentially simul…
SchedulingVideo GenerationVideo PredictionVideo-to-Video SynthesisUnified Framework for Open-World Compositional Zero-shot Learning
Open-World Compositional Zero-Shot Learning (OW-CZSL) addresses the challenge of recognizing novel compositions of known primitives and entities. Even though prior works utilize language knowledge for recognition, such a…
Compositional Zero-Shot LearningZero-Shot LearningLearning Attention Propagation for Compositional Zero-Shot Learning
Compositional zero-shot learning aims to recognize unseen compositions of seen visual primitives of object classes and their states. While all primitives (states and objects) are observable during training in some combin…
Compositional Zero-Shot LearningZero-Shot LearningGenZI: Zero-Shot 3D Human-Scene Interaction Generation
Can we synthesize 3D humans interacting with scenes without learning from any 3D human-scene interaction data? We propose GenZI, the first zero-shot approach to generating 3D human-scene interactions. Key to GenZI is our…