paper-with-me

홈 › Papers

Multi-Frame Content Integration with a Spatio-Temporal Attention Mechanism for Person Video Motion Transfer

2019-08-12 · Kun Cheng, Hao-Zhi Huang, Chun Yuan, Lingyiqing Zhou, Wei Liu

Existing person video generation methods either lack the flexibility in controlling both the appearance and motion, or fail to preserve detailed appearance and temporal consistency. In this paper, we tackle the problem of motion transfer for generating person videos, which provides controls on both the appearance and the motion. Specifically, we transfer the motion of one person in a target video to another person in a source video, while preserving the appearance of the source person. Besides only relying on one source frame as the existing state-of-the-art methods, our proposed method integrates information from multiple source frames based on a spatio-temporal attention mechanism to preserve rich appearance details. In addition to a spatial discriminator employed for encouraging the frame-level fidelity, a multi-range temporal discriminator is adopted to enforce the generated video to resemble temporal dynamics of a real video in various time ranges. A challenging real-world dataset, which contains about 500 dancing video clips with complex and unpredictable motions, is collected for the training and testing. Extensive experiments show that the proposed method can produce more photo-realistic and temporally consistent person videos than previous methods. As our method decomposes the syntheses of the foreground and background into two branches, a flexible background substitution application can also be achieved.

📄 PDF Abstract BibTeX arXiv:1908.04013

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Spatiotemporal Object Detection for Improved Aerial Vehicle Detection in Traffic Monitoring

2024-10-17 · Kristina Telegraph, Christos Kyrkou

This work presents advancements in multi-class vehicle detection using UAV cameras through the development of spatiotemporal object detection models. The study introduces a Spatio-Temporal Vehicle Detection Dataset (STVD…

Objectobject-detectionObject Detectionvehicle detection

Spatio-temporal Prompting Network for Robust Video Feature Extraction

2024-02-04 · ICCV 2023 1 · Guanxiong Sun, Chi Wang, Zhaoyu Zhang, Jiankang Deng 외

Frame quality deterioration is one of the main challenges in the field of video understanding. To compensate for the information loss caused by deteriorated frames, recent approaches exploit transformer-based integration…

Instance Segmentationobject-detectionObject DetectionObject Tracking+5

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

2026-03-13 · Yuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu 외 arxiv

Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While current Multimodal Large Language Models (MLL…

Visual Question Answering

M4V: Multi-Modal Mamba for Text-to-Video Generation

2025-06-12 · Jiancheng Huang, Gengwei Zhang, Zequn Jie, Siyu Jiao 외

Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal space remains computationally demanding, pa…

MambaText-to-Video GenerationVideo Generation

LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration

2025-12-26 · Wen Jiang, Li Wang, Kangyao Huang, Wei Fan 외 arxiv

Unmanned aerial vehicles (UAVs) are crucial tools for post-disaster search and rescue, facing challenges such as high information density, rapid changes in viewpoint, and dynamic structures, especially in long-horizon na…

Image Compression