paper-with-me

홈 › Papers

DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation

2025-03-08 · Runze Zhang, Guoguang Du, Xiaochuan Li, Qi Jia, Liang Jin, Lu Liu, Jingjing Wang, Cong Xu, Zhenhua Guo, YaQian Zhao, Xiaoli Gong, RenGang Li, Baoyu Fan

Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying viewpoints. Prior research, especially in open-source projects, primarily focuses on either temporal or spatial consistency, or their basic combination, such as appending a description of a camera movement after a prompt without constraining the outcomes of this movement. However, camera movement may introduce new objects to the scene or eliminate existing ones, thereby overlaying and affecting the preceding narrative. Especially in videos with numerous camera movements, the interplay between multiple plots becomes increasingly complex. This paper introduces and examines integral spatio-temporal consistency, considering the synergy between plot progression and camera techniques, and the long-term impact of prior content on subsequent generation. Our research encompasses dataset construction through to the development of the model. Initially, we constructed a DropletVideo-10M dataset, which comprises 10 million videos featuring dynamic camera motion and object actions. Each video is annotated with an average caption of 206 words, detailing various camera movements and plot developments. Following this, we developed and trained the DropletVideo model, which excels in preserving spatio-temporal coherence during video generation. The DropletVideo dataset and model are accessible at https://dropletx.github.io.

📄 PDF Abstract BibTeX arXiv:2503.06053

Code (1)

IEIT-AGI/DropletVideo pytorch

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Spontaneous Facial Micro-Expression Recognition using Discriminative Spatiotemporal Local Binary Pattern with an Improved Integral Projection

2016-08-07 · Xiaohua Huang, Su-Jing Wang, Xin Liu, Guoying Zhao 외

Recently, there are increasing interests in inferring mirco-expression from facial image sequences. Due to subtle facial movement of micro-expressions, feature extraction has become an important and critical issue for sp…

Attributefeature selectionMicro Expression RecognitionMicro-Expression Recognition

Nonlocal operator learning for fMRI encoding and decoding tasks

2026-05-19 · Andreas Kramer, Saugat Acharya, Alice Giola, Emanuele Zappala arxiv

Functional MRI data exhibit high-dimensional spatiotemporal structure, making both prediction and decoding challenging. In this work, we investigate neural integral-operator-based models for encoding and decoding tasks i…

Representation Learning

Physics-Informed Neural PDE Solvers via Spatio-Temporal MeanFlow

2026-05-09 · Hanru Bai, Yuncheng Zhou, Difan Zou arxiv

Deep learning paradigms, such as PINNs and neural operators, have significantly advanced the solving of PDEs. However, they often struggle to capture the continuous integral nature of physical systems, relying either on …

Automatic Integration for Spatiotemporal Neural Point Processes

2023-10-09 · NeurIPS 2023 11 · ZiHao Zhou, Rose Yu

Learning continuous-time point processes is essential to many discrete event forecasting tasks. However, integration poses a major challenge, particularly for spatiotemporal point processes (STPPs), as it involves calcul…

Point Processes

Spatio-Temporal Covariance Descriptors for Action and Gesture Recognition

2013-03-25 · Andres Sanin, Conrad Sanderson, Mehrtash T. Harandi, Brian C. Lovell

We propose a new action and gesture recognition method based on spatio-temporal covariance descriptors and a weighted Riemannian locality preserving projection approach that takes into account the curved space formed by …

General ClassificationGesture RecognitionInterest Point Detection