paper-with-me

Papers

Temporal Regularization Makes Your Video Generator Stronger

2025-03-19 · Harold Haodong Chen, Haojian Huang, Xianfeng Wu, Yexin Liu, Yajing Bai, Wen-Jie Shu, Harry Yang, Ser-Nam Lim

Temporal quality is a critical aspect of video generation, as it ensures consistent motion and realistic dynamics across frames. However, achieving high temporal coherence and diversity remains challenging. In this work, we explore temporal augmentation in video generation for the first time, and introduce FluxFlow for initial investigation, a strategy designed to enhance temporal quality. Operating at the data level, FluxFlow applies controlled temporal perturbations without requiring architectural modifications. Extensive experiments on UCF-101 and VBench benchmarks demonstrate that FluxFlow significantly improves temporal coherence and diversity across various video generation models, including U-Net, DiT, and AR-based architectures, while preserving spatial fidelity. These findings highlight the potential of temporal augmentation as a simple yet effective approach to advancing video generation quality.

📄 PDF Abstract BibTeX arXiv:2503.15417

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityVideo Generation

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

Intrinsic Temporal Regularization for High-resolution Human Video Synthesis

2020-12-11 · Lingbo Yang, Zhanning Gao, Peiran Ren, Siwei Ma 외

Temporal consistency is crucial for extending image processing pipelines to the video domain, which is often enforced with flow-based warping error over adjacent frames. Yet for human video synthesis, such scheme is less…

Motion EstimationVocal Bursts Intensity Prediction

Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation

2024-09-02 · Qihua Chen, Yue Ma, Hongfa Wang, Junkun Yuan 외

This paper explores higher-resolution video outpainting with extensive content generation. We point out common issues faced by existing methods when attempting to largely outpaint videos: the generation of low-quality co…

GPU

Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models

2023-04-18 · CVPR 2023 1 · Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn 외

Latent Diffusion Models (LDMs) enable high-quality image synthesis while avoiding excessive compute demands by training a diffusion model in a compressed lower-dimensional latent space. Here, we apply the LDM paradigm to…

Image GenerationSuper-ResolutionText-to-Video GenerationVideo Generation+2

Detecting AI-Generated Videos with Spiking Neural Networks

2026-05-07 · Minsuk Jang, Yujin Yang, Heeseon Kim, Minseok Son 외 arxiv

Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detection. Existing detectors typically handle this temporal evidence in three ways: fe…

Dance Your Latents: Consistent Dance Generation through Spatial-temporal Subspace Attention Guided by Motion Flow

2023-10-20 · Haipeng Fang, Zhihao Sun, Ziyao Huang, Fan Tang 외

The advancement of generative AI has extended to the realm of Human Dance Generation, demonstrating superior generative capacities. However, current methods still exhibit deficiencies in achieving spatiotemporal consiste…