paper-with-me

홈 › Papers

Video ControlNet: Towards Temporally Consistent Synthetic-to-Real Video Translation Using Conditional Image Diffusion Models

2023-05-30 · Ernie Chu, Shuo-Yen Lin, Jun-Cheng Chen

In this study, we present an efficient and effective approach for achieving temporally consistent synthetic-to-real video translation in videos of varying lengths. Our method leverages off-the-shelf conditional image diffusion models, allowing us to perform multiple synthetic-to-real image generations in parallel. By utilizing the available optical flow information from the synthetic videos, our approach seamlessly enforces temporal consistency among corresponding pixels across frames. This is achieved through joint noise optimization, effectively minimizing spatial and temporal discrepancies. To the best of our knowledge, our proposed method is the first to accomplish diverse and temporally consistent synthetic-to-real video translation using conditional image diffusion models. Furthermore, our approach does not require any training or fine-tuning of the diffusion models. Extensive experiments conducted on various benchmarks for synthetic-to-real video translation demonstrate the effectiveness of our approach, both quantitatively and qualitatively. Finally, we show that our method outperforms other baseline methods in terms of both temporal consistency and visual quality.

📄 PDF Abstract BibTeX arXiv:2305.19193

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Flow EstimationTranslation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models

2024-07-12 · Jeongho Kim, Min-Jung Kim, Junsoo Lee, Jaegul Choo

Pose-driven human-image animation diffusion models have shown remarkable capabilities in realistic human video synthesis. Despite the promising results achieved by previous approaches, challenges persist in achieving tem…

Image Animation

DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution

2025-07-01 · Zhe Kong, Le Li, Yong Zhang, Feng Gao 외 arxiv

Real-world video super-resolution (VSR) presents significant challenges due to complex and unpredictable degradations. Although some recent methods utilize image diffusion models for VSR and have shown improved detail ge…

Image Super-ResolutionVideo Super-Resolution

HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

2026-03-31 · Mingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee 외 arxiv

Recent methods have made notable progress in the visual quality of hand-object interaction video synthesis. However, most approaches rely on 2D control signals that lack spatial expressiveness and limit the utilization o…

Video Generation

DiVE: DiT-based Video Generation with Enhanced Control

2024-09-03 · Junpeng Jiang, Gangyi Hong, Lijun Zhou, Enhui Ma 외

Generating high-fidelity, temporally consistent videos in autonomous driving scenarios faces a significant challenge, e.g. problematic maneuvers in corner cases. Despite recent video generation works are proposed to tack…

Autonomous DrivingVideo Generation

HARIVO: Harnessing Text-to-Image Models for Video Generation

2024-10-10 · Mingi Kwon, Seoung Wug Oh, Yang Zhou, Difan Liu 외

We present a method to create diffusion-based video models from pretrained Text-to-Image (T2I) models. Recently, AnimateDiff proposed freezing the T2I model while only training temporal layers. We advance this method by …

DiversityVideo Generation