paper-with-me

Papers

Fitting Image Diffusion Models on Video Datasets

2025-09-04 · Juhun Lee, Simon S. Woo arxiv

Image diffusion models are trained on independently sampled static images. While this is the bedrock task protocol in generative modeling, capturing the temporal world through the lens of static snapshots is information-deficient by design. This limitation leads to slower convergence, limited distributional coverage, and reduced generalization. In this work, we propose a simple and effective training strategy that leverages the temporal inductive bias present in continuous video frames to improve diffusion training. Notably, the proposed method requires no architectural modification and can be seamlessly integrated into standard diffusion training pipelines. We evaluate our method on the HandCo dataset, where hand-object interactions exhibit dense temporal coherence and subtle variations in finger articulation often result in semantically distinct motions. Empirically, our method accelerates convergence by over 2$\text{x}$ faster and achieves lower FID on both training and validation distributions. It also improves generative diversity by encouraging the model to capture meaningful temporal variations. We further provide an optimization analysis showing that our regularization reduces the gradient variance, which contributes to faster convergence.

📄 PDF Abstract BibTeX arXiv:2509.03794

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Score-Guided Diffusion for 3D Human Recovery

2024-03-14 · CVPR 2024 1 · Anastasis Stathopoulos, Ligong Han, Dimitris Metaxas

We present Score-Guided Human Mesh Recovery (ScoreHMR), an approach for solving inverse problems for 3D human pose and shape reconstruction. These inverse problems involve fitting a human body model to image observations…

DenoisingHuman Mesh Recovery

OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on

2024-03-04 · Yuhao Xu, Tao Gu, Weifeng Chen, Chengcai Chen

We present OOTDiffusion, a novel network architecture for realistic and controllable image-based virtual try-on (VTON). We leverage the power of pretrained latent diffusion models, designing an outfitting UNet to learn t…

DenoisingImage GenerationVirtual Try-on

LiveSVG: Zero-Shot SVG Animation via Video Generation

2026-05-28 · Matan Levy, Ran Margolin, Bar Cavia, Dvir Samuel 외 arxiv

We introduce LiveSVG, a zero-shot approach for generating Scalable Vector Graphics (SVG) animations using video diffusion models. Current SVG animation methods struggle with complex motions: LLM-based code synthesis fail…

Video Generation

VideoPure: Diffusion-based Adversarial Purification for Video Recognition

2025-01-25 · Kaixun Jiang, Zhaoyu Chen, Jiyuan Fu, Lingyi Hong 외

Recent work indicates that video recognition models are vulnerable to adversarial examples, posing a serious security risk to downstream applications. However, current research has primarily focused on adversarial attack…

Adversarial DefenseAdversarial PurificationAdversarial RobustnessDenoising+1

ODPG: Outfitting Diffusion with Pose Guided Condition

2025-01-12 · Seohyun Lee, Jintae Park, Sanghyeok Park

Virtual Try-On (VTON) technology allows users to visualize how clothes would look on them without physically trying them on, gaining traction with the rise of digitalization and online shopping. Traditional VTON methods,…

DenoisingVirtual Try-on