paper-with-me

홈 › Papers

YODA: Yet Another One-step Diffusion-based Video Compressor

2026-01-03 · Xingchen Li, Junzhe Zhang, Junqi Shi, Ming Lu, Zhan Ma arxiv

While one-step diffusion models have recently excelled in perceptual image compression, their application to video remains limited. Prior efforts typically rely on pretrained 2D autoencoders that generate per-frame latent representations independently, thereby neglecting temporal dependencies. We present YODA--Yet Another One-step Diffusion-based Video Compressor--which embeds multiscale features from temporal references for both latent generation and latent coding to better exploit spatial-temporal correlations for more compact representation, and employs a linear Diffusion Transformer (DiT) for efficient one-step denoising. YODA achieves state-of-the-art perceptual performance, consistently outperforming traditional and deep-learning baselines on LPIPS, DISTS, FID, and KID. Source code will be publicly available at https://github.com/NJUVISION/YODA.

📄 PDF Abstract BibTeX arXiv:2601.01141

Code (0)

등록된 구현이 없습니다.

Tasks

Image Compression

Similar Papers 제목 키워드 기반

Dynamic Attention-Guided Diffusion for Image Super-Resolution

2023-08-15 · Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio 외

Diffusion models in image Super-Resolution (SR) treat all image regions uniformly, which risks compromising the overall image quality by potentially introducing artifacts during denoising of less-complex regions. To addr…

DenoisingImage Super-ResolutionSSIMSuper-Resolution

Regression is all you need for medical image translation

2025-05-04 · Sebastian Rassmann, David Kügler, Christian Ewert, Martin Reuter

The acquisition of information-rich images within a limited time budget is crucial in medical imaging. Medical image translation (MIT) can help enhance and supplement existing datasets by generating synthetic images from…

AllHallucinationImage Generationregression+1

Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video Generation

2023-06-06 · Aram Davtyan, Paolo Favaro

We propose a novel unsupervised method to autoregressively generate videos from a single frame and a sparse motion input. Our trained model can generate unseen realistic object-to-object interactions. Although our model …

ObjectVideo Generation

YODAS: Youtube-Oriented Dataset for Audio and Speech

2024-06-02 · Xinjian Li, Shinnosuke Takamichi, Takaaki Saeki, William Chen 외

In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both lab…

Self-Supervised Learningspeech-recognitionSpeech Recognition

RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping

2025-06-10 · Yang Bai, Liudi Yang, George Eskandar, Fengyi Shen 외

Recent advancements in generative models have revolutionized video synthesis and editing. However, the scarcity of diverse, high-quality datasets continues to hinder video-conditioned robotic learning, limiting cross-pla…

Video Editing