paper-with-me

Papers

Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance

2025-12-08 · Naifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li, Zihan Zheng, Yuan Zhang, Yan Lu arxiv

While traditional and neural video codecs (NVCs) have achieved remarkable rate-distortion performance, improving perceptual quality at low bitrates remains challenging. Some NVCs incorporate perceptual or adversarial objectives but still suffer from artifacts due to limited generation capacity, whereas others leverage pretrained diffusion models to improve quality at the cost of heavy sampling complexity. To overcome these challenges, we propose S2VC, a Single-Step diffusion based Video Codec that integrates a conditional coding framework with an efficient single-step diffusion generator, enabling realistic reconstruction at low bitrates with reduced sampling cost. Recognizing the importance of semantic conditioning in single-step diffusion, we introduce Contextual Semantic Guidance to extract frame-adaptive semantics from buffered features. It replaces text captions with efficient, fine-grained conditioning, thereby improving generation realism. In addition, Temporal Consistency Guidance is incorporated into the diffusion U-Net to enforce temporal coherence across frames and ensure stable generation. Extensive experiments show that S2VC delivers state-of-the-art perceptual quality with an average 52.73% bitrate saving over prior perceptual methods, underscoring the promise of single-step diffusion for efficient, high-quality video compression.

📄 PDF Abstract BibTeX arXiv:2512.07480

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface Networks

2025-03-21 · Bhishma Dedhia, David Bourgin, Krishna Kumar Singh, Yuheng Li 외

Diffusion Transformers (DiTs) can generate short photorealistic videos, yet directly training and sampling longer videos with full attention across the video remains computationally challenging. Alternative methods break…

DenoisingOptical Flow EstimationVideo Generation

A Causal Diffusion Model for Video Reconstruction from Ultra-Low-Bitrate Representations

2026-02-14 · Cem Eteke, Batuhan Tosun, Martin Piccolrovazzi, Alexander Griessel 외 arxiv

We study video reconstruction from ultra-low-bitrate representations, where the primary challenge shifts from encoding to decoding. In this regime, reconstruction with classical and neural codecs introduces blur, while g…

Video Reconstruction

Diffusion Models for Joint Audio-Video Generation

2026-03-17 · Alejandro Paredes La Torre arxiv

Multimodal generative models have shown remarkable progress in single-modality video and audio synthesis, yet truly joint audio-video generation remains an open challenge. In this paper, I explore four key contributions …

Video Generation

DiffVC-OSD: One-Step Diffusion-based Perceptual Neural Video Compression Framework

2025-08-11 · Wenzhuo Ma, Zhenzhong Chen arxiv

In this work, we first propose DiffVC-OSD, a One-Step Diffusion-based Perceptual Neural Video Compression framework. Unlike conventional multi-step diffusion-based methods, DiffVC-OSD feeds the reconstructed latent repre…

Phased One-Step Adversarial Equilibrium for Video Diffusion Models

2025-08-28 · Jiaxiang Cheng, Bing Ma, Xuhua Ren, Hongyi Henry Jin 외 arxiv

Video diffusion generation suffers from critical sampling efficiency bottlenecks, particularly for large-scale models and long contexts. Existing video acceleration methods, adapted from image-based techniques, lack a si…

Video Generation