paper-with-me

홈 › Papers

Sliding Window Attention for Learned Video Compression

2025-10-04 · Alexander Kopte, André Kaup arxiv

To manage the complexity of transformers in video compression, local attention mechanisms are a practical necessity. The common approach of partitioning frames into patches, however, creates architectural flaws like irregular receptive fields. When adapted for temporal autoregressive models, this paradigm, exemplified by the Video Compression Transformer (VCT), also necessitates computationally redundant overlapping windows. This work introduces 3D Sliding Window Attention (SWA), a patchless form of local attention. By enabling a decoder-only architecture that unifies spatial and temporal context processing, and by providing a uniform receptive field, our method significantly improves rate-distortion performance, achieving Bjørntegaard Delta-rate savings of up to 18.6 % against the VCT baseline. Simultaneously, by eliminating the need for overlapping windows, our method reduces overall decoder complexity by a factor of 2.8, while its entropy model is nearly 3.5 times more efficient. We further analyze our model's behavior and show that while it benefits from long-range temporal context, excessive context can degrade performance.

📄 PDF Abstract BibTeX arXiv:2510.03926

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fast Video Generation with Sliding Tile Attention

2025-02-06 · Peiyuan Zhang, Yongqi Chen, Runlong Su, Hangliang Ding 외

Diffusion Transformers (DiTs) with 3D full attention power state-of-the-art video generation, but suffer from prohibitive compute cost -- when generating just a 5-second 720P video, attention alone takes 800 out of 945 s…

Video Generation

Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation

2026-04-11 · Ruibin Li, Tao Yang, Fangzhou Ai, Tianhe Wu 외 arxiv

Streaming video generation (SVG) distills a pretrained bidirectional video diffusion model into an autoregressive model equipped with sliding window attention (SWA). However, SWA inevitably loses distant history during l…

Computational EfficiencyModel CompressionVideo Generation

Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies

2025-11-02 · Yuxuan Hu, Jianchao Tan, Jiaqi Zhang, Wen Zan 외 arxiv

In this work, we conduct a systematic analysis of Native Sparse Attention (NSA) and propose targeted improvements that enhance long-context modeling. A key insight is that alternating between local (sliding-window) and g…

Long-Context Understanding

WorldKV: Efficient World Memory with World Retrieval and Compression

2026-05-21 · Jung Yi, Minjae Kim, Paul Hyunbin Cho, Wooseok Jang 외 arxiv

Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpoint yields consistent content, remains a…

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression

2026-05-01 · Shen Han, Yuyang Wu, Junpu Yu, Olexandr Isayev arxiv

Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and incurs high decoding latency and limited throughput. To address these issues, KV ca…