paper-with-me

홈 › Papers

Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs

2025-05-25 · Xuan Zhang, Cunxiao Du, Sicheng Yu, Jiawei Wu, Fengzhuo Zhang, Wei Gao, Qian Liu

Due to the auto-regressive nature of current video large language models (Video-LLMs), the inference latency increases as the input sequence length grows, posing challenges for the efficient processing of video sequences that are usually very long. We observe that during decoding, the attention scores of most tokens in Video-LLMs tend to be sparse and concentrated, with only certain tokens requiring comprehensive full attention. Based on this insight, we introduce Sparse-to-Dense (StD), a novel decoding strategy that integrates two distinct modules: one leveraging sparse top-K attention and the other employing dense full attention. These modules collaborate to accelerate Video-LLMs without loss. The fast (sparse) model speculatively decodes multiple tokens, while the slow (dense) model verifies them in parallel. StD is a tuning-free, plug-and-play solution that achieves up to a 1.94$\times$ walltime speedup in video processing. It maintains model performance while enabling a seamless transition from a standard Video-LLM to a sparse Video-LLM with minimal code modifications.

📄 PDF Abstract BibTeX arXiv:2505.19155

Code (0)

등록된 구현이 없습니다.

Tasks

Video Understanding

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
STD The Spatial-Channel Token Distillation method is proposed to improve the spatial and channel mixing from a novel knowledge distillation (KD) perspective. To be specific, we…

Similar Papers 제목 키워드 기반

Accelerating Prefilling via Decoding-time Contribution Sparsity

2025-07-29 · Zhiyuan He, Yike Zhang, Chengruidong Zhang, Huiqiang Jiang 외 arxiv

Large Language Models (LLMs) incur quadratic attention complexity with input length, creating a major time bottleneck in the prefilling stage. Existing acceleration methods largely exploit attention score sparsity by est…

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload

2026-05-19 · Zhiben Chen, Youpeng Zhao, Yang Sui, Jun Wang 외 arxiv

Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive (AR) models, offering better hardware utilization and bidirectional context through parallel block-level decoding. Howev…

Free Lunch in the Forest: Functionally-Identical Pruning of Boosted Tree Ensembles

2024-08-28 · Youssouf Emine, Alexandre Forel, Idriss Malek, Thibaut Vidal

Tree ensembles, including boosting methods, are highly effective and widely used for tabular data. However, large ensembles lack interpretability and require longer inference times. We introduce a method to prune a tree …

FREE: Uncertainty-Aware Autoregression for Parallel Diffusion Transformers

2025-11-25 · Xinwan Wen, Bowen Li, Jiajun Luo, Ye Li 외 arxiv

Diffusion Transformers (DiTs) achieve state-of-the-art generation quality but require long sequential denoising trajectories, leading to high inference latency. Recent speculative inference methods enable lossless parall…

Just-in-Time: Training-Free Spatial Acceleration for Diffusion Transformers

2026-03-11 · Wenhao Sun, Ji Li, Zhaoqiang Liu arxiv

Diffusion Transformers have established a new state-of-the-art in image synthesis, but the high computational cost of iterative sampling severely hampers their practical deployment. While existing acceleration methods of…