paper-with-me

홈 › Papers

From Hours to Minutes: Lossless Acceleration of Ultra Long Sequence Generation up to 100K Tokens

2025-02-26 · Tong Wu, Junzhe Shen, Zixia Jia, Yuxuan Wang, Zilong Zheng

Generating ultra-long sequences with large language models (LLMs) has become increasingly crucial but remains a highly time-intensive task, particularly for sequences up to 100K tokens. While traditional speculative decoding methods exist, simply extending their generation limits fails to accelerate the process and can be detrimental. Through an in-depth analysis, we identify three major challenges hindering efficient generation: frequent model reloading, dynamic key-value (KV) management and repetitive generation. To address these issues, we introduce TOKENSWIFT, a novel framework designed to substantially accelerate the generation process of ultra-long sequences while maintaining the target model's inherent quality. Experimental results demonstrate that TOKENSWIFT achieves over 3 times speedup across models of varying scales (1.5B, 7B, 8B, 14B) and architectures (MHA, GQA). This acceleration translates to hours of time savings for ultra-long sequence generation, establishing TOKENSWIFT as a scalable and effective solution at unprecedented lengths. Code can be found at https://github.com/bigai-nlco/TokenSwift.

📄 PDF Abstract BibTeX arXiv:2502.18890

Code (1)

bigai-nlco/tokenswift 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Hyperacute pathophysiology of traumatic and vascular brain injury captured by ultrasound, photoacoustic, and magnetic resonance imaging

2022-02-24 · Ali Kamali, Laurel Dieckhaus, Emily C. Peters, Collin A. Preszler 외

Cerebrovascular dynamics and pathomechanisms that evolve in the minutes and hours following traumatic vascular injury in the brain remain largely unknown. We investigated the pathophysiology evolution within the first th…

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

2026-05-21 · Haichen He, Jiayi Zhou, Sifeng Shang, Yihan Hu 외 arxiv

Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video durations. Mastering this intense cognitiv…

GPAIR: Gaussian-Kernel-Based Ultrafast 3D Photoacoustic Iterative Reconstruction

2026-02-03 · Yibing Wang, Shuang Li, Tingting Huang, Yu Zhang 외 arxiv

Although the iterative reconstruction (IR) algorithm can substantially correct reconstruction artifacts in photoacoustic (PA) computed tomography (PACT), it suffers from long reconstruction times, especially for large-sc…

Image Reconstruction

The Backbone Method for Ultra-High Dimensional Sparse Machine Learning

2020-06-11 · Dimitris Bertsimas, Vassilis Digalakis Jr

We present the backbone method, a generic framework that enables sparse and interpretable supervised machine learning methods to scale to ultra-high dimensional problems. We solve sparse regression problems with $10^7$ f…

BIG-bench Machine LearningregressionVocal Bursts Intensity Prediction

Fast-MC-PET: A Novel Deep Learning-aided Motion Correction and Reconstruction Framework for Accelerated PET

2023-02-14 · Bo Zhou, Yu-Jung Tsai, Jiazhen Zhang, Xueqi Guo 외

Patient motion during PET is inevitable. Its long acquisition time not only increases the motion and the associated artifacts but also the patient's discomfort, thus PET acceleration is desirable. However, accelerating P…