paper-with-me

홈 › Papers

Just-in-Time: Training-Free Spatial Acceleration for Diffusion Transformers

2026-03-11 · Wenhao Sun, Ji Li, Zhaoqiang Liu arxiv

Diffusion Transformers have established a new state-of-the-art in image synthesis, but the high computational cost of iterative sampling severely hampers their practical deployment. While existing acceleration methods often focus on the temporal domain, they overlook the substantial spatial redundancy inherent in the generative process, where global structures emerge long before fine-grained details are formed. The uniform computational treatment of all spatial regions represents a critical inefficiency. In this paper, we introduce Just-in-Time (JiT), a novel training-free framework that addresses this challenge by acceleration in the spatial domain. JiT formulates a spatially approximated generative ordinary differential equation (ODE) that drives the full latent state evolution based on computations from a dynamically selected, sparse subset of anchor tokens. To ensure seamless transitions as new tokens are incorporated to expand the dimensions of the latent state, we propose a deterministic micro-flow, a simple and effective finite-time ODE that maintains both structural coherence and statistical correctness. Extensive experiments on the state-of-the-art FLUX.1-dev model demonstrate that JiT achieves up to a 7x speedup with nearly lossless performance, significantly outperforming existing acceleration methods and establishing a new and superior trade-off between inference speed and generation fidelity.

📄 PDF Abstract BibTeX arXiv:2603.10744

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Training-free Mixed-Resolution Latent Upsampling for Spatially Accelerated Diffusion Transformers

2025-07-11 · Wongi Jeong, Kyungryeol Lee, Hoigi Seo, Se Young Chun arxiv

Diffusion transformers (DiTs) offer excellent scalability for high-fidelity generation, but their computational overhead poses a great challenge for practical deployment. Existing acceleration methods primarily exploit t…

Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models

2026-05-29 · Jiacheng Lu, Haoyi Zhu, Sipei Yi, Enze Xie 외 arxiv

Interactive video world models generate video chunk by chunk in response to user-controlled camera movements, enabling applications such as real-time game simulation, virtual scene navigation, and embodied AI training. H…

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

2025-07-10 · Jeongseok Hyun, Sukjun Hwang, Su Ho Han, Taeoh Kim 외 arxiv

Video large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with token count. To address this, we propose …

Parallel Jacobi Decoding for Fast Autoregressive Image Generation

2026-06-04 · Boya Liao, Ying Li, Siyong Jian, Huan Wang arxiv

Autoregressive (AR) models have demonstrated remarkable performance in generating high-fidelity images. However, their inherently sequential next-token prediction leads to significantly slower inference. Recent studies h…

Image Generation

Training-Free Acceleration of ViTs with Delayed Spatial Merging

2023-03-04 · Jung Hwan Heo, Seyedarmin Azizi, Arash Fayyazi, Massoud Pedram

Token merging has emerged as a new paradigm that can accelerate the inference of Vision Transformers (ViTs) without any retraining or fine-tuning. To push the frontier of training-free acceleration in ViTs, we improve to…

Transfer Learning