paper-with-me

홈 › Papers

SoftCap: Soft-Budget Control for Diffusion Transformer Acceleration

2026-05-26 · Yuhang Zhang, Junxiang Qiu, Huixia Ben, Zhenhua Tang, Shuo Wang, Yanbin Hao arxiv

Diffusion Transformers (DiTs) achieve strong visual quality, but their iterative denoising process requires many costly Transformer evaluations. Training-free acceleration methods reduce this cost by caching, forecasting, or verifying intermediate features, yet the runtime decision of when to execute a Full step is often driven by fixed schedules or hand-tuned thresholds. We propose \textbf{SoftCap}, a training-free control layer for cache-based DiT inference. SoftCap couples a Trajectory Drift Observer, which estimates local cache risk from lightweight hidden-state statistics, with a Soft-Budget PI Controller, which adjusts the Full-triggering threshold from realized compute relative to a fixed reference profile. The budget is a soft ceiling: it shapes the threshold but does not require a run to spend a prescribed number of Full evaluations. On FLUX.1-dev, SoftCap improves over SpeCa at a comparable middle-compute operating point, raising ImageReward from 0.967 to 0.981 and reducing LPIPS-Full from 0.518 to 0.498 at nearly identical FLOPs, while target-sweep diagnostics show the intended soft-ceiling behavior as the budget is relaxed.

📄 PDF Abstract BibTeX arXiv:2605.27075

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Training Transformers with Enforced Lipschitz Constants

2025-07-17 · Laker Newhouse, R. Preston Hess, Franz Cesista, Andrii Zahorodnii 외

Neural networks are often highly sensitive to input and weight perturbations. This sensitivity has been linked to pathologies such as vulnerability to adversarial examples, divergent training, and overfitting. To combat …

Benchmarking

Budgeted Attention Allocation: Cost-Conditioned Compute Control for Efficient Transformers

2026-05-07 · Amrit Nidhi arxiv

Transformers usually expose one inference cost per trained model, while deployed systems often need multiple cost-quality operating points. We study Budgeted Attention Allocation, a monotone head-gating mechanism conditi…

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

2026-07-27 · Haopeng Li, Yitong Li, Junsong Chen, Tian Ye 외 hf

Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by compu…

Video Generation

Towards Precise Scaling Laws for Video Diffusion Transformers

2024-11-25 · CVPR 2025 1 · Yuanyang Yin, Yaqi Zhao, Mingwu Zheng, Ke Lin 외

Achieving optimal performance of video diffusion transformers within given data and compute budget is crucial due to their high training costs. This necessitates precisely determining the optimal model size and training …

Chest-Diffusion: A Light-Weight Text-to-Image Model for Report-to-CXR Generation

2024-06-30 · Peng Huang, Xue Gao, Lihong Huang, Jing Jiao 외

Text-to-image generation has important implications for generation of diverse and controllable images. Several attempts have been made to adapt Stable Diffusion (SD) to the medical domain. However, the large distribution…

DenoisingImage GenerationText to Image GenerationText-to-Image Generation