paper-with-me

홈 › Papers

Consistent Accelerated Inference via Confident Adaptive Transformers

2021-04-18 · EMNLP 2021 11 · Tal Schuster, Adam Fisch, Tommi Jaakkola, Regina Barzilay

We develop a novel approach for confidently accelerating inference in the large and expensive multilayer Transformers that are now ubiquitous in natural language processing (NLP). Amortized or approximate computational methods increase efficiency, but can come with unpredictable performance costs. In this work, we present CATs -- Confident Adaptive Transformers -- in which we simultaneously increase computational efficiency, while guaranteeing a specifiable degree of consistency with the original model with high confidence. Our method trains additional prediction heads on top of intermediate layers, and dynamically decides when to stop allocating computational effort to each input using a meta consistency classifier. To calibrate our early prediction stopping rule, we formulate a unique extension of conformal prediction. We demonstrate the effectiveness of this approach on four classification and regression tasks.

📄 PDF Abstract BibTeX arXiv:2104.08803

Code (1)

TalSchuster/CATs 공식 구현 pytorch

Tasks

Computational EfficiencyConformal PredictionPredictionregression

Similar Papers 제목 키워드 기반

WorldCache: Content-Aware Caching for Accelerated Video World Models

2026-03-23 · Umair Nawaz, Ahmed Heakl, Ufaq Khan, Abdelrahman Shaker 외 arxiv

Diffusion Transformers (DiTs) power high-fidelity video world models but remain computationally expensive due to sequential denoising and costly spatio-temporal attention. Training-free feature caching accelerates infere…

Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs

2025-09-23 · Marcin Chrapek, Marcin Copik, Etienne Mettaz, Torsten Hoefler arxiv

Large Language Models (LLMs) are increasingly deployed on converged Cloud and High-Performance Computing (HPC) infrastructure. However, as LLMs handle confidential inputs and are fine-tuned on costly, proprietary dataset…

YuriiFormer: A Suite of Nesterov-Accelerated Transformers

2026-01-30 · Aleksandr Zimin, Yury Polyanskiy, Philippe Rigollet arxiv

We propose a variational framework that interprets transformer layers as iterations of an optimization algorithm acting on token embeddings. In this view, self-attention implements a gradient step of an interaction energ…

Fusion: A Framework for Unified Sequential Token AdaptatIon in VisiOn TraNsformers

2026-07-01 · Aravind Pradeep, Samira Nazari, Mahdi Taheri, Christian Herglotz arxiv

Vision Transformers achieve strong image classification accuracy but process all image regions with nearly the same computation, even when many regions are redundant or uninformative. Recent adaptive inference methods re…

Image Classification

Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning

2026-01-06 · Junseok Kim, Nakyeong Yang, Kyungmin Min, Kyomin Jung arxiv

Self-Consistency improves reasoning reliability through multi-sample aggregation, but incurs substantial inference cost. Adaptive self-consistency methods mitigate this issue by adjusting the sampling budget; however, th…