paper-with-me

홈 › Papers

Dissecting FLOPs along input dimensions for GreenAI cost estimations

2021-07-26 · Andrea Asperti, Davide Evangelista, Moreno Marzolla

The term GreenAI refers to a novel approach to Deep Learning, that is more aware of the ecological impact and the computational efficiency of its methods. The promoters of GreenAI suggested the use of Floating Point Operations (FLOPs) as a measure of the computational cost of Neural Networks; however, that measure does not correlate well with the energy consumption of hardware equipped with massively parallel processing units like GPUs or TPUs. In this article, we propose a simple refinement of the formula used to compute floating point operations for convolutional layers, called {\alpha}-FLOPs, explaining and correcting the traditional discrepancy with respect to different layers, and closer to reality. The notion of {\alpha}-FLOPs relies on the crucial insight that, in case of inputs with multiple dimensions, there is no reason to believe that the speedup offered by parallelism will be uniform along all different axes.

📄 PDF Abstract BibTeX arXiv:2107.11949

Code (1)

asperti/alpha_flops_dataset 공식 구현

Tasks

Computational Efficiency

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

2024-04-02 · David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap 외

Transformer-based language models spread FLOPs uniformly across input sequences. In this work we demonstrate that transformers can instead learn to dynamically allocate FLOPs (or compute) to specific positions in a seque…

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

2026-05-09 · Kechen Fang, Yihua Qin, Chongyi Wang, Wenshuo Ma 외 arxiv

Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. The prevailing practice typically adopts global encoding followed by …

Dissecting User-Perceived Latency of On-Device E2E Speech Recognition

2021-04-06 · Yuan Shangguan, Rohit Prabhavalkar, Hang Su, Jay Mahadeokar 외

As speech-enabled devices such as smartphones and smart speakers become increasingly ubiquitous, there is growing interest in building automatic speech recognition (ASR) systems that can run directly on-device; end-to-en…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation

2026-01-31 · Yebin Yang, Huaijin Wu, Fu Guo, Lin Yao 외 arxiv

LLMs have traditionally scaled along dense dimensions, where performance is coupled with near-linear increases in computational cost. While MoE decouples capacity from compute, it introduces large memory overhead and har…

Multi-Dimensional Model Compression of Vision Transformer

2021-12-31 · Zejiang Hou, Sun-Yuan Kung

Vision transformers (ViT) have recently attracted considerable attentions, but the huge computational cost remains an issue for practical deployment. Previous ViT pruning methods tend to prune the model along one dimensi…

modelModel Compression